Spdiering the Internet 07 Dec 03

I have started to document what I have been doing to construct the spiders. It is not really a tutorial, it”s more about what I did and how I did it. I doubt it is even close to how it should be done but I am enjoying doing it and I get to research some interesting areas of information retrieval an processing while doing it so what the hell.

Finishing the Robots 03 Dec 03

I have been quite busy lately dipping my toes in various waters hence the lack of entries just lately. I have actually finished with the robots for the short term and now moved onto the Search engine part of the project.
I am enjoying building the search engine because I get to work with C++ again, which is another language I enjoy using. I like it because I feel as if I am as close to the hardware as I am when using C but have various High Level Tools at hand when I need them. I picked C++ over C because it has the STL which I have used before. I imagine that most commercial search engines are using either C or C++ for

Off to wales 20 Nov 03

I have left the robots gathering links and pages for the last few days and the results are as follows. I am off for a couple of weeks one of which will be spent in a cottage in Wales which should be nice so there will be little or no action for quite a while here.
62.0 Million links found
11.9 Million unique links found

Weeding the database 12 Nov 03

You will see that the database has been reduced in size quite a bit. I have been running out of space so I decided to do some weeding. What I have done is fix all all the Url’s that had a fragment part. Url’s come in the following format.

The fragment part of the URL is not really required by us because it indicates a position in a document. This level od granularity is not required or any use to us, we are only interested in the document itself. I wrote a simple Perl script in conjunction with a Postgres Function to weed these out. During the process I deleted all links that where found by following the original URL with the fragment. This is what has led to the reduction in total links found. If you have a look at the latest robot code you will see that I now cater for this fragment art and strip it off before requestint the document.
55.0 Million links found
11.9 Million unique links found

Re-writing the spiders 08 Nov 03

I have been very busy lately re-writing the spiders for the search engine. I have decided to write up what I did to build the spider in the vain hope that someone may find it useful one day. I digressed several times and had some fun writing a recursive one but I eventually settled on writing an iterative robot that uses Postgres to store the links. This was partly due to already having a database with several million links in it already. Please see the link above for more details. I have also managed to download a few thousand documents for the search engine, hence the increase in the links found, this was caused by me parsing the documents that I had found when experimenting with the new robots .
85.0 Million links found

Sorting out the code 27 Oct 03

The last couple of days have been spent sorting out some of the perl and C++. I have also expanded the stop list quite a bit. The Perl script that I was using to produce the file to build the term document matrix also got a bit of a working over.
I have increased the document list to 15700 which is still relatively small for an internet search engine but it is now a respectable amount of text to search for a small intranet site, like a small law firm. I will gradually increase this as I go along and testing to see what kind of results I get.
I have decided to write up what I have done with example code and put it on another few pages. Hopefully someone will be able to make some use of it.
Please see my:
Vector Space Search Engine
page for more details of what I have done to get this working, I am using the term working in its weakest sense here since I have been unable to test it properly yet.
83.1 Million links found
10.9 Million unique links found

Getting the vector space search engine runing 25 Oct 03

I have spent the last few days trying to get the Vector Space Search engine running. The code is in a bit of a mess at the moment but, it’s comming along. All I can say is thank god for the STL, without it I would been in for a hell of a job. I have now managed to create a Sparse Vector Space Matrix from 2397 documents. This needs to be increased before I can really start testing any weighting algorithms.
At the moment this is using 30Mb of memory. This is the max used during the entire process. I did have it running at 256Mb but this was my first round at designing the matrix. I then showed my program a copy of Knuth volume 3, it cowered in fear and its shoesize quickly dropped to a more respectable size. I am pretty sure that I could drop this even further by writing my own data structure without using the STL but I am happy with it at the moment.
I am not entirely happy with the output of the program yet becasue the inner product routine is not producing the correct output but this should be reletively easy to fix. I really need to do a review to make sure that I am not missing any points in my methodology.
I also need to try and compile a better stop list. The one I am using is not particularly good. This is a sure fire way of reducing the RAM footprint.
83.1 Million links found
10.9 Million unique links found

Joys of DOS 5 20 Oct 03

Well I managed to get to play with DOS 5 today to get the PC working, don’t ask. I then had trouble trying to get the data off the old drive that was in the original PC because it insisted that ‘it’ was the operating system and that it was going to boot no matter where on the IDE channels you put it. It also seemed to ignore various jumper settings which was very odd. A quick format /mbr taught it who the boss was and we had no more trouble from the cheap seats.
After not having much success with www.realemail.co.uk due to spurious line disconnections we eventually got online with uklinux which was an ISP I used to use when Jason Clifford ran it. Unfortunately he is no longer concerned with it for various reasons, but he has started www.ukfsn.org which I can highly recommend.
I then downloaded ZoneAlarm and AVG for them. Unfortunately I forgot my Video cable or they could have been using a 19″ Iiyama Vision Master instead of the old 14″ jurassic valve driven thing they had been using previously.
35.876 Million links found
5.476 Million unique links found

20 Oct 03

After all the shuffling that I did two days ago to get the 20Gb disk out of the machine, I now need to check all the Postgres files to make sure that I have them in all the correct places and restart the robots to get the count up a little. Moving Postgres File
35.876 Million links found
5.476 Million unique links found

C++ Libraries for sparse matrix manipulations are not easy to find 20 Oct 03

After much hunting for a library that I can use to impliment an LSI search engine I have had little luck. The library that seems to be the job for this is called SVDPACK it is written in Fortran and has been ported to C++. However, it has not been ported to the humble x86 architecture. It looks like I will have to run with writing the Vector Space search engine instead.
I managed to write a C++ routine to get the output of the Perl Term Document parser. This is a very simple parser that splits all the words in the document on whitespace. I know that there are reasons for not doing it like this so if I get time I will come up with a better method later but for now it will do.
My next task now is to take the input of the C++ program and create a Term Document Matrix from that that I can manipulate easily. I need to be able to carry out the following actions and quite a few more.
1. Count all occourances of each word in the entire document set.
2. Calculate mean values for each word.
3. Come up with some method to ran words in the document matrix. This is to avoid the typical abuses that you see where website saturate pages with keywords to try and manipulate the results of a search engine.
67.6 Million links found
8.529 Million unique links found