Showing posts with label search engines. Show all posts
Showing posts with label search engines. Show all posts

Tuesday, 8 September 2020

Monday, 14 May 2007

, ,

Uninventing the search engine

Don't you hate it when stuff just works? It's predictable and boring, and if you ask me, anything falling into this category should be sabotaged immediately to spice things up a little.
Many web coders clearly share my view because this is precisely what they've been doing with their previously accurate, efficient and dependable search engines.

Take Googles' Image Search for example. Imagine your typical day; you're surfing the web when a sudden impulse to track down a picture of Spider-man wrestling a T-Rex grips you with full force. You visit Google Image Search and type in the keywords 'spiderman', 'wrestles' and 'trex'. Now you wouldn't imagine there would be all that many depictions of such a scene so it would be reasonable to expect a return of say less than half a dozen hits at the most. Well you'd be wrong; supposedly Google currently indexes 1030 images of the web-shooting wonder getting down and dirty with the "last and largest known carnosaur".

It's curious that amongst these 'hits' are images of King Kong, random politicians, The Simpsons, Bambi, fish corpses and Wacko Jacko's face embedded in a slice of toast, but none of them remotely resemble what I actually searched for.

The fact that there are lots of pictures containing isolated wrestlers, spidermen and dinosaurs might indicate that Google has applied the OR Boolean search operator to my query rather than the more useful AND one. This isn't the case, however; if you click on the 'Advanced Image Search' link you'll see that the keywords are automatically entered into the "find results related to all of the words" box to demonstrate which kind of search I performed prior to reaching this page. Just to confirm, clicking the search button again at this point returns exactly the same set of irrelevant flotsam.

It could be that I'll never ascertain for certain if Spider-Man (yes, I know that's the correct way to write it) ever unleashed the Pumphandle Michinoku driver II on a 43 foot long, 7.5 tonne 'tyrant lizard king'. It's no laughing matter.

That's just a drop in the ocean. All kinds of search engines across the board are falling prey to Boolean vandalism; software repositories, forums, recipe databases - the list is endless. The digg coders are prime suspects. Try probing it for stories involving two of the widest prevailing bedfellows, the 'llama' and 'blamange'. Go on, guess how many hits there are for this keyword combo. 89! That's eighty-nine, EIGHTY-NINE!

It appears that contrary to the norm, you can crowbar the AND operator in between them to narrow down the field, but why wouldn't this be the default setting to begin with as it is with Google? (well, the text search element of Google anyway). You wouldn't dial 999 to report a crime and when asked, "which service do you require" reply police... OR a florist please, either will do. So why would it make sense in any other context?

Sunday, 17 December 2006

,

Filtering Google search results by date range

Supposedly we are able to use the operator 'date:3/6/9/12' to limit search results to only those added to Google's index within the last 3 months, 6 months and so on. In practice you may as well not bother because all this tweak does is return pages which include keywords such as "Date: 12 December 2006". Chocolate fireguard anyone?

An alternative, undocumented, super-secret operator you can use is 'daterange:[julian date]-[julian date]'. Huh? As defined by Wikipedia: "The Julian day or Julian day number (JDN) is the number of days that have elapsed since 12 noon Greenwich Mean Time (UT or TT) on Monday, January 1, 4713 BC in the proleptic Julian calendar . That day is counted as Julian day zero. The Julian day system was intended to provide astronomers with a single system of dates that could be used when working with different calendars and to unify different historical chronologies."

Right so Stephen Hawking has his bases covered, but how are the rest of us going to do the maths in our heads? We don't need to. We can use the Gmacker date range search, which will plug in the correct calculations automagically based on the number of days or dates entered. The problem is, using the daterange operator doesn't make a scrap of difference to your search results either. Great tip this is turning out to be, eh! I hope the likes of Likehacker are taking note. The recipe for a top tech tip: identify problem, offer solution, decide solution is rubbish and shrug shoulders.

Take a major, recent news story, for example, and apply the only-the-last-30-days modifier to the keywords entered: ipswich prostitutes "serial killer" "paula clennell" daterange:2454055-2454085. Now try the same search without the daterange operator. Either way you get 44,500 hits. That's precision fine-tuning at work.

Actually I shouldn't call the victims 'prostitutes' so we're told by the politically correct, feminist mob because it belittles the tragedy and demeans the women involved. According to these pedants it isn't useful to identify them in this way so that other sex workers will know to be wary, employ safety-in-numbers tactics, or get off the streets altogether. Also it doesn't help the police to be able to draw correlations between the targets enabling them - with the help of criminal psychologists - to build a profile of the killer.

They argue that if all the victims had been McDonald's employees, this facet of the case wouldn't have featured so prominently, or received so much media attention. I think the rest of McDonald's staff working in the area would beg to differ.

One commentator ratcheted the farce up another notch when she tried to sugar-coat the reasons some prostitutes were still walking the streets in Ipswich despite the heightened risks: like any other doting mothers they need to put in extra hours at this time of year to be able to afford Christmas presents for their children. Paints a cosy picture doesn't it, but in reality most of them are compelled to put their lives in jeopardy to feed their addiction to hard drugs. According to the BBC's victim profiles page, only one of them was a mother, and a heroin user.

Of course the sum of these women's lives shouldn't be defined solely by their chosen career path, but surely a dead spade is still a spade? Why does truth have to be the casualty of news reporting in this era of politically correct doublespeak?

Friday, 10 November 2006

, ,

Taking out the G-Trash

It's funny how you can put up with niggling annoyances and learn to muddle through, and then as soon as you throw a tantrum a solution presents itself entirely out of the blue.

Today I stumbled upon what is known as Gumbmug, phonetically speaking. No idea where the name comes from, but being down with the latest web trends I'd guess it stems from the tendency to drop the vowels from words so you've got more chance of bagging a unique trademark. Well whatever, what it does is gives you back your Google by blacklisting notorious e-drool such as Shopbot, Dooyoo and so on, thereby tipping the spam-to-genuine-content ratio in your favour.

Seasoned Googlers will know you can achieve the same thing with Google Classic all on your lonesome, but then who wants to append "-inurl:(kelkoo | bizrate | pixmania | dealtime | pricerunner | dooyoo | pricegrabber | pricewatch | resellerratings | ebay | shopbot | comparestoreprices | ciao | unbeatable | shopping | epinions | nextag | buy)" to every single search query? This by the way is advanced Google operator shorthand for 'don't link me to any sites which contain these words in the web address'. Even keeping this string close at hand for copy/paste purposes is no substitute for Gumbmug seeing as a static list wouldn't take into account the emergence of new webscurge upstarts, or remove banished sites if they one day decided to provide information that anyone cared about.

Let's have a tinker then shall we. A search for "wireless mp3 player" returns 91,300 results in plain old Google, while the same search generates only 25,400 hits via Gumbmug. Eureka, that's what I call progress! I have a new home page. In the rare event of actually wanting to run a price comparison check, I'll pick one 'screen scraper' and visit it directly. They're extremely useful in the right context of course.

I'm not usually one to lose myself in a tirade of strong language, but gosh darn it, sometimes a webapp gets me so excited I just can't help it. My apologies for the four letter words.

Thursday, 9 November 2006

, , , ,

MiggyTrack


When it comes to offering personalised search tools, Rollyo are no longer the only game in town. Google are hungry for a slice of the pie and aim to claim a sizeable portion by way of their newly uncorked Co-Op web app.

With Co-Op you get pretty much the same deal, except under the bonnet (or 'hood' I suppose :p) you'll find Google's own search engine rather than Yahoo's, you're supplied with a wacky, instantly forgettable URL to link to your widgets and you're given more scope to categorise your web-foraging offspring.

To check out the hue of the grass on the other side of the fence I've thrown together a custom search widget which queries eleven of the top-ranking Amiga game database web sites. Not that I'm obsessed or anything silly like that.

I was pleasantly surprised to find it's a nice shiny emerald green. The start pages are distinctly uncluttered as you'd expect from a Google Gooey, you can opt to eradicate all adverts (as long as you're not operating as a commercial organisation) and there are plenty of advanced options to keep the most demanding tweakers happy. For now the race is too close to call. Do the front-runners have any competition?

Wednesday, 8 November 2006

, , ,

My First Search Engine. Porn and spam sold separately.

Statistics show that 98.54% of the content on the internet is worthless dross, yet we still have to wade our way through it to get to the good stuff. Perfect example: whenever you search Google to locate a trustworthy review of a piece of tech gear you're considering purchasing, it spews out wads of irrelevant shopping spam sites which - purely by chance of course - contain the word 'review', even though no opinions, positive or negative, are imparted within their pages. Typically this fluff populates the first few pages of Google's output, pushing the genuine content deep into obscurity.

One workaround would be to identify a handful of reliable sources for each kind of information you require, bookmark and search them individually. Better still is Rollyo; a newish, startup web gizmo which provides the means to tailor your search results to suit your personal preferences. It does this by allowing you to 'Roll Your Own' categorised search filters. For instance, you could create a health 'Searchroll' by entering the URLs of up to 25 top-rated health-focused web sites, which when queried would only return content produced by these previously vetted sources.

Wednesday, 12 October 2005

, , ,

Google launches bootleg search engine

Three and a half seconds ago (if you could arrange for your jaw to drop in awe of my finger-on-the-pulseness it would be much appreciated, thanks) search engine monolith, Google, unveiled the latest widget in their web-taming repertoire.

Google Swag Bag (TM) allows users to locate no-nonsense index listings of illegal booty such as MP3 music files and movies. The service operates by tapping into Google's traditional search engine technology while excluding common web site documents with the extensions html, htm, php, asp and so on to return results consisting only of binary content, aka just the juicy stuff.

To those of you apt to perusing Google's search modifier cheat sheet, this is old news; this feat has previously been accomplished by entering little known combinations of operator strings into the standard Google search box. What's new is the user-friendly, streamlined interface which takes the hassle out of digging for multimedia content.

By default Swag Bag forages for MP3 files while filtering out distracting text documents of various formats and keyword red herrings. To re-focus its search beam you can click on the 'toggle advanced settings' link and check/uncheck the boxes adjacent to the media type(s) you would like Google to ferret out.

Disclaimer: Kookosity does not endorse piratey shenanigans. If, as a direct consequence of reading this post, your soul is irreparably corrupted leading to eternal damnation (and singed eyebrows), it's not my fault. God made me do it. Hey, if it's good enough for George W., it's good enough for me.

Thursday, 25 August 2005

, ,

Google fills in the blanks

I couldn't tell you if this search refining feature is new, or just new to me, but it's one well worth adding to your info mining arsenal.

If you want Google to forage for a particular phrase, though can't bring it to mind in its entirety, you can replace the tip-of-the-tongue, missing words with stars and let Google fill in the blanks. This might be a useful way to look up song lyrics, amongst other things. Note that you aren't required to enclose the words in speech marks to instruct Google to search for them in the order they were entered.

Stars are useful for finding quick answers to concise questions too: try entering the text the capital of paraguay is * and you'll be told in no uncertain terms - two million times no less - that the answer you seek is 'Asuncion'. Great for pub quizzes then... if you happen to have a laptop with you, and your local boozer is equipped with WAP, and the other participants are too drunk to notice you furiously bashing away at your keyboard, coincidentally right after each question has been posed.

You could also use stars to quickly assess the general consensus of opinion on any given topic. For instance, if you submitted the text george bush is an * you might be given the impression that Darth Bush isn't exactly dynamite in the popularity stakes - in fact you'd have to click through to page five before you struck upon a positive adjective... and even these ones look conspicuously sarcastic/ironic.

Other Google search modifiers to have escaped my notice until now include:-

filetype: (or ext:) - extremely useful for tracking down PDF journal articles or technical manuals e.g. ipod user manual filetype:pdf

allintitle: - limits search results to those containing your specified keywords in the title of the pages e.g. allintitle:charlie and the chocolate factory

allinurl: - limits search results to those containing your specified keywords in the web address e.g. allinurl:extras ricky gervais

Still thirsty for more? Try Google Guide's advanced operators reference page.

Tuesday, 20 July 2004

, , ,

Forum Googleism

If you're looking for quick answers, the best place to find them is on web forums because the questions to those answers are likely to have already been asked a multitude of times. We all have our favourite forums bookmarked and get into the habit of returning to the same ones for information, though being this selective severely limits the scope of the resources available.

Typing the same query into the search box of each one sequentially isn't practical, which is why Board Reader is such a miraculous web widget. Board Reader's spiders tirelessly crawl the web looking for vBulletin and UBB forums. Forums of all shapes and sizes are unearthed, their contents are indexed and then made available to visitors of the Board Reader web site via the search box. Relevant hits are displayed in an easy to read, uniform fashion as with more traditional search engines, and it even provides cached versions of the pages in case the original sources have been moved, deleted or are temporarily unavailable.

Thursday, 15 November 2001

, ,

Honing your search engine technique

Be more specific when using web search engines. First of all make sure you're using Google as it's undoubtedly the most comprehensive search engine ever to have existed, and what's more, the hits it returns are actually relevant to your search queries - surprisingly a feature which all too many search engines lack!

Now that Google is set as your home page remember to use Boolean queries whenever you use it to search the web. A Boolean query is a logical term or operator, which can be added to your keywords or phrases to refine your search. The most common ones include the words AND, OR and NOT. It's worth remembering though that some search engines will allow you to replace the words AND and NOT with the symbols + and -, and that Google dispenses with the + operator altogether as it assumes you want all the keywords entered to be included in the results - makes perfect sense if you ask me. Notice that these are written in uppercase. I haven't accidentally left my caps lock button turned on; I've done this deliberately to indicate that this is the way they must be typed into your search engine.

As already noted, the AND operator is largely superfluous today, but to demonstrate how this would work in less advanced search engines, consider the following example. If you were looking for information regarding the author Joseph Heller (you know, the guy who wrote Catch-22?), you might try typing only the two words comprising his name into a search box. The hits your search engine returned would lead to any web pages containing either the word 'Joseph' or the word 'Heller', but not necessarily ones that contain both.

In effect, you could be directed to sites that revolve around John Heller or Joseph Smith (whoever they are!). If you wanted to filter out all the irrelevant hits you could use the AND operator. This ensures that the keywords you specify must both appear in the search results (or hits). Any sites containing just one of the terms will be ignored. To do this you would type Joseph AND Heller into your search engine. Try it now and you'll see how effective this can be.

If you don't want to rule out too many possible matches you could use the OR operator instead. Typing in two keywords separated by the OR operator will return web pages that contain either one of your keywords. Again, this switch is fairly redundant in most cases because it will already be the default setting, but since most search engines support it I thought I'd give it a mention.

The NOT operator is much more useful. This can be used to specify words which should not appear in the results (I apologise if this is already obvious!). For example, typing in soap operas NOT Neighbours will give you a list of all the web pages that concern soap operas, except the ones about Neighbours in particular. So you may be directed towards Home and Away, Days of Their Lives or Family Affairs fan sites. Disclaimer: my knowledge of these program's existence does indicate that I watch them, thank you very much. ;)

Perhaps the most effective measure you can take to refine a search query is to use phrases enclosed by inverted commas. For instance, one way to find a selection of Monkey Island related sites would be to type "The Curse of Monkey Island" or "Escape from Monkey Island" into your favourite search engine (the case isn't important, I just like to be grammatically correct). This will only return web pages that include all the keywords listed in the correct sequence, i.e. any sites containing just one or two of these keywords, or containing all of the words, but in a different order will be filtered out.

For more search engine fine-tuning tips refer to the official Google cheat sheet.

So what are you waiting for? Go and practice with Google. Remember, 'tis all in de Booleans mon!... or something like that.

Sunday, 9 July 2000

,

One search engine or ten? Why settle for half measures?

When searching for software, clip art, MP3s, or anything for that matter, use a 'meta' search engine. Meta search engines take a query and submit it to a multitude of diverse search engines simultaneously, amalgamate the results and then present them to you in a logical, standardised format.

A single search engine cannot possibly index everything the web has to offer and as a result they miss many relevant hits. Because meta search engines have access to the leading search engines they are able to offer much more comprehensive results than any single search engine.

The number of search engines utilised by meta search engines varies considerably, but obviously try to find one which searches the highest number of portals and will therefore harvest the most results. Metacrawler crawls metas like no other if you're open to suggestions.