Semji

Updated

Search engines: how do they work?

Graphic illustrating a search engine

Thanks to search engines, people can research a topic, find a restaurant, or pick out a pair of shoes in a single click. These tools give access to a huge volume of information through all kinds of queries. But do you know how a search engine actually works? This article walks through how Google works, then Qwant (a French search engine), Bing, and DuckDuckGo, before closing with a snapshot of the search market.

What is a search engine?

Definition of a search engine

Search engines are web applications built to search the web. Results appear according to the phrases users type. Google is still the best known of all. But there are plenty of capable engines: DuckDuckGo, Bing, Qwant, Yahoo. You will also find specialized engines: Google Scholar for education, Yahoo Kids for children, Ecosia for the environment.

What can you find with search engines?

You can use search engines to look things up in a specific field. Google, for example, can surface informative web pages, images, ecommerce listings, documents, or videos.

The Maps feature acts as a world map and uses satellite imagery to pinpoint a location. Other alternatives to Google such as Bing, DuckDuckGo, Yahoo, or Qwant are powerful search engines too. Each one leads with a different argument. Qwant, for instance, is a search engine that respects privacy. It does not try to find out who you are or where you are in order to serve results.

How do you access a search engine?

Search engines are reached through a browser. Most browsers now use an omnibox so the user can run a search. Omnibox is the new name for the old address bar.

How do you succeed with search engines?

That is the question every SEO agency asks. Ranking in the first results of the SERP has a real financial stake. SEO is both a strategic and a technical discipline. A few simple optimizations still help search engines look on you kindly.

You can, for example:

  • Avoid “occultation.” Better known as “cloaking,” this means showing one page to Googlebot and another to human visitors in order to rank higher. The web server is programmed to serve a different page depending on who made the request (Google’s robot or a person).

  • Build a site with a clear hierarchy and provide a sitemap.

  • Create relevant internal linking.

  • Build a useful site that is rich in information. Structure your content with H1, H2, and H3 tags. Your tags and your copy should include the keyword you want to rank for in the SERP.

  • Do not neglect ALT attributes and meta descriptions. These should be precise and contain your keyword.

  • Earn quality backlinks to give your site more authority.

What Bing suggests:

  • Put your keywords in the URL. Domain names that contain keywords seem to sit well with Bing. Keep URLs short. Bing leans on user experience. A short URL is easier for people to remember, so Bing will tend to favor your ranking if you follow that logic.

  • Build relevant internal linking.

  • Run a quality backlink strategy. As with Google, the more your links come from well-known sites, the stronger your site’s SEO.

  • Create topic-focused content. Bing’s principle is simple: one page = one topic!

  • Keep the site architecture clear and simple. Bing recommends 3 levels. Pages should not sit more than 3 clicks from the homepage. As with Google, a sitemap is recommended.

  • Make sure content is not buried in rich-media systems (Adobe Flash Player, JavaScript, Ajax, and so on).

  • Avoid putting keywords only inside images. If you want your company name indexed, make sure it is not displayed solely inside your logo.

Graphic for a search engine

How search engines work

Crawling and indexing

Search engines exist for one job: answering users’ questions. To return relevant results, they go through two stages:

  • crawling: finding pages on the internet

  • indexing: ranking results by relevance

Crawling

Crawling is the first job of a search engine. It is a systematic inspection of websites on the internet. Done before the user’s query, this step gathers as much information as possible from web platforms. It is carried out by robots called “spiders” or “crawlers.” At the end of this step, they send the collected information to the index so it can be indexed.

Indexing

When the index (the engine’s brain) receives the information from the robots, it evaluates it. That way, every time a user searches, the engine can serve relevant results.

How do search engines decide a result is relevant?

Evaluating relevance is not just matching the query to a web page. Other factors come into play. Search engines assume that the more popular a site is, the more relevant the information it holds. That assumption is how they try to keep users satisfied with the results.

Myths and reality around search engines

The myths

Submitting to search engines In the 90s, search engines used submission forms. Webmasters submitted their sites and their keywords. The idea was to flag the site so the engines would crawl and index it. That system was quickly revised and then dropped. Today the robots come on their own to crawl sites and index them on key phrases.

Ranking by meta tag Meta tags (especially the meta keywords tag) used to be crucial for ranking. Every major engine has abandoned that criterion. Meta tags no longer have any effect on rankings.

Paid search (SEA) pushes pages to the top of the SERP Some theories claim that sites paying for search ads (SEA) automatically rank better organically. That assumption has no basis. Google, Qwant (a French tool), and Yahoo have even put safeguards in place to stop that kind of talk. At Google, advertisers who spend millions of dollars a month on ads have noticed they get no special treatment from the search engine.

If those are the myths, what are the realities?

The reality

The crawl budget The web holds trillions of pieces of data. To make the robots’ work manageable, search engines cap how much they crawl. Crawl budget is the time the robots grant your site. Search engines need to find your pages as quickly as possible. That is a real stake. You have to make the robots’ job easier so they can crawl and index as much of your site as they can. If they cannot, part of your site will stay invisible to the engines and to users.

To make the robots’ job easier, you can already apply a few good practices:

  • Avoid broken links. Robots dislike them. They may stop the crawl.

  • Avoid low-quality content. Error pages, duplicate content, faceted navigation, and the like.

  • Limit 301/302 redirects

  • Optimize your page load time. A long load time is bad for your SEO, and for the visitor too. They will tend to go to another site to answer their query if yours takes too long. You lose prospects that way.

  • Keep your sitemap up to date. It will guide the robots more easily when they index your pages.

A regular crawl of your site You have just launched your site and you see it indexed in search engines. You figure the work is done? The robots come back regularly. A site that is updated often will see robots more often than a static one. Every day, search engines run a keyword analysis of pages in order to index them.

Cloaking detection Cloaking means showing different content to search engines and to visitors. The server recognizes whether a person or a robot made the request. On that basis it serves different content. For a robot, it might serve a more optimized page that would be unpleasant for a person to read. Google penalizes this technique.

Filtering low-value content Engines all use robots to judge the added value of a piece of content for readers. The content most often filtered is

  • affiliate content,

  • duplicate content

  • generated pages with very little text.

Engines evaluate a domain on its originality and on the visitor experience it offers. Sites that publish poor-quality content will struggle to rank at the top, even if they are otherwise well optimized. If you have a high bounce rate from the SERP, for example, the engines will demote you. It means people are not finding an answer to their query and the content is not relevant.

The launch of Google Panda in 2011 also showed the engine’s intent to reward quality content. The algorithm was introduced after a large wave of spam and low-quality sites. How does the penalty work? Panda penalizes poor-quality content and sometimes the whole site. The pages in question are then deindexed.

Ranking based on the trust your site earns Several signals are used to evaluate your site and place it in the SERP. One important criterion for the search engine is backlinks. To measure how reliable your site is, Google looks at the number of links pointing to it. In short, the engine treats your site as relevant because many other sites point to it. The search engine does not measure quantity of backlinks alone. The quality of those links is essential. The more your links come from authority sites, the more the engines will favor you. If you have spammy, low-quality links, the search tool will treat that as fraud and apply a penalty. The Penguin algorithm was created to clean Google’s indexes of low-quality sites that game SEO with fraudulent linking techniques.

The fight against search-engine spam Spam is very present on the web. Rising since the mid-1990s, the practice lets spammers take over well-ranked sites in order to promote low-quality destinations. A single day at the top of Google can bring in as much as €20,000 in net revenue. With stakes like that, it is no surprise the practice became so popular. Today, thanks to Google’s technical progress, it is harder and harder to pull off.

Search engine news

Who holds the largest share of the global search market?

The 2017 global ranking put Google first with a net share of 74.54%. It was followed by the search engines Yahoo, Baidu, Bing, and Qwant (a French search tool), whose shares sit around 7 to 10%. It is worth noting that even though Google holds the largest share, that share slowly declined from the second quarter of 2017, while Baidu’s share reached 14.69%.

How many searches are run each day on search engines?

In 2017, 46.8% of the world’s population had access to the internet. By 2021, that figure was expected to reach 53.7%. According to the statistics, Google receives 3.5 billion queries a day, or 1.2 trillion a year. Google moves fast. In 1999 it took Google a month to crawl and index 50 million pages; in 2012 that task was done in less than a minute.

Search engines are therefore powerful, complex applications. Every day, millions of queries are typed by users. Beyond the information stake, search engines also carry a marketing and financial one. To face competition and generate revenue on the web, ranking well in the SERP is essential. Knowing how your audience actually searches matters even more. In 2009, only 0.7% of worldwide web traffic came from mobile phones. In 2017, mobile accounted for 50.3% of global web traffic. In 10 countries, including the United States and Japan, mobile searches have largely overtaken those made on a computer.

Related guides