Computers help us find things. 
Computers help us find things. 
Computers help us find what we need. This is called information retrieval. It is the science of searching for things. You might search for a book or a song. You can even search for images or videos. 
The way it works starts with a query. A query is a formal statement of what you want. You might type words into a search engine. The system then looks through a database. This database is a large collection of information. The system compares your query to the data. It looks for things that match your words.
Most systems give each result a score. This score shows how well it matches. The system then ranks the results. It puts the best matches at the top. This helps you find the right thing fast. Some systems use metadata to help. Metadata is data that describes other data.
In the past, people used film to store documents. Now, we use the web. Search engines like Google use special rules. These rules help them find the most important pages. New tools now use deep learning. This helps computers understand how words work together.
Information retrieval is the science of finding things we need. It helps us find resources like books, journals, or sounds. This science is very important because it helps reduce information overload. That is a hard job because there is so much data to look through. We use these systems every day to find specific things. 
The way it works starts with a user entering a query. A query is a formal statement of what you need. You might type a string of words into a search engine. The system then looks through a database to find matches. These matches are called objects, which can be text, images, or videos. Most systems do not store the whole document directly. Instead, they use metadata, which is data that describes other data. The system gives each object a numeric score based on its match. It then ranks the objects from best to worst. 
People have been thinking about this for a long time. In 1945, Vannevar Bush wrote an article called "As We May Think." He suggested using computers to find relevant information. He was inspired by patents from the 1920s and 1930s. These patents were for a machine that searched documents on film. In 1948, a person named Holmstrom described a computer searching for info. By the 1960s, Gerard Salton formed a research group at Cornell. 
Many important milestones helped search engines grow. In 1992, the US government helped start the Text Retrieval Conference. This helped researchers learn how to search huge collections of text. In 1994, Yahoo! became a popular search tool. Then, Google was founded in 1998. Google used the PageRank algorithm to rank pages by importance. In 2009, Microsoft launched Bing to offer new features. Later, in 2018, Google used a model called BERT. BERT helps computers understand the context of words in a sentence. 
You can see these tools in many parts of your life. They help with music retrieval or finding news stories. They even help scientists find genomic information or chemical structures. Some systems use machine learning to learn from how people click. Researchers now study three main types of models. These are sparse models, dense models, and hybrid models. Sparse models look for exact word matches. Dense models look for deeper meaning. Hybrid models try to use the best of both ways. 
Information retrieval, or IR, is a specialized field within computing and information science. Its primary goal is to identify and retrieve resources that match a specific information need. This process is essential in our modern world to prevent information overload. Information overload occurs when the amount of available data is too large for a person to manage alone. IR systems act as tools to help users navigate massive collections of digital data. These systems can search through text, images, sounds, and even complex metadata. Metadata is simply data that describes other data, such as the author of a book or the date a photo was taken.
The mechanism of an IR system begins with a user entering a query. A query is a formal statement of what the user is looking for. In a web search engine, this is usually a string of words. Unlike a standard database query, an IR query does not usually identify a single, perfect object. Instead, the system finds several objects that might be relevant to the user. These objects can be text documents, videos, or audio files. To manage this, the system often uses document surrogates. These are representations of the documents rather than the full files themselves. The system calculates a numeric score for each object based on its match to the query. It then ranks these objects from most relevant to least relevant. This ranking process is a defining characteristic of information retrieval.
Researchers categorize IR models using different mathematical and structural dimensions. One way to group them is by their mathematical basis. Set-theoretic models treat documents as sets of words or phrases. These use operations to find similarities between those sets. Algebraic models represent documents and queries as mathematical structures like vectors or matrices. In these models, similarity is shown as a single scalar value. Probabilistic models treat retrieval as a matter of chance. They calculate the probability that a document is relevant to a query. Finally, feature-based models use various functions to create a single relevance score. These models can be further classified by how they handle term interdependencies. Some treat words as independent, while others account for how words relate to each other.

The history of information retrieval is tied to the evolution of computing. In 1945, Vannevar Bush popularized the idea of computer-aided searching in his article "As We May Think." He was likely inspired by Emanuel Goldberg, who filed patents for a statistical machine in the 1920s and 1930s. These machines searched for documents stored on film. By the 1960s, Gerard Salton formed a major research group at Cornell University. In the 1970s, large-scale systems like the Lockheed Dialog system began to appear. A major turning point occurred in 1992 when the US Department of Defense and NIST co-sponsored the Text Retrieval Conference (TREC). This event provided the infrastructure needed to test how retrieval methods work on massive collections of text.
The rise of the World Wide Web fundamentally changed the field in the late 1990s. Early search engines like Yahoo! and AltaVista used simple keyword-based retrieval. However, they were limited in how they ranked results. In 1998, the founding of Google introduced a major breakthrough with the PageRank algorithm. PageRank used the structure of hyperlinks to judge the importance of a webpage. In 2009, Microsoft launched Bing, which later integrated semantic web technologies through its Satori knowledge base. A massive leap happened in 2018 when Google deployed BERT. BERT, which stands for Bidirectional Encoder Representations from Transformers, uses deep neural language models. This allowed computers to understand the context and meaning of words more like a human would.
Today, IR technology is used in many specialized areas. In general applications, it powers web search engines, digital libraries, and recommender systems. It is also used for media searches involving music, video, and images. In scientific fields, IR is used for genomic information retrieval and searching for chemical structures. It is also vital in legal and software engineering fields. Modern research often focuses on three specific types of deep learning models. Sparse models rely on inverted indexes to find exact term matches. Dense models use continuous vector embeddings to find semantic similarity. Hybrid models combine these two approaches to balance precision and depth.

As these systems become more complex, the focus of the scientific community is shifting. While efficiency and relevance remain important, new concerns are emerging. Researchers are now studying the impact of bias and fairness in retrieval algorithms. There is also a growing need for explainability in how these systems make decisions. This means making the logic behind a search result more transparent to the user. The goal is to build systems that foster user trust and accountability. As we move forward, the integration of machine learning and natural language processing will continue to transform how we interact with the world's information.
🖼️ Images & Media (1)
More to explore
✨ What else?
Related topics you might enjoy
🪜 Step back
Simpler topics to build understanding
What is Nepedia?
A free, ad-free encyclopedia for children. Every article is written at five reading levels, so the same page works for a five-year-old and a fifteen-year-old — use the level switcher above to see this one change. No account needed to read.