Last updated: 2026-10-07
Knowledge Graph Fundamentals
A first encounter with the idea that underlies search engines, recommendation systems, and the atlas on this site
For new readers
You have probably heard the term "knowledge graph" from Google search results, from a chatbot, or from semantic search. This page does not assume you know graph theory, databases, or anything about RDF, ontologies, triples or embeddings. It builds the idea from the ground up, so that later pages on retrieval-augmented generation, agent memory and this site's own concept atlas have something to stand on.RAG often uses a KG to ground facts
Imagine you wanted to represent everything you know about a university. You could create folders: students, courses, staff, buildings. That would organise the things, but it would miss what usually matters more, which is how the things connect. Students take courses. Staff teach courses. Buildings contain rooms. Researchers collaborate with each other. A knowledge graph is built around those connections rather than around the folders.
A working definition: a knowledge graph is a collection of entities and the relationships between them, represented as a network.
The Basic Building Blocks
Three ideas do almost all of the work.
Nodes are the things. A person, an institution, a subject, a course. In the university example: Pat, the University of Reading, Artificial Intelligence, COMP101.
Relationships, usually drawn as edges, are the connections between things. "Pat teaches COMP101" is a relationship. Chains of relationships build up a picture: Pat teaches COMP101, and COMP101 covers Algorithms.edges carry direction and type
Properties are attributes attached to a thing, rather than a connection to another thing. Pat's name and role are properties, not separate nodes: "name = Pat Parslow", "role = Lecturer". Not every fact needs its own node. A property is the right choice when the value is just data about one entity, and a node is the right choice when the value is itself something you want to connect to other things.
From Tables to Networks
A spreadsheet or a database table can hold the same facts. A table with columns for lecturer and module can say that Pat teaches COMP101 and COMP220. Tables are efficient when you already know the question you want to ask, such as "which modules does Pat teach?", because the answer is a lookup.tables need a fixed schema; graphs grow organically
A graph holds the same facts differently.
The graph becomes more useful than the table once the question is not known in advance, or once it involves a chain of relationships rather than a single lookup. "Which other modules cover topics that COMP101 also covers?" or "which staff are two steps away from each other through shared modules?" are awkward to answer from a table and natural to answer by following edges. Tables answer known questions efficiently. Graphs are especially useful for discovering unexpected connections.
Paths Through Knowledge
Once a graph exists, a path through it is a chain of relationships, and meaning often comes from the chain rather than from any single edge.
Read on its own, "COMP101 contains Algorithms" is a small fact. Read as part of a path, it explains why a student needs COMP101 before Machine Learning makes sense. This site's own concept atlas works the same way. The Wayfarer route that appears as you move between pages is a path through a graph of concepts, and the Graph Traversal and Pathway Interrogation page describes how that traversal works in detail.see how the site walks these paths for you
Different Types of Knowledge Graph
Introductory material often implies there is one kind of knowledge graph. There are several, and they are built for different purposes.
Simple concept maps
Concept maps are usually human-built, informal, and small. A classroom concept map showing that "Programming includes Algorithms" is a concept map. They are useful for teaching and for planning a learning path, and they are not usually expected to be complete or precise.
Semantic knowledge graphs
Here, every relationship has a precise, agreed meaning. The standard unit is the triple: subject, predicate, object. "London, capital_of, UK" is a triple. The Resource Description Framework, or RDF, is the main standard for writing triples in a way that different systems can share, and it grew out of the semantic web vision that Berners-Lee, Hendler and Lassila set out in 2001[1]. Before any triple gets written, though, something has already decided what counts as a valid subject, predicate and object in that domain in the first place. That decision is an ontology: a formal specification of which classes and relations a domain recognises, checkable by a machine rather than left as an unstated assumption. The site's Informatics, Frame Analysis, and Ontologies page covers how that specification gets built, and why making it explicit turns an argument about whose "customer" class is right into a concrete, inspectable design choice.
Enterprise knowledge graphs
Organisations such as universities, and companies including Microsoft, Google and Amazon, build knowledge graphs that integrate many internal data sources: employee records, project systems, department structures. "Employee works_on Project, Project belongs_to Department" is the kind of chain an enterprise graph is built to answer. The value here is less about discovering new facts and more about connecting facts that already exist in separate systems.
Search knowledge graphs
Google's Knowledge Graph, launched in 2012, is the best-known example. Google described the shift as moving from "strings" to "things": instead of matching the text of a query, the search understands that "Alan Turing" is an entity with properties and relationships, such as being born in London and having worked on computing[2]. The purpose is to improve search, answer direct questions, and disambiguate entities that share a name.
AI-augmented knowledge graphs
The most recent variant combines a graph with a large language model. The graph supplies structured, checkable facts. The model supplies fluent language. Put together, a system can look up a fact in the graph and use the model to explain it, rather than asking the model to recall the fact from its training and risk it being wrong. This is one of the main reasons graphs have become relevant to AI again: they give an agent something to check its answer against.this is the core of RAG architecture
Knowledge Graphs versus Vector Databases
These two are often confused, and the confusion is understandable, because both are used to find things that are "related". The relationship is established differently.
In a knowledge graph, the relationship is explicit. "Pat teaches Algorithms" is a specific, named connection that someone or something asserted.
In a vector database, items are placed in a space according to similarity, usually computed from an embedding model, and "related" items are simply nearby. Nothing says why Algorithms and Data Structures are close together, only that they are. The site's page on Understanding Large Language Models covers where these embeddings come from.
The short version: graphs represent known relationships. Vectors represent statistical similarity. A system that combines both, using the graph for facts it can check and vectors for finding plausibly relevant material, tends to be more capable than one that relies on either alone. The site's page on Retrieval-Augmented Generation describes a system built mainly on the vector side of that combination.vectors capture 'vibes', graphs capture facts
A Different Kind of Graph: Bayesian Networks
It is worth setting knowledge graphs apart from a graph structure that looks similar but means something different: the Bayesian network. Judea Pearl's foundational account treats a Bayesian network as a directed graph in which the nodes are random variables and the edges represent probabilistic dependence, used to reason about uncertain situations[3].
The difference matters. An edge in a knowledge graph asserts a specific, usually categorical fact: "London capital_of UK" is true or it is not. An edge in a Bayesian network says that one variable's probability depends on another's, and the network as a whole lets you compute how evidence about one variable should shift your belief about the others. "Rain" influencing the probability of "Wet Grass" is a Bayesian network edge, not a knowledge graph edge, because the point is the conditional probability, not a named relationship between two fixed facts.edges are conditional probabilities
The two structures are often useful together. A knowledge graph can hold what is known. A Bayesian network built over some of the same entities can hold how confident a system should be, and how that confidence should change as new evidence arrives. Confusing the two is a common mistake, because both are drawn as nodes and edges, and only one of them is meant to be read as a set of probabilities.
A Third Point of Comparison: Trust as a Learned Relationship
Knowledge-graph edges are categorical, and Bayesian-network edges are conditional probabilities. Worth placing alongside both is a third kind of edge that turns up in exactly the same nodes-and-edges shape while meaning something different again: this site's page on A Distributed Embedding Calculus of Trust represents trust between two parties not as a fixed fact or a probability table, but as an experiential embedding โ a vector, learned from the specific history of interactions between that pair, that drifts as new evidence arrives.
The resemblance to the paths described in "Paths Through Knowledge" above is closer than it first looks. Composing two hops of trust, \(T(A, C) = T(A, B) \times T(B, C)\), is the trust calculus's own version of following a chain of knowledge-graph edges into a path. Where a graph path simply holds or it doesn't, a trust path degrades: each additional hop discounts the estimate, and an untrustworthy intermediary suppresses what they vouch for rather than passing it through unchanged. Closer still to the vector-database side of the earlier comparison than to either graph, trust embeddings let two entities with no direct history still be assigned an estimated score from how similar their embeddings are โ the same way a vector database returns items that are merely near each other with no named edge explaining why.
Knowledge Graphs and Agents
An AI agent answering a question can draw on a graph in several ways at once: as memory of what has happened before, as context for the current task, as a world model of how things relate, and as an organisational knowledge base that does not depend on what the underlying language model happened to learn during training.
This is likely where the idea becomes practically important rather than merely tidy. A language model can produce a fluent answer to almost anything, whether or not the answer is correct. A graph lookup can only return what is actually recorded. An agent that checks its draft answer against a graph has a way of catching the cases where fluency and correctness have come apart, which is the central concern of the page on How LLMs Reason, Self-Correct and Get Checked.ground truth vs. probable next token
Limitations
None of this makes a knowledge graph a solved problem. Building one is expensive, and keeping it current is harder still, because the world keeps changing after the graph is built. Real graphs are usually incomplete, sometimes inconsistent where different sources disagree, and always shaped by the modelling decisions of whoever built them: what counts as an entity, which relationships were worth recording, and which were left out.
A knowledge graph is always a model of reality, not reality itself. Hogan and colleagues' survey of the field treats this as a starting assumption rather than a flaw to be engineered away: a knowledge graph is defined by the modelling choices behind it as much as by the facts inside it[4]. Treating a graph as a finished, authoritative picture of the world is the mistake to avoid, whether the graph was built by a university, a search engine, or a single enthusiastic researcher.
Two knowledge graphs covering the same domain can disagree, invisibly, about what counts as an entity in the first place โ whether a cancelled account is still a customer, or a visiting lecturer counts as staff. That disagreement is not a bug waiting to be coded around. It is a genuine difference in what each graph's underlying ontology commits to, and the site's page on meaning, ontology, and the limits of fixed definitions sets out why a classification like this is a disciplined proposal for a purpose, not the discovery of a category's hidden essence.
A Closing Analogy
Documents store knowledge in containers. Databases store knowledge in tables. Knowledge graphs store knowledge in relationships. None of the three is the right choice for everything, and the more the relationships between things matter to the question you are asking, the more a graph has to offer over the alternatives.
Related Topics
- Retrieval-Augmented Generation โ a system built mainly on the vector-similarity side of the comparison in this page.
- Understanding Large Language Models โ where embeddings and vector similarity come from.
- How LLMs Reason, Self-Correct and Get Checked โ why an agent benefits from a graph it can check its own answers against.
- Graph Traversal and Pathway Interrogation โ how this site's own atlas turns a knowledge graph into a navigable map.
- K-Blades vs. Self-Organizing Maps โ how the live concept atlas behind this site is actually built.
- The Caring Learning Agent โ a learner model that plays a similar role to a knowledge graph, as a structured account an agent can reason over.
- A Distributed Embedding Calculus of Trust โ a third kind of edge, learned and continuous rather than categorical or conditional-probability.
- Informatics, Frame Analysis, and Ontologies โ the formal specification that decides what counts as a valid node, relation and triple before any knowledge graph gets built.
- Meaning, Ontology, and the Limits of Fixed Definitions โ why two knowledge graphs covering the same domain can disagree about what counts as an entity in the first place.
References
- Berners-Lee, T., Hendler, J., & Lassila, O. (2001). The Semantic Web. Scientific American, 284(5), 34โ43.
- Singhal, A. (2012, May 16). Introducing the Knowledge Graph: Things, not strings. The Keyword, Google. https://blog.google/products-and-platforms/products/search/introducing-knowledge-graph-things-not/
- Pearl, J. (1988). Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann.
- Hogan, A., Blomqvist, E., Cochez, M., d'Amato, C., de Melo, G., Gutiรฉrrez, C., Kirrane, S., Labra Gayo, J. E., Navigli, R., Neumaier, S., et al. (2021). Knowledge graphs. ACM Computing Surveys, 54(4), Article 71. https://doi.org/10.1145/3447772