Memgraph
Back to blog
You Built a Knowledge Graph From 1 Million Documents. Can You Trust It?

You Built a Knowledge Graph From 1 Million Documents. Can You Trust It?

By Sara Tilly, Toni Lastre
11 min readOctober 8, 2026

Building a knowledge graph from a million documents is expensive. This is not especially surprising. Most things become expensive once you add “a million” in front of them. But there’s another problem that gets less attention:

Once you’ve built the graph, how do you know it’s actually correct?

That question came up after we read a recent LinkedIn post by Gurbinder Gill, co-founder of Corvic AI, looking at the cost of building a knowledge graph over a very large document corpus. His example was one million pages. And his post makes a useful point that you probably shouldn’t send every word of every document through an expensive LLM if some of the information you want is already sitting there in the document structure.

Pages have an order. Sections contain paragraphs. Tables belong to pages. Chunks follow other chunks. You don’t need an LLM to discover that page 12 comes after page 11.

Somewhere during the discussion, though, Toni Lastre, our Head of Platform Engineering at Memgraph, noticed something interesting. The post referred to a “structural graph build” in one of its tables. But elsewhere, including the title, it talked about building a knowledge graph. Those two things can overlap. They are not automatically the same thing.

The author of that LinkedIn post clearly understands knowledge graphs. His other writing describes semantic extraction, typed entities and relationships, and deduplication across documents correctly. So this wasn’t really about catching somebody using the wrong word.

It raised a more interesting engineering question:

What exactly are we building when we build a graph over documents and where does the difficult part begin?

So we started a debate internally, of course.

A Graph Over Documents Is Not Necessarily a Knowledge Graph

As Toni explains it, let’s say you have a document describing a company and its products. You split it into chunks and represent the structure like this:

Document → page → chunk 1 → chunk 2 → chunk 3.

That is a graph. It can also be a very useful graph. You’ve preserved information about how the source material is organized: which chunk belongs to which document, what comes before or after it, and perhaps which documents link to one another. Those chunks can still have embeddings and participate in vector search. In fact, the structural graph can enrich vector retrieval by giving it relationships to follow in addition to similarity.

But the graph is still mostly describing the document.

A knowledge graph is trying to describe the things the document is talking about. So, instead, you might have:

Company → builds → product x
Company → builds → product y

And somewhere else, in an entirely different document, you might have:

Customer A → uses → product x

Now we have a different challenge. The two documents might not link to each other at all. The graph connects them because product x means the same thing in both places.

As Toni put it:

“Having chunks of a document in a graph where one chunk points to the next one is not the same thing as representing the knowledge in that document.”

This distinction matters because the second system has to infer things.

And inference is where things start getting interesting.

By which we mean expensive.

Structure Is Cheap-ish, Semantics Are Not

A document already gives you quite a lot for free: page order, headings, sections, tables, paragraph boundaries, sometimes metadata.

You can preserve much of that without asking an LLM what it thinks anything means.

A semantic knowledge graph has a harder job.

It may figure out that, for example, “Acme Widget Platform”, “the Widget product”, “AWP”, and “our flagship platform” all refer to the same thing.

Then it needs to determine what that thing is, which company owns it, which customers use it, what it depends on, and how facts scattered across thousands or millions of chunks relate to each other.

Now you’re doing things like entity extraction, relationship extraction, identity resolution, deduplication, and schema or ontology alignment. Depending on the architecture, an LLM may be doing a significant amount of that work.

This is where the cost jumps.

Toni used a deliberately rough example. Imagine one million documents. Each document has 50 pages. Each page has around 2,000 characters.

That’s Gill’s million pages, fifty times over. That gets you to roughly 100 billion characters of text. Somewhere in the tens of billions of input tokens before you start worrying about output, retries, multiple extraction passes, or the inevitable moment when someone decides the schema should have been slightly different.

The exact dollar figure isn’t terribly useful because it depends on model choice, caching, batching, tokenization, and pipeline design.

The more useful point is:

Once you ask a model to understand what the text means, every page becomes inference work.

Mind you, there’s also a small matter of time.

A million documents means a lot of model calls. You hit rate limits. You wait. You batch things. You retry failed calls. Someone discovers that yesterday’s extraction job quietly stopped at 3:14 AM.

A laptop fan somewhere begins to sound concerned.

Eventually, though, let’s say you make it through. The graph exists, congratulations!

Now comes the awkward part.

Is Any of It Actually Right?!

As one comment on the original discussion put it:

“sure, it’s big and fast, but is the resulting graph any good?”

Exactly.

Because at some point, counting things stops being the same as evaluating them.

A pipeline can tell you that it created 8,000 entities and 12,000 relationships. This is useful information. It proves, among other things, that the pipeline has been extremely busy.

It doesn’t tell you whether those entities and relationships are correct.

Suppose system A extracts twice as many entities as system B. Which one built the better KG?

There is no way to answer that from the counts alone.

Maybe system A found significantly more useful knowledge.

Excellent.

Or maybe it enthusiastically created four versions of the same company and merged two entirely different John Smiths into one unusually accomplished individual.

Less excellent.

Once you’re building a semantic graph, “good” starts meaning something different.

How many of the things we extracted are actually correct? How much important information did we miss? Did we correctly recognize when two mentions refer to the same real-world entity or accidentally merge two different entities because their names looked similar?

And does this edge actually represent what the source says?

Unlike counting nodes, these questions don't have particularly convenient answers.

Computers are very good at counting things. They are also perfectly happy to count things that are wrong. And the larger the corpus gets, the less practical it becomes to inspect the result yourself.

With ten documents, you can look.

With a hundred, you can probably still inspect enough of it to convince yourself that nothing catastrophic is happening.

With a million documents, “we looked at a few and they seemed fine” starts to feel less like quality assurance and more like optimism with a sampling strategy.

Which brings us to the nastier question.

Ten Documents Looked Great. Then We Processed the Other 999,990

Imagine you build an extraction pipeline. You test it on a small, randomly selected sample. Everything looks good. Everything makes sense. Wonderful.

The exact sample size isn’t really the point here. You’re still validating a tiny fraction of what is about to become a very large corpus.

So you run the pipeline over the remaining documents.

At some point, somewhere around document 643,817 the model gets something wrong. Maybe it extracts a relationship the source never states. Maybe it decided that the product x acquired company y, which would certainly be news to both of them.

Now comes the hard part.

Toni's concern wasn't simply that an LLM can hallucinate. We already know that. His concern was what happens afterward.

You can sample outputs. You can have domain experts review random sections of the graph. There are other validation approaches you can add on top.

All useful.

But suppose you know that somewhere inside millions of automatically generated relationships there is a bad one.

How do you find it?

And once you find it, what else did the same extraction run produce? Did other facts depend on it? Should you delete the edge? Reprocess the document?

This is no longer an extraction problem. It’s a recovery problem. And recovery tends to get considerably more interesting when nobody designed for it before the graph becomes enormous.

Quality Becomes an Operational Problem

We’re pretty good at measuring infrastructure. Latency, throughput, token usage, memory, cost.

Computers enjoy these numbers. They are well-behaved numbers. You ask how many milliseconds something took and you get a number of milliseconds back.

“Is this relationship actually true?” is a different sort of question.

That matters because a knowledge graph is not normally built for decorative purposes. Nobody spends months extracting entities and relationships so that everyone can stand admiring the density of the edges.

Something is going to use it.

Applications query it. GraphRAG systems retrieve from it. Agents reason over it. Analytics runs on top of it.

And once other systems begin treating the graph as knowledge, a wrong relationship stops being merely an unfortunate edge and starts getting a career.

It turns up in retrieval results. It joins other facts. It becomes context. Something reasons over it. Before long, one confidently incorrect relationship has acquired dependents.

Fast wrong answers are, after all, still wrong answers. They have simply arrived sooner.

Which means graph quality isn’t an academic concern. It becomes application quality.

Building the Graph May Be the Easy Part

None of this is an argument for giving up on knowledge graphs and defaulting to vector search. Although, understandably, that is sometimes what happens.

But the two approaches solve different problems. If the answer depends on relationships between things, not merely finding text that looks similar to the question, avoiding the graph doesn’t make the relational problem disappear. It just means you’ve decided not to model it.

The more useful direction is to make building graphs less intimidating in the first place.

The hard parts are real.

They’re also engineering problems.

And engineering problems have an irritating tendency to become easier once enough engineers get annoyed by them.

There is still, however, a surprisingly large gap between:

We successfully generated ten million relationships.

and:

We have ten million relationships we trust.

The first statement is relatively easy to prove. There are ten million of them. You can count them. The database will help.

The second statement tends to produce a longer meeting.

Keep the Original Source Close Enough to the Extracted Knowledge

You can review samples. You can improve extraction. You can keep the original source close enough to the extracted knowledge that, when something looks suspicious, you have somewhere sensible to start digging.

The cheapest time to answer the question is before you build the graph. Afterwards, it gets expensive. Then it gets archeological.

Every extracted relationship should carry a pointer back to the chunk it came from, plus a record of the run that produced it: model, prompt version, date. That won't make a relationship true, but it will make it checkable. When someone finds a bad edge at document 643,817, you can pull up the sentence behind it, find every other edge from the same run, and decide what to do next.

All of that helps.

What you do not have is a large green button marked: VALIDATE KNOWLEDGE GRAPH

At least, not one we'd trust.

So perhaps the hardest question in large-scale knowledge graph construction isn't how many pages you can process, how many entities you can extract, or even how much the whole thing costs.

It's what happens six months later when someone points at one relationship among several million and asks:

“Why does the graph think this is true?”

If you can answer that, you're getting somewhere.

If you can't, you may still have a very impressive graph.

You just also have a mystery.

Further Reading

Join us on Discord!
Find other developers performing graph analytics in real time with Memgraph.
© 2026 Memgraph Ltd. All rights reserved.