Why plain RAG misses
Classic retrieval-augmented generation (RAG) cuts documents into chunks, turns each chunk into a vector and fetches the chunks closest to the question. That works when the answer sits inside one chunk. Real documents are rarely that tidy: a product's price is in one table, its stock in another and the delivery rule on a different page. Each chunk that comes back is only partly relevant, and the model fills the gaps by guessing.
Arabic makes it harder. The same word can be written several ways (with or without diacritics, different forms of hamza and taa marbuta, dialect next to Modern Standard Arabic), so a vector search can miss the one chunk that holds the answer.
How Graph RAG works
Graph RAG adds a step when documents are ingested:
- Extract. A language model reads each document and pulls out entities (products, people, places, rules) and the relationships between them.
- Store. Those entities and relationships go into a graph database such as Neo4j, linked back to the passages they came from.
- Retrieve. When a question arrives, the system finds the entities it mentions and walks the graph to collect the facts connected to them.
- Answer. The model gets that small, connected context, usually together with a few vector-search passages for exact wording.
Why we use it
- Answers that span documents. A question like "is this item in stock and can you deliver it today?" needs facts from several places. The graph brings them back together instead of hoping one chunk contains everything.
- Arabic spelling and dialect. The question only has to reach the right entity, not match a chunk's exact wording, so spelling variation matters much less.
- Smaller, cheaper models. When the context is already connected and precise, the model has less to infer. In our research, a rich knowledge graph brought a 7B model close to the answer quality of a 70B one, which makes answers cheaper and faster and lets the model run on modest hardware.
- Traceable answers. Every fact in the graph points back to its source, so an answer can be checked and cited instead of taken on trust.
- Knowing when not to answer. A question that touches no entity in the graph is very likely off-topic. That is a clearer signal for "politely decline" than a similarity score alone.
Making it work well
- Check the relationships. A single extraction pass misses links or attaches them to the wrong entity; a second pass that reviews and repairs relationships is worth the extra model call.
- Combine graph and vectors. The graph finds connected facts; vector search catches phrasing. Together they beat either alone.
- Treat ingestion as engineering. Batch writes to the graph database and process documents concurrently so large imports stay practical.
When not to use it
If your content is short, uniform and already well structured (an FAQ page, a single policy), plain vector RAG is simpler and good enough. Graph RAG earns its extra cost when facts are spread across many documents and questions need several of them at once.