Baraa Said العربية Message

Retrieval · RA Development

Graph RAG: a 7B model approaching 70B quality

Graph RAG pipelines that turn Arabic and English documents into a Neo4j knowledge graph, so a small 7B model answers close to the quality of a 70B model.

A knowledge graph of connected entities, the core of a Graph RAG pipeline
A knowledge graph: entities as points, relationships as lines
answer quality with a rich knowledge graph
7B ≈ 70B
languages compared: Arabic and English
2
RAG styles compared: naive, graph and agentic
3

The problem

Plain vector search cuts documents into chunks and hopes the right chunk comes back. On messy, real-world Arabic documents it often doesn't: facts that belong together land in different chunks, and the model fills the gaps by guessing.

What I built

As an AI Engineer at RA Development I built Graph RAG pipelines on Neo4j and LangChain:

  • Gemini 2.5 Flash extracts entities and relationships from Arabic text and writes them into a knowledge graph.
  • The same pipeline in English runs on Groq Llama models, so the two languages can be compared side by side.
  • A two-pass relationship fix: a second pass repairs links the first extraction missed or got wrong.
  • Hybrid retrieval: graph traversal brings back connected facts, vector search catches the wording.
  • Benchmarks of Pinecone against ChromaDB as the vector layer, a public API design and a scaling plan for a multi-tenant SaaS.

The result

With a rich knowledge graph behind it, a 7B model came close to the answer quality of a 70B model. Better retrieval buys a smaller, cheaper and faster model, which matters when every answer has a cost.

Where it went next

The same approach runs in production inside Faheem, the Arabic assistant businesses train on their own documents.

Questions about Graph RAG

What is Graph RAG?

Graph RAG is retrieval-augmented generation that stores the facts from your documents as a knowledge graph of entities and relationships. At question time it retrieves connected facts from the graph, often together with vector search, and gives them to the language model as context.

Why does a knowledge graph help a small model?

A small model is weakest when it has to infer missing links. A knowledge graph hands it the links directly: who did what, which product belongs to which category, which rule applies where. With that context a 7B model came close to a 70B model in Baraa's tests.

Does Graph RAG work in Arabic?

Yes. The pipelines use Gemini 2.5 Flash for Arabic entity extraction and were compared directly against English pipelines. The approach now runs in production in Faheem, which answers in Arabic.

More work

Questions?

Hiring for an AI or full-stack role, or need a system built? One message is enough.