The problem
Plain vector search cuts documents into chunks and hopes the right chunk comes back. On messy, real-world Arabic documents it often doesn't: facts that belong together land in different chunks, and the model fills the gaps by guessing.
What I built
As an AI Engineer at RA Development I built Graph RAG pipelines on Neo4j and LangChain:
- Gemini 2.5 Flash extracts entities and relationships from Arabic text and writes them into a knowledge graph.
- The same pipeline in English runs on Groq Llama models, so the two languages can be compared side by side.
- A two-pass relationship fix: a second pass repairs links the first extraction missed or got wrong.
- Hybrid retrieval: graph traversal brings back connected facts, vector search catches the wording.
- Benchmarks of Pinecone against ChromaDB as the vector layer, a public API design and a scaling plan for a multi-tenant SaaS.
The result
With a rich knowledge graph behind it, a 7B model came close to the answer quality of a 70B model. Better retrieval buys a smaller, cheaper and faster model, which matters when every answer has a cost.
Where it went next
The same approach runs in production inside Faheem, the Arabic assistant businesses train on their own documents.


