About this project
ReasonTrace is an agentic RAG system that decides when to retrieve, when to retrieve again, when to call a tool, when to verify — and when to refuse to answer. Nine LangGraph nodes and six conditional-edge functions keep the control flow in plain, readable Python instead of hiding it inside a prompt: the agent first enumerates the atomic facts a question needs, then after each search computes which facts are still missing and writes a follow-up query targeting them — so a genuine multi-hop question is answered from both documents instead of half the evidence. Sufficiency is computed rather than asked for (the LLM only answers a narrow, checkable “does this passage state fact X?”), arithmetic is routed to a deterministic calculator whose operands must already appear in retrieved evidence, and a verifier audits the draft against five checks — citations resolving to real chunks, every number traceable — before the answer is published; an adversarial lying provider is caught and the run ends in abstention. Contradictions are settled by document front-matter (effective_date, status, supersedes) and disclosed in the answer, iteration/retrieval/tool budgets with duplicate-query detection prevent loops, and every run exposes a structured trace of observable actions and routing decisions — never model chain-of-thought. Evaluated on the path, not the string: 6/6 eval cases, 100% route accuracy and citation grounding, 0 wasted tool calls, 229 tests. Hybrid retrieval (dense vectors + BM25 → reciprocal rank fusion → reranking), section-aware chunking with citation metadata, and providers, embeddings and vector stores that swap by environment variable — including a deterministic offline engine that runs the whole graph with no API key.
Technologies: Python, LangGraph, FastAPI, Streamlit, ChromaDB, Pydantic v2, sentence-transformers, Docker