Flux
Flux answers questions about biological pathways, such as “how does the PI3K-AKT pathway regulate cell survival?”, from Reactome, KEGG, and OpenAlex rather than from a model’s memory. It stores pathway text in Postgres with pgvector and pathway structure in a Neo4j graph of about 15,000 nodes and 40,000 relationships. A LangGraph agent retrieves from both, has a second model judge whether the evidence is good enough, searches again if it is not, and returns an answer as validated JSON with a citation behind every claim.
The problem
Biology questions are a bad fit for a plain chat model in two ways. The facts are specific — which genes a pathway contains, which reactions it includes — and models blur them. And the useful answers are often relational: what connects two pathways, what sits downstream of a protein. Text search over descriptions can find a pathway that mentions AKT; it cannot walk from that pathway to everything it contains. I wanted a system that could do both, and that would refuse to present an answer it could not back with a source.
One schema from three sources
Each source has its own client, parser, and normalizer under data/ingest/. Reactome returns nested pathway records with summaries, compartments, and contained events. KEGG returns flat entries with gene, compound, and reaction lists. OpenAlex returns paper abstracts tagged with concepts. The normalizers turn all three into one document shape: a type, a source, an external_id, a title, content, and source-specific metadata. Every later stage handles one shape instead of three.
In Postgres, documents has a UNIQUE(source, external_id) constraint and chunks has UNIQUE(document_id, chunk_index), so re-running ingestion updates rows instead of duplicating them. Chunk embeddings are 768-dimensional vectors behind an HNSW index with cosine distance. Reactome descriptions are split with a semantic chunker that breaks where consecutive sentences’ embeddings diverge, at the 75th percentile, rather than at a fixed character count.
The graph, and what it does not merge
The Neo4j loaders write every node with MERGE on a stable identifier, never CREATE, behind uniqueness constraints on Pathway.id, Reaction.id, Gene.id, and Compound.id. That makes the load idempotent: running it twice produces the same graph. Shared concepts like species and cellular compartments are merged by name, so every human pathway from both sources points at one Species node instead of a hundred copies.
I want to be precise about the limit. Reactome and KEGG pathways keep their own ids, R-HSA for one and hsa for the other, and the loader does not resolve a Reactome pathway and a KEGG pathway to the same node, even when they describe the same biology. Doing that properly means mapping through shared gene or protein identifiers, such as the UniProt mapping file already in data/pathways/, and deciding how much overlap counts as “the same pathway”. That is the most valuable thing missing from the graph.
Two retrievers, then a judge
The agent is a LangGraph state machine. A router first decides what the question needs. A question over 30 words, with more than one question mark, or joining clauses with “and”, “or”, or “as well as”, is decomposed into subquestions. A short question, under eight words, or one with unexplained acronyms like PI3K or p53, is expanded; a long or noisy one, over 25 words or containing error text, is compressed. Anything else goes straight to embedding.
Both retrievers then run on every question; the graph does not choose between them. The Neo4j retriever finds the five pathways most similar to the query and expands each into its subgraph to depth two, rendering the related reactions, genes, and compounds as text. The Postgres retriever runs pgvector search and BM25, 20 candidates each, unions them by a deterministic source_externalid_chunkindex id so a chunk found by both is scored once, scores every candidate with the ms-marco-MiniLM-L-6-v2 cross-encoder, and keeps the top five.
Then a judge model reads the question and the merged evidence and returns a verdict of good, partial, or bad, a confidence, and a list of what is missing. That list is the useful part. On a weak verdict the agent turns each missing aspect into a targeted subquery, searches again, and judges again, up to three attempts.
def should_reretrieve(state: dict) -> str:
validated = GraphState(**state)
judge_result = validated.get_judge_result()
retrieval_metadata = state.get("retrieval_metadata", {})
reretrieval_attempts = retrieval_metadata.get("reretrieval_attempts", 0)
max_attempts = 3
if reretrieval_attempts >= max_attempts:
return "generate"
if judge_result is None:
return "generate"
if judge_result.verdict == "good" and judge_result.confidence >= 0.7:
return "generate"
if judge_result.verdict == "bad" or (judge_result.verdict == "partial" and judge_result.confidence < 0.7):
return "reretrieve"
return "generate"Answers as data, not prose
The final answer is not free text. The generator must return JSON that validates against a Pydantic FinalAnswer model: the answer, a list of citations that each point at a chunk id, a list of claims that each name the citations supporting them, a confidence between 0 and 1, the judge’s retrieval-quality verdict, and a one-to-three-sentence reasoning summary. The dataset and pipeline versions are stamped on every answer, so a result can be traced to the data that produced it.
Validation runs in three passes. First the JSON has to parse. Then it has to match the schema. Then every citation has to refer to a chunk the model was actually given, and every claim has to cite something. A citation to a chunk that does not exist is the structured form of a hallucinated source, and it is caught here instead of shown to the user. If parsing or the schema fails, the user gets an explicit fallback answer with zero confidence and the error in its warnings, never a half-parsed one.
class FinalAnswer(BaseModel):
answer: str
citations: List[Citation] = Field(default_factory=list)
claims: List[Claim] = Field(default_factory=list)
confidence: float = Field(ge=0.0, le=1.0)
retrieval_quality: Literal["good", "partial", "bad"]
used_sources: UsedSources = Field(default_factory=UsedSources)
warnings: List[str] = Field(default_factory=list)
reasoning_summary: strCaching what is expensive
Embedding a query, running the cross-encoder, and calling the judge are the slow steps, and users repeat themselves. Flux keeps a Redis cache in front of each step, with a time-to-live matched to how fast the result goes stale: seven days for query embeddings, one hour for retrieval and reranking results, thirty days for answers. Keys are a SHA-256 of the normalized query, so case and whitespace differences hit the same entry. Retrieval keys include an index version, so rebuilding the index invalidates old results without flushing anything. Requests that carry user or session filters are never cached.
The answer cache is semantic: a new question is compared to cached ones by cosine similarity of their embeddings, and a match at 0.95 or above returns the cached answer, so a light rephrasing still hits. If Redis is down, the agent runs uncached instead of failing.
The code is at github.com/andrewzapps/flux.