← Projects

Flux

Personal project2026


Flux answers questions about biological pathways, such as “how does the PI3K-AKT pathway regulate cell survival?”, from Reactome, KEGG, and OpenAlex rather than from a model’s memory. It stores pathway text in Postgres with pgvector and pathway structure in a Neo4j graph of about 15,000 nodes and 40,000 relationships. A LangGraph agent retrieves from both, has a second model judge whether the evidence is good enough, searches again if it is not, and returns an answer as validated JSON with a citation behind every claim.

The problem

Biology questions are a bad fit for a plain chat model in two ways. The facts are specific — which genes a pathway contains, which reactions it includes — and models blur them. And the useful answers are often relational: what connects two pathways, what sits downstream of a protein. Text search over descriptions can find a pathway that mentions AKT; it cannot walk from that pathway to everything it contains. I wanted a system that could do both, and that would refuse to present an answer it could not back with a source.

One schema from three sources

Each source has its own client, parser, and normalizer under data/ingest/. Reactome returns nested pathway records with summaries, compartments, and contained events. KEGG returns flat entries with gene, compound, and reaction lists. OpenAlex returns paper abstracts tagged with concepts. The normalizers turn all three into one document shape: a type, a source, an external_id, a title, content, and source-specific metadata. Every later stage handles one shape instead of three.

In Postgres, documents has a UNIQUE(source, external_id) constraint and chunks has UNIQUE(document_id, chunk_index), so re-running ingestion updates rows instead of duplicating them. Chunk embeddings are 768-dimensional vectors behind an HNSW index with cosine distance. Reactome descriptions are split with a semantic chunker that breaks where consecutive sentences’ embeddings diverge, at the 75th percentile, rather than at a fixed character count.

SOURCES Reactome · pathways KEGG · genes, compounds OpenAlex · abstracts NORMALIZE one document schema · keyed by (source, external_id) STORES Postgres · pgvector chunks, HNSW Neo4j · MERGE on stable ids
The orange lane is the whole trick. Three source formats converge on one schema before anything is stored, so both stores and every retriever only ever see one shape of data.

The graph, and what it does not merge

The Neo4j loaders write every node with MERGE on a stable identifier, never CREATE, behind uniqueness constraints on Pathway.id, Reaction.id, Gene.id, and Compound.id. That makes the load idempotent: running it twice produces the same graph. Shared concepts like species and cellular compartments are merged by name, so every human pathway from both sources points at one Species node instead of a hundred copies.

I want to be precise about the limit. Reactome and KEGG pathways keep their own ids, R-HSA for one and hsa for the other, and the loader does not resolve a Reactome pathway and a KEGG pathway to the same node, even when they describe the same biology. Doing that properly means mapping through shared gene or protein identifiers, such as the UniProt mapping file already in data/pathways/, and deciding how much overlap counts as “the same pathway”. That is the most valuable thing missing from the graph.

REACTOME Pathway · R-HSA id Reaction Compartment KEGG Pathway · hsa id Gene · Compound Reaction SHARED Species · one Homo sapiens node, MERGE by name, unique constraint
Read across, then down. Each source keeps its own pathway ids; only the bottom lane is truly shared, which is why cross-source entity resolution is the first thing on the list of what is left.

Two retrievers, then a judge

The agent is a LangGraph state machine. A router first decides what the question needs. A question over 30 words, with more than one question mark, or joining clauses with “and”, “or”, or “as well as”, is decomposed into subquestions. A short question, under eight words, or one with unexplained acronyms like PI3K or p53, is expanded; a long or noisy one, over 25 words or containing error text, is compressed. Anything else goes straight to embedding.

Both retrievers then run on every question; the graph does not choose between them. The Neo4j retriever finds the five pathways most similar to the query and expands each into its subgraph to depth two, rendering the related reactions, genes, and compounds as text. The Postgres retriever runs pgvector search and BM25, 20 candidates each, unions them by a deterministic source_externalid_chunkindex id so a chunk found by both is scored once, scores every candidate with the ms-marco-MiniLM-L-6-v2 cross-encoder, and keeps the top five.

Then a judge model reads the question and the merged evidence and returns a verdict of good, partial, or bad, a confidence, and a list of what is missing. That list is the useful part. On a weak verdict the agent turns each missing aspect into a targeted subquery, searches again, and judges again, up to three attempts.

python
def should_reretrieve(state: dict) -> str:
    validated = GraphState(**state)
    judge_result = validated.get_judge_result()

    retrieval_metadata = state.get("retrieval_metadata", {})
    reretrieval_attempts = retrieval_metadata.get("reretrieval_attempts", 0)
    max_attempts = 3

    if reretrieval_attempts >= max_attempts:
        return "generate"
    if judge_result is None:
        return "generate"
    if judge_result.verdict == "good" and judge_result.confidence >= 0.7:
        return "generate"
    if judge_result.verdict == "bad" or (judge_result.verdict == "partial" and judge_result.confidence < 0.7):
        return "reretrieve"
    return "generate"
QUERY question route · enhance, decompose, skip embed RETRIEVE Neo4j · subgraph pgvector + BM25 · rerank merge JUDGE judge · verdict, confidence ≥ 0.7 generate · validated JSON RETRY subqueries · from missing_aspects → search again, at most 3 times
The dashed lane is the corrective part of corrective RAG. The judge does not just say no; it says what is missing, and that becomes the next search.

Answers as data, not prose

The final answer is not free text. The generator must return JSON that validates against a Pydantic FinalAnswer model: the answer, a list of citations that each point at a chunk id, a list of claims that each name the citations supporting them, a confidence between 0 and 1, the judge’s retrieval-quality verdict, and a one-to-three-sentence reasoning summary. The dataset and pipeline versions are stamped on every answer, so a result can be traced to the data that produced it.

Validation runs in three passes. First the JSON has to parse. Then it has to match the schema. Then every citation has to refer to a chunk the model was actually given, and every claim has to cite something. A citation to a chunk that does not exist is the structured form of a hallucinated source, and it is caught here instead of shown to the user. If parsing or the schema fails, the user gets an explicit fallback answer with zero confidence and the error in its warnings, never a half-parsed one.

python
class FinalAnswer(BaseModel):
    answer: str
    citations: List[Citation] = Field(default_factory=list)
    claims: List[Claim] = Field(default_factory=list)
    confidence: float = Field(ge=0.0, le=1.0)
    retrieval_quality: Literal["good", "partial", "bad"]
    used_sources: UsedSources = Field(default_factory=UsedSources)
    warnings: List[str] = Field(default_factory=list)
    reasoning_summary: str

Caching what is expensive

Embedding a query, running the cross-encoder, and calling the judge are the slow steps, and users repeat themselves. Flux keeps a Redis cache in front of each step, with a time-to-live matched to how fast the result goes stale: seven days for query embeddings, one hour for retrieval and reranking results, thirty days for answers. Keys are a SHA-256 of the normalized query, so case and whitespace differences hit the same entry. Retrieval keys include an index version, so rebuilding the index invalidates old results without flushing anything. Requests that carry user or session filters are never cached.

The answer cache is semantic: a new question is compared to cached ones by cosine similarity of their embeddings, and a match at 0.95 or above returns the cached answer, so a light rephrasing still hits. If Redis is down, the agent runs uncached instead of failing.

The code is at github.com/andrewzapps/flux.