← Projects

NeuroMANCER-GPT

RAG Systems Developer Intern · Johns Hopkins2026


NeuroMANCER is PNNL’s open-source scientific machine learning library for constrained optimization, system identification, and model-based control, with more than 37,000 users. NeuroMANCER-GPT is a retrieval-augmented assistant I built for it: you ask a question in plain English, it searches the library’s own source, docs, and examples, and the model answers only from what it found, with the file paths it drew from listed under the answer. It runs fully local through Ollama or against an OpenAI-compatible API.

I submitted it upstream as pull request #290: +3,576 lines across 30 files, 70 pytest tests, and no changes to the library’s own src/. It is open and awaiting maintainer review. The working branch is assistant/neuromancer-gpt on my fork.

This page is the design write-up. Most of it is about retrieval, because that is where the hard decisions were: how the repository gets chunked, why every documented function is indexed two ways, how three searches are fused and reranked, and what the model is told to do with the result. After that it covers how I kept the pull request small enough to review, how it is tested, and how it was made fast enough to run on a laptop.

Showcase video · 0:35

The problem

Ask a general-purpose chat model how to set up a Problem with a soft constraint in NeuroMANCER and it will usually answer confidently and wrongly. It invents constructors and keyword arguments that look plausible and do not exist, because a specialized research library is barely represented in training data. The failure is not that the model says “I don’t know”; it is that it never does.

The fix is to stop asking the model to remember the library and make it read it instead. Retrieve the relevant code and docs from the repository at question time, put them in front of the model, and forbid it from naming anything that is not in front of it. Then every claim in an answer traces back to a file a user can open, and when the repository has no answer, the assistant says so.

Where the original script stopped

The repository already shipped a preprocessing script when I started. It walked the checkout, converted notebooks to commented Python with pandoc, and concatenated everything into three files: docs.txt, src.txt, and examples.txt, with a line of dashes between sources. The README said the intent plainly: upload these to a RAG-as-a-service product and let it do the rest.

That design pushes the hard decisions onto whichever service receives the files, and those services make the same decision every time: split on a fixed character window. A window does not know where a function ends, so a docstring lands in one chunk and its signature in the next. Nothing carries a file path or a symbol name, so an answer cannot cite anything more specific than the whole 5 MB file. The rest of this page is what it took to make the pipeline local, structured, and specific.

FEB 2025 repo ingest.py · concatenate 3 text blobs · upload somewhere else NOW repo ingest.py · chunk by symbol JSONL · ids 2 indexes
The top lane ends at a file someone else has to index. The bottom lane ends at two indexes the app owns, which is what makes citations and offline use possible.

What was already there, and what I built

I did not start from nothing, and the pull request says so. The repository’s assistant/ directory already had the preprocessing script above, a carefully written system prompt describing NeuroMANCER’s concepts, and a stated goal: lower the barrier to using LLM assistants with the library. What it did not have was the retrieval pipeline it assumed someone would run elsewhere.

My pull request supplies that pipeline and keeps the original intent. The system prompt is preserved unchanged as prompt.txt and is still what the model reads first. ingest.py and ipynb_filter.py are replaced, because three concatenated text files are not a usable unit of retrieval. Everything else is new: the structural chunker, the index builder, hybrid retrieval and reranking, query rewriting, the grounding prompt, the Streamlit app, the setup scripts, and the tests. I designed and wrote all of it; the review I have asked for is from the library’s maintainers.

Chunking along the code's own seams

The new ingest.py keeps the original walker and ignore lists, but instead of appending file text it asks each corpus how it is structured. Python source goes through the standard library’s ast and becomes one record per top-level class or function and one per method, with the class name prepended so Trainer.fit is a real symbol rather than a bare fit. Private names are indexed only if they have a docstring or take arguments; dunder methods other than __init__ are skipped.

Documentation splits at headings. The script reads both reStructuredText underlines and Markdown hashes, and it drops files that are nothing but Sphinx automodule directives, since a page with eight lines of prose and a dozen directives embeds as noise. Notebooks are the interesting case. A markdown cell is glued to the code cell that follows it, and a markdown-only stretch shorter than 200 characters is carried forward into the next chunk, so a heading never becomes its own orphan record. Large example scripts over 200 lines split per top-level definition; anything still over 6,000 characters is cut on line boundaries, because the embedding model would only see the front of it anyway.

Every record gets an id of the form path:symbol:line, and the file path, symbol, source type, and line range travel with it as metadata. That metadata is what the UI later prints as a citation. On a current checkout this yields about 1,900 chunks.

SRC module.py ast · one chunk per class and def api · Trainer.fit DOCS .rst / .md headings · autodoc shells dropped doc · one section EXAMPLES .ipynb / .py cells · prose stays with its code example · one topic
Each corpus has its own seam. The middle column is the rule; the right column is what one resulting chunk looks like.

Two views of every function

The first version of the source chunker embedded the full body of every function. It retrieved badly, and the reason was subtle: a signature-plus-docstring chunk and a full-implementation chunk for the same function are near-duplicates, so they crowd each other in the top of every result list and push out the tutorial the user actually needed. The fix was to stop pretending they are two documents.

For a documented symbol, only the signature and docstring are embedded, because that is the text a question resembles. The implementation is attached to the same record as an impl field, stored in the vector database as metadata, and never searched. When that chunk is retrieved, the prompt builder appends the implementation inside a code fence, so the model still sees the body. Undocumented symbols have no docstring to embed, so they are indexed on their code instead and tagged impl rather than api.

python
if doc and doc.strip():
    # embed only the signature + docstring;
    # carry the implementation along for display instead of indexing
    # a near-duplicate chunk that competes in search
    api_content = extract_signature_and_doc(node, source_lines)
    record = make_record(rel_path, "api", symbol_name, start, end, api_content)
    if len(impl_content) > MAX_IMPL_CHARS:
        impl_content = impl_content[:MAX_IMPL_CHARS] + "\n# ... truncated ..."
    record["impl"] = impl_content
    records.append(record)
else:
    # no docstring: a bare signature won't match queries, index the code
    records.append(
        make_record(rel_path, "impl", symbol_name, start, end, impl_content)
    )

Implementations are capped at 12,000 characters in the record and 2,000 in the prompt. The cap in the prompt matters more than it looks: eight excerpts with full bodies would not fit the 8,192-token context the local model runs with.

ONE DEF def fit(self, ...) · with docstring api chunk · signature + docstring, indexed impl · full body as payload, never searched
One record, two roles. Only the top branch competes in search; the bottom branch rides along for the model.

One corpus, two indexes

load_vector_store.py reads the three JSONL files into three Chroma collections, one per corpus, so a query can weight them separately later. Embeddings come from nomic-embed-text through Ollama. Nomic models expect a task prefix, so every document is embedded as search_document: followed by a provenance header (File: ... | Symbol: ... | Type: ...) and then the text, while queries get search_query:. The header is deliberately part of what is embedded and not part of what is stored: it lets a question that mentions a file or symbol name pull the right chunk, without polluting the excerpt the model reads.

The same header-plus-text is tokenized into lowercase word tokens and fed to BM25Okapi. The pickle that comes out stores only chunk ids, collection names, and the BM25 statistics; document text is fetched back from Chroma at query time. That keeps the file small and cold starts fast. It also stores the chunk count it was built against, and retrieval refuses to use it if the live collections disagree, which is the failure you get when someone re-ingests and forgets to rebuild.

DENSE search_document: File | Symbol | Type + text nomic-embed-text Chroma · 3 SPARSE same header + text · \w+ tokens BM25Okapi pickle · ids only
Both indexes see the provenance header. Only the dense store keeps the text, and only the plain text, so the header never leaks into an excerpt.

Why one search is not enough

retrieve.py runs three searches for every question and they exist because each one fails differently. Dense search handles paraphrase: “how do I add a constraint” finds the Constraint docs even though the word “add” appears nowhere in them. But vector similarity blurs rare identifiers; DictDataset and DictDatasetLoader embed almost identically. Sparse search is the opposite: BM25 will not connect “penalty” to “soft constraint”, but it will find the one tutorial that says DictDataset six times.

Symbol search covers the case the other two both get wrong, which is an API name typed verbatim. At startup the app reads every symbol_name from the source collection into a set. A query is tokenized on identifier characters, including dotted paths, and any token that is in the set becomes an exact metadata lookup. Someone who types LitTrainer.fit gets that definition in the pool no matter what the embedding thinks the question is about.

Dense and sparse each return 24 candidates. Symbol hits are added afterwards, deduplicated by id, so they never displace a fused result before reranking gets a look at them.

QUERY question · plus chat history REWRITE contextualize · resolve "it", then clean SEARCH dense · vectors, 24 sparse · BM25, 24 symbol · exact name FUSE reciprocal rank fusion · k = 60, pool of 24 RANK cross-encoder rerank · drop below 2% of the best · keep 8 · re-seat named symbols
Read it top to bottom. Note the symbol lane bypasses fusion on the right: an exact name match is not a ranked list and should not be treated as one.

Fusing ranks, not scores

Dense search returns cosine distances and BM25 returns term-frequency scores. They are not on the same scale and there is no principled way to add them. Reciprocal rank fusion sidesteps the problem by throwing the scores away: a chunk’s fused score is the sum over lists of 1/(k + rank), with k = 60. A chunk that both searches agree on collects two terms; a chunk one search loved and the other never saw collects one.

The consequence is easiest to see with numbers. A chunk ranked fourth by both searches scores 0.0313. A chunk ranked second by dense search alone scores 0.0161. Agreement beats a single high rank, which is what you want when one of the two searches is known to be confidently wrong on some queries. The fused pool is capped at 24 before reranking, because the reranker is the slow step.

DENSE RANK SPARSE RANK FUSED · 1/(60+r) SUMMED 1  Trainer.fit 2  Problem 3  fit tutorial 4  Node 1  fit tutorial 2  Trainer.fit 3  Loss 4  Node Trainer.fit .0325 fit tutorial .0323 Node .0313 Problem .0161 Loss .0159
Follow Node across the three columns. Fourth place twice beats second place once, which is the whole argument for fusing by rank.

Reranking, then trimming by relative score

Bi-encoder retrieval compares a query vector to document vectors that were computed without seeing the query. A cross-encoder reads the query and the candidate together and is much better at judging relevance, at the cost of one forward pass per pair. The pool of 24 goes through bge-reranker-base, with the provenance header prepended to each candidate so the model can use the file path as evidence. Candidate text is cut at 2,000 characters because the model truncates at 512 tokens anyway and tokenizing the rest is wasted time.

The reranker returns raw logits, often negative. I squash them through a sigmoid so the next rule is meaningful: any chunk scoring below 2% of the best chunk is dropped, subject to keeping at least three. A fixed top-k pads a narrow question with irrelevant filler and gives the model more to misread; a relative floor lets “what does Trainer.fit return” come back with two sources and a broad design question come back with eight.

One correction sits on top of that. The cross-encoder consistently prefers tutorials over raw API definitions, even when the user named the API. So if a symbol hit was demoted out of the final list, up to two of them are swapped back in, replacing the lowest non-symbol entries. It is a small rule, but without it the assistant would explain Trainer.fit from a notebook that calls it instead of from its signature.

SIGMOID SCORE KEPT · 6 OF 10 floor = 2% of best .92 .71 .40 .18 .09 .05 .012 .008 .004 .001
Six of ten survive here. The floor moves with the best score, so a query with one strong match and nine weak ones sends the model one excerpt, not ten.

Follow-ups, prompts, and citations

Chat makes retrieval harder, because “how do I change it” has no search terms. Before searching, generate.py asks a model to rewrite a follow-up into a standalone query, with a prompt whose default is to change nothing. It rewrites only to resolve a pronoun or a bare noun, must copy identifiers character for character since they are matched literally, and is told not to carry a previous topic into a new one. The rewrite runs at temperature zero; in API mode it uses a small model rather than the answering model, since it runs before every search. If it fails or returns something over 300 characters, the original question is searched as typed and the UI says so.

The answering prompt is two files. prompt.txt describes NeuroMANCER’s concepts so the model can help someone frame a problem. mechanics.txt is the grounding contract: every class, function, and keyword argument must appear verbatim in the excerpts, a code fence may open only around code adapted from the excerpts, and if the excerpts do not answer the question the model says so and stops. It also forbids talking about chunks or retrieval at all. The user should read “as described in INSTALLATION.md”, never “as mentioned in excerpt 4”.

Excerpts go into the final user message, each under a heading with its file path and symbol, with the implementation fenced below when there is one. The answer streams token by token from either llama3.1:8b through Ollama or an OpenAI-compatible endpoint, and the file paths behind the excerpts are listed under it in the order they were retrieved.

MESSAGES system · prompt.txt + mechanics.txt history · last three turns user · 8 excerpts + impl, then the question llama3.1:8b or an OpenAI model streamed answer · file paths cited
The excerpts live in the user turn, not the system prompt, so the grounding rules in the system turn are always the first thing the model reads.

Keeping the pull request reviewable

The hardest decision on this project was not technical. A contribution to someone else’s library is only useful if the maintainers can review it, trust it, and keep it, and every extra dependency or touched file is a reason to say no. So I set constraints before writing much code. Nothing under src/ changes. Everything new lives under assistant/. The only edits outside it are one optional dependency group and two .gitignore lines for the generated index.

toml
[project.optional-dependencies]
assistant = [
    "streamlit>=1.40,<2",
    "chromadb>=1.0,<2",
    "rank-bm25>=0.2.2,<0.3",
    "sentence-transformers>=3.0,<6",
    "ollama>=0.4,<1",
    "openai>=1.40,<3",
    "tqdm>=4.66",
]

That decision cost me things. ChromaDB and current Streamlit need Python 3.11, while NeuroMANCER still supports 3.9, so the assistant documents a separate environment instead of raising the library’s floor. A full local install is about 6.3 GB, mostly the 4.9 GB answering model, which is a lot to ask of someone who just wants to try it. That is why setup.sh --api exists: it skips the local answering model, brings the install to about 1.4 GB, and takes an API key in the app’s sidebar. Ollama is still required on that path, because query embeddings always run locally and the index has to be embedded by the same model that built it.

Setup also re-chunks and rebuilds the index from whatever branch it is run on, so the assistant follows the code instead of a snapshot of it. A maintainer who merges a new API gets it in the assistant the next time they run setup, with no separate data to update.

Testing

The pull request has 70 pytest tests. None of them need Ollama, a GPU, or the network: models, Chroma, and the API client are replaced with small fakes. That is enforced rather than hoped for. An autouse fixture in conftest.py patches socket.connect to fail the test with “stub the client” if anything tries to open a real connection. A second fixture clears every module-level cache before and after each test, including the query-embedding LRU, so no test passes because an earlier one warmed something up.

I wrote the tests against the decisions on this page rather than against coverage, so most test names read as the rule they protect.

Ingestion has 22: chunk boundaries for each corpus, record ids and line ranges, oversized-chunk splitting that must preserve content exactly, and a source file with a syntax error that must produce no records instead of crashing the build. Index building has 8, including test_add_batches_embeds_prefixed_header_but_stores_raw_document and test_load_collection_aborts_on_count_mismatch. Retrieval has 17, covering the fusion math, test_hybrid_merge_does_not_mutate_inputs, the score floor, symbol re-seating, and test_load_bm25_pickle_rejects_stale_index. Generation and the UI have 23, including fallbacks when a rewrite fails or comes back too long, and guards for a missing API key.

INGEST ingest/ · 22 tests seams · ids, line ranges, oversize splits INDEX load_vector_store · 8 header · embedded, never stored RETRIEVE retrieve* · 17 tests ranking · RRF, floor, symbol re-seat GENERATE generate, ui · 23 tests contract · message order, fallbacks
The counts add to 70. The right-hand column is the point: each group pins a decision described earlier on this page, so changing that decision breaks a named test.

One test, in full

The symbol re-seating rule from the reranking section is a good example, because it is easy to get almost right. The obvious implementation appends the demoted symbol to the result, which silently returns more chunks than asked for. The rule is to overwrite the lowest-ranked non-symbol slot. This test stubs every stage, gives the named symbol the lowest rerank score in the pool, and checks that it evicts n2 instead of being appended.

python
# a symbol named in the query is forced in even when the reranker demotes it, by
# overwriting the lowest-ranked non-symbol slot -- so n2 is evicted, not appended
def test_retrieve_reinjects_demoted_symbol_hit(monkeypatch):
    pool = [candidate(f"n{i}") for i in range(5)]
    scores = {"n0": 0.9, "n1": 0.8, "n2": 0.7, "n3": 0.6, "n4": 0.5, "sym1": 0.1}
    stub_stages(monkeypatch, pool, scores, symbol_hits=[candidate("sym1")])

    result = retrieve.retrieve("what does LitTrainer do", top_k=3)

    assert [c.id for c in result] == ["n0", "n1", "sym1"]

A companion test gives four symbol hits and asserts that at most two are re-seated and that no id appears twice, which is the edge case where the replacement loop could otherwise overwrite one symbol with another.

Making it fast enough to run locally

A local assistant that takes a long time per answer does not get used, so I went after the time spent on each question. The answering model runs with an 8,192-token context to fit eight excerpts and their implementations. On a laptop, the key-value cache for that context competes with the model weights for memory. The launcher starts Ollama with an 8-bit quantized KV cache and flash attention, which Ollama requires before it will quantize the cache. That roughly halves the cache’s memory relative to 16-bit at a small cost in precision.

bash
OLLAMA_KV_CACHE_TYPE="${OLLAMA_KV_CACHE_TYPE:-q8_0}" \
OLLAMA_FLASH_ATTENTION="${OLLAMA_FLASH_ATTENTION:-1}" \
    nohup ollama serve >"${OLLAMA_LOG}" 2>&1 &

Both are defaults, not overrides, so someone with a large GPU can turn them off. On the retrieval side, query embeddings are memoized with an LRU cache of 256 entries, because the same cleaned query is embedded again on a retry or a repeated question. The Chroma collections, the BM25 index, the reranker, and the symbol vocabulary are all loaded once at startup by a warmup() call, so the first question does not pay for loading four things. fetch_models.py downloads the reranker and then loads it again with local_files_only=True, so a broken cache fails at setup instead of on someone’s first question.

These settings only affect local mode. In API mode the provider controls the model’s caching, and the speed-ups that remain are the retrieval-side caches.

The code and the discussion are in pull request #290. The assistant/README.md in that branch has the setup steps and its own diagrams of the indexing and answering flows.