NeuroMANCER-GPT
NeuroMANCER is PNNL’s open-source scientific machine learning library for constrained optimization, system identification, and model-based control, with more than 37,000 users. NeuroMANCER-GPT is a retrieval-augmented assistant I built for it: you ask a question in plain English, it searches the library’s own source, docs, and examples, and the model answers only from what it found, with the file paths it drew from listed under the answer. It runs fully local through Ollama or against an OpenAI-compatible API.
I submitted it upstream as pull request #290: +3,576 lines across 30 files, 70 pytest tests, and no changes to the library’s own src/. It is open and awaiting maintainer review. The working branch is assistant/neuromancer-gpt on my fork.
This page is the design write-up. Most of it is about retrieval, because that is where the hard decisions were: how the repository gets chunked, why every documented function is indexed two ways, how three searches are fused and reranked, and what the model is told to do with the result. After that it covers how I kept the pull request small enough to review, how it is tested, and how it was made fast enough to run on a laptop.
The problem
Ask a general-purpose chat model how to set up a Problem with a soft constraint in NeuroMANCER and it will usually answer confidently and wrongly. It invents constructors and keyword arguments that look plausible and do not exist, because a specialized research library is barely represented in training data. The failure is not that the model says “I don’t know”; it is that it never does.
The fix is to stop asking the model to remember the library and make it read it instead. Retrieve the relevant code and docs from the repository at question time, put them in front of the model, and forbid it from naming anything that is not in front of it. Then every claim in an answer traces back to a file a user can open, and when the repository has no answer, the assistant says so.
Where the original script stopped
The repository already shipped a preprocessing script when I started. It walked the checkout, converted notebooks to commented Python with pandoc, and concatenated everything into three files: docs.txt, src.txt, and examples.txt, with a line of dashes between sources. The README said the intent plainly: upload these to a RAG-as-a-service product and let it do the rest.
That design pushes the hard decisions onto whichever service receives the files, and those services make the same decision every time: split on a fixed character window. A window does not know where a function ends, so a docstring lands in one chunk and its signature in the next. Nothing carries a file path or a symbol name, so an answer cannot cite anything more specific than the whole 5 MB file. The rest of this page is what it took to make the pipeline local, structured, and specific.
What was already there, and what I built
I did not start from nothing, and the pull request says so. The repository’s assistant/ directory already had the preprocessing script above, a carefully written system prompt describing NeuroMANCER’s concepts, and a stated goal: lower the barrier to using LLM assistants with the library. What it did not have was the retrieval pipeline it assumed someone would run elsewhere.
My pull request supplies that pipeline and keeps the original intent. The system prompt is preserved unchanged as prompt.txt and is still what the model reads first. ingest.py and ipynb_filter.py are replaced, because three concatenated text files are not a usable unit of retrieval. Everything else is new: the structural chunker, the index builder, hybrid retrieval and reranking, query rewriting, the grounding prompt, the Streamlit app, the setup scripts, and the tests. I designed and wrote all of it; the review I have asked for is from the library’s maintainers.
Chunking along the code's own seams
The new ingest.py keeps the original walker and ignore lists, but instead of appending file text it asks each corpus how it is structured. Python source goes through the standard library’s ast and becomes one record per top-level class or function and one per method, with the class name prepended so Trainer.fit is a real symbol rather than a bare fit. Private names are indexed only if they have a docstring or take arguments; dunder methods other than __init__ are skipped.
Documentation splits at headings. The script reads both reStructuredText underlines and Markdown hashes, and it drops files that are nothing but Sphinx automodule directives, since a page with eight lines of prose and a dozen directives embeds as noise. Notebooks are the interesting case. A markdown cell is glued to the code cell that follows it, and a markdown-only stretch shorter than 200 characters is carried forward into the next chunk, so a heading never becomes its own orphan record. Large example scripts over 200 lines split per top-level definition; anything still over 6,000 characters is cut on line boundaries, because the embedding model would only see the front of it anyway.
Every record gets an id of the form path:symbol:line, and the file path, symbol, source type, and line range travel with it as metadata. That metadata is what the UI later prints as a citation. On a current checkout this yields about 1,900 chunks.
Two views of every function
The first version of the source chunker embedded the full body of every function. It retrieved badly, and the reason was subtle: a signature-plus-docstring chunk and a full-implementation chunk for the same function are near-duplicates, so they crowd each other in the top of every result list and push out the tutorial the user actually needed. The fix was to stop pretending they are two documents.
For a documented symbol, only the signature and docstring are embedded, because that is the text a question resembles. The implementation is attached to the same record as an impl field, stored in the vector database as metadata, and never searched. When that chunk is retrieved, the prompt builder appends the implementation inside a code fence, so the model still sees the body. Undocumented symbols have no docstring to embed, so they are indexed on their code instead and tagged impl rather than api.
if doc and doc.strip():
# embed only the signature + docstring;
# carry the implementation along for display instead of indexing
# a near-duplicate chunk that competes in search
api_content = extract_signature_and_doc(node, source_lines)
record = make_record(rel_path, "api", symbol_name, start, end, api_content)
if len(impl_content) > MAX_IMPL_CHARS:
impl_content = impl_content[:MAX_IMPL_CHARS] + "\n# ... truncated ..."
record["impl"] = impl_content
records.append(record)
else:
# no docstring: a bare signature won't match queries, index the code
records.append(
make_record(rel_path, "impl", symbol_name, start, end, impl_content)
)Implementations are capped at 12,000 characters in the record and 2,000 in the prompt. The cap in the prompt matters more than it looks: eight excerpts with full bodies would not fit the 8,192-token context the local model runs with.
One corpus, two indexes
load_vector_store.py reads the three JSONL files into three Chroma collections, one per corpus, so a query can weight them separately later. Embeddings come from nomic-embed-text through Ollama. Nomic models expect a task prefix, so every document is embedded as search_document: followed by a provenance header (File: ... | Symbol: ... | Type: ...) and then the text, while queries get search_query:. The header is deliberately part of what is embedded and not part of what is stored: it lets a question that mentions a file or symbol name pull the right chunk, without polluting the excerpt the model reads.
The same header-plus-text is tokenized into lowercase word tokens and fed to BM25Okapi. The pickle that comes out stores only chunk ids, collection names, and the BM25 statistics; document text is fetched back from Chroma at query time. That keeps the file small and cold starts fast. It also stores the chunk count it was built against, and retrieval refuses to use it if the live collections disagree, which is the failure you get when someone re-ingests and forgets to rebuild.
Why one search is not enough
retrieve.py runs three searches for every question and they exist because each one fails differently. Dense search handles paraphrase: “how do I add a constraint” finds the Constraint docs even though the word “add” appears nowhere in them. But vector similarity blurs rare identifiers; DictDataset and DictDatasetLoader embed almost identically. Sparse search is the opposite: BM25 will not connect “penalty” to “soft constraint”, but it will find the one tutorial that says DictDataset six times.
Symbol search covers the case the other two both get wrong, which is an API name typed verbatim. At startup the app reads every symbol_name from the source collection into a set. A query is tokenized on identifier characters, including dotted paths, and any token that is in the set becomes an exact metadata lookup. Someone who types LitTrainer.fit gets that definition in the pool no matter what the embedding thinks the question is about.
Dense and sparse each return 24 candidates. Symbol hits are added afterwards, deduplicated by id, so they never displace a fused result before reranking gets a look at them.
Fusing ranks, not scores
Dense search returns cosine distances and BM25 returns term-frequency scores. They are not on the same scale and there is no principled way to add them. Reciprocal rank fusion sidesteps the problem by throwing the scores away: a chunk’s fused score is the sum over lists of 1/(k + rank), with k = 60. A chunk that both searches agree on collects two terms; a chunk one search loved and the other never saw collects one.
The consequence is easiest to see with numbers. A chunk ranked fourth by both searches scores 0.0313. A chunk ranked second by dense search alone scores 0.0161. Agreement beats a single high rank, which is what you want when one of the two searches is known to be confidently wrong on some queries. The fused pool is capped at 24 before reranking, because the reranker is the slow step.
Reranking, then trimming by relative score
Bi-encoder retrieval compares a query vector to document vectors that were computed without seeing the query. A cross-encoder reads the query and the candidate together and is much better at judging relevance, at the cost of one forward pass per pair. The pool of 24 goes through bge-reranker-base, with the provenance header prepended to each candidate so the model can use the file path as evidence. Candidate text is cut at 2,000 characters because the model truncates at 512 tokens anyway and tokenizing the rest is wasted time.
The reranker returns raw logits, often negative. I squash them through a sigmoid so the next rule is meaningful: any chunk scoring below 2% of the best chunk is dropped, subject to keeping at least three. A fixed top-k pads a narrow question with irrelevant filler and gives the model more to misread; a relative floor lets “what does Trainer.fit return” come back with two sources and a broad design question come back with eight.
One correction sits on top of that. The cross-encoder consistently prefers tutorials over raw API definitions, even when the user named the API. So if a symbol hit was demoted out of the final list, up to two of them are swapped back in, replacing the lowest non-symbol entries. It is a small rule, but without it the assistant would explain Trainer.fit from a notebook that calls it instead of from its signature.
Follow-ups, prompts, and citations
Chat makes retrieval harder, because “how do I change it” has no search terms. Before searching, generate.py asks a model to rewrite a follow-up into a standalone query, with a prompt whose default is to change nothing. It rewrites only to resolve a pronoun or a bare noun, must copy identifiers character for character since they are matched literally, and is told not to carry a previous topic into a new one. The rewrite runs at temperature zero; in API mode it uses a small model rather than the answering model, since it runs before every search. If it fails or returns something over 300 characters, the original question is searched as typed and the UI says so.
The answering prompt is two files. prompt.txt describes NeuroMANCER’s concepts so the model can help someone frame a problem. mechanics.txt is the grounding contract: every class, function, and keyword argument must appear verbatim in the excerpts, a code fence may open only around code adapted from the excerpts, and if the excerpts do not answer the question the model says so and stops. It also forbids talking about chunks or retrieval at all. The user should read “as described in INSTALLATION.md”, never “as mentioned in excerpt 4”.
Excerpts go into the final user message, each under a heading with its file path and symbol, with the implementation fenced below when there is one. The answer streams token by token from either llama3.1:8b through Ollama or an OpenAI-compatible endpoint, and the file paths behind the excerpts are listed under it in the order they were retrieved.
Keeping the pull request reviewable
The hardest decision on this project was not technical. A contribution to someone else’s library is only useful if the maintainers can review it, trust it, and keep it, and every extra dependency or touched file is a reason to say no. So I set constraints before writing much code. Nothing under src/ changes. Everything new lives under assistant/. The only edits outside it are one optional dependency group and two .gitignore lines for the generated index.
[project.optional-dependencies]
assistant = [
"streamlit>=1.40,<2",
"chromadb>=1.0,<2",
"rank-bm25>=0.2.2,<0.3",
"sentence-transformers>=3.0,<6",
"ollama>=0.4,<1",
"openai>=1.40,<3",
"tqdm>=4.66",
]That decision cost me things. ChromaDB and current Streamlit need Python 3.11, while NeuroMANCER still supports 3.9, so the assistant documents a separate environment instead of raising the library’s floor. A full local install is about 6.3 GB, mostly the 4.9 GB answering model, which is a lot to ask of someone who just wants to try it. That is why setup.sh --api exists: it skips the local answering model, brings the install to about 1.4 GB, and takes an API key in the app’s sidebar. Ollama is still required on that path, because query embeddings always run locally and the index has to be embedded by the same model that built it.
Setup also re-chunks and rebuilds the index from whatever branch it is run on, so the assistant follows the code instead of a snapshot of it. A maintainer who merges a new API gets it in the assistant the next time they run setup, with no separate data to update.
Testing
The pull request has 70 pytest tests. None of them need Ollama, a GPU, or the network: models, Chroma, and the API client are replaced with small fakes. That is enforced rather than hoped for. An autouse fixture in conftest.py patches socket.connect to fail the test with “stub the client” if anything tries to open a real connection. A second fixture clears every module-level cache before and after each test, including the query-embedding LRU, so no test passes because an earlier one warmed something up.
I wrote the tests against the decisions on this page rather than against coverage, so most test names read as the rule they protect.
Ingestion has 22: chunk boundaries for each corpus, record ids and line ranges, oversized-chunk splitting that must preserve content exactly, and a source file with a syntax error that must produce no records instead of crashing the build. Index building has 8, including test_add_batches_embeds_prefixed_header_but_stores_raw_document and test_load_collection_aborts_on_count_mismatch. Retrieval has 17, covering the fusion math, test_hybrid_merge_does_not_mutate_inputs, the score floor, symbol re-seating, and test_load_bm25_pickle_rejects_stale_index. Generation and the UI have 23, including fallbacks when a rewrite fails or comes back too long, and guards for a missing API key.
One test, in full
The symbol re-seating rule from the reranking section is a good example, because it is easy to get almost right. The obvious implementation appends the demoted symbol to the result, which silently returns more chunks than asked for. The rule is to overwrite the lowest-ranked non-symbol slot. This test stubs every stage, gives the named symbol the lowest rerank score in the pool, and checks that it evicts n2 instead of being appended.
# a symbol named in the query is forced in even when the reranker demotes it, by
# overwriting the lowest-ranked non-symbol slot -- so n2 is evicted, not appended
def test_retrieve_reinjects_demoted_symbol_hit(monkeypatch):
pool = [candidate(f"n{i}") for i in range(5)]
scores = {"n0": 0.9, "n1": 0.8, "n2": 0.7, "n3": 0.6, "n4": 0.5, "sym1": 0.1}
stub_stages(monkeypatch, pool, scores, symbol_hits=[candidate("sym1")])
result = retrieve.retrieve("what does LitTrainer do", top_k=3)
assert [c.id for c in result] == ["n0", "n1", "sym1"]A companion test gives four symbol hits and asserts that at most two are re-seated and that no id appears twice, which is the edge case where the replacement loop could otherwise overwrite one symbol with another.
Making it fast enough to run locally
A local assistant that takes a long time per answer does not get used, so I went after the time spent on each question. The answering model runs with an 8,192-token context to fit eight excerpts and their implementations. On a laptop, the key-value cache for that context competes with the model weights for memory. The launcher starts Ollama with an 8-bit quantized KV cache and flash attention, which Ollama requires before it will quantize the cache. That roughly halves the cache’s memory relative to 16-bit at a small cost in precision.
OLLAMA_KV_CACHE_TYPE="${OLLAMA_KV_CACHE_TYPE:-q8_0}" \
OLLAMA_FLASH_ATTENTION="${OLLAMA_FLASH_ATTENTION:-1}" \
nohup ollama serve >"${OLLAMA_LOG}" 2>&1 &Both are defaults, not overrides, so someone with a large GPU can turn them off. On the retrieval side, query embeddings are memoized with an LRU cache of 256 entries, because the same cleaned query is embedded again on a retry or a repeated question. The Chroma collections, the BM25 index, the reranker, and the symbol vocabulary are all loaded once at startup by a warmup() call, so the first question does not pay for loading four things. fetch_models.py downloads the reranker and then loads it again with local_files_only=True, so a broken cache fails at setup instead of on someone’s first question.
These settings only affect local mode. In API mode the provider controls the model’s caching, and the speed-ups that remain are the retrieval-side caches.
The code and the discussion are in pull request #290. The assistant/README.md in that branch has the setup steps and its own diagrams of the indexing and answering flows.