Cite the page or say nothing: 7 RAG refusals that were pipeline bugs
Forwarded this? The FAQ at the end is the two-minute version, and the live demo takes two minutes more.
Picture a project engineer asking the archive: "What did the low bidders bid for structural steel on our last two water jobs?"
A wrong unit price is worse than no number. It gets copied into the estimate and the bid, and nobody checks it until it costs money. So when I built Corpus by SalemWise, a hybrid search RAG system that answers from an engineering firm's project archive and cites the page, I started from one rule:
Every answer cites a document and page, or the system says "Nothing on file."
The rule sounds like a prompt instruction. It turned out to be mostly retrieval and plumbing: in the first weeks, almost every wrong refusal I traced came from the pipeline, not the model. This post covers 7 such bugs from two rounds of testing on the demo archive, plus one bug that did the opposite: it showed an answer that should have been a refusal. The lesson under all of them: a wrong refusal is the cheapest retrieval bug report you will get.
You can try everything here on the live demo: 57 public engineering documents, 7,225 PDF pages, 7,107 of them searchable.

Why are engineering archives hard for RAG?
- Literal tokens matter. Users ask for clause IDs ("Section 2-09.3(1)E"), item numbers and dollar values. Embeddings blur exactly these.
- Decades of scans. Many plan sets and old reports are images with no text layer.
- Access is not uniform. A project under a litigation hold must stay invisible to a staff engineer, even indirectly.
How does page-cited hybrid retrieval work?
Indexing: PDF pages, OCR only where there is no text layer
overlapping chunks in two indexes:
vector index (HNSW, recall checked)
full-text index (BM25 keyword ranking)
Asking: follow-up questions rewritten to stand alone
projects this user may not see removed before search, not after
full-text and vector search, merged by reciprocal-rank fusion
the top passages to the model, the first few pages whole
self-hosted LLM answers as JSON, citing passage numbers
a value is shown only when two readings agree
no answer in the passages: "Nothing on file."
The answer model is open-weight, on our own NVIDIA GPU server (CUDA sm_120 series), at temperature 0.
Hybrid retrieval with reciprocal-rank fusion. Full-text search catches clause IDs and numbers; vector search catches paraphrases. The fusion started as plain reciprocal-rank fusion with k = 60, from Cormack, Clarke and Büttcher, "Reciprocal Rank Fusion outperforms Condorcet and individual rank learning methods" (SIGIR 2009). Each chunk scores 1/(k + rank) in each list that contains it. With k = 60 and short candidate lists, that has a side effect: any chunk found by both searches outscores every chunk found by only one, however high it ranked there. Bug 6 is what that cost, and it added one rule: each search gets a few guaranteed places, and the rest merge by fused score.
The citation check is narrow. The rule is the design goal, not a guarantee. The code enforces only two things. Citations that point outside the passage list are dropped. And if the model claims an answer but cites no usable passage, the answer still ships, with the top retrieved passages attached as its sources, so it can carry a citation the model did not choose. Nothing checks the inline [n] markers, and nothing checks that a cited passage supports the sentence (an entailment check). The highlighted sentence on the cited page is how a person verifies the claim.
Permissions trimmed before retrieval, not filtered after. Rank first and hide later, and three things break: restricted chunks crowd visible ones out of the top-k; fewer hits hint that something is hidden; an exact-phrase lookup that runs before the main search can "succeed" on a restricted chunk, so the main search never runs and the user gets nothing. Filtering inside the SQL query and the vector metadata filter removes all three. One caveat: BM25's corpus statistics (term and length counts) still include hidden rows, so they can nudge the order of visible results; nothing from a hidden row is returned or shown. Where that matters, use a separate index per group. The demo deliberately shows each user how many documents were searched, so visitors can see the permission trim working.

Round 1: running the demo
1. Table-of-contents chunks outranked content
On the first live run, three demo questions came back as refusals. The retrieval cause: table-of-contents chunks. A ToC page lists every section title, so it matches almost any question about its document but holds no content. About 2.9% of chunks were ToC pages. A smaller one: fusion identified passages by document and page only, so two passages from one dense page collapsed into one, and sometimes the lost one held the answer. Now the key also includes the start of the text; from scratch I would use a chunk id.
The fix dropped ToC chunks, pulled deeper candidate lists before fusion and used the finer key. More fundamentally, some of the three questions asked for content the public PDFs do not contain; those were bad test questions, not bugs, so I replaced them.
The gate was built with this fix: real questions run against the live server, each checked for the right citation or the right refusal. After the fix it passed 7 of 7. More on it after bug 3.
Lesson: when the system refuses, read the retrieved passages before touching the prompt.
2. The price was in the part we cut off
A question about the structural-steel price trend got a cited answer about 2 runs in 5. It looked like model randomness; it was truncation. The step that assembled passages for the prompt trimmed each chunk shorter than it was indexed. In a flattened bid tabulation, the row with the prices ($1.42 and $3.45 per pound) sat in the tail that was cut, every time. Whether the model answered anyway depended on sampling; the temperature was 0.1 then.
Sending whole chunks fixed the cause: the bid row was in the passages on 8 of 8 runs. The same commit also set the temperature to 0 and added a prompt rule for flattened tables, so I can't say which change fixed the answers. The passage fix alone is proven: the bid row now reaches the model every time, whatever the temperature or prompt. After it, all 7 gate questions held over 8 runs each (56 calls).
Lesson: "flaky" answers are often deterministic truncation plus sampling. Log exactly what the model saw.
3. The cache made a borderline answer drift
One demo question started refusing on one server process only, with a byte-identical prompt. A restart fixed it. llama.cpp's server reuses the KV cache of a matching prompt prefix. The likely mechanism: the reused prefix was computed under a different batch shape than a fresh prefill, and logits are not bit-identical across batch sizes, so even at temperature 0 a borderline answer could tip into a refusal. The demo's inference server also runs other workloads, and cross-request batching can still vary the numerics; the cache was the part I could remove. With prompt caching off for these calls, a repeat run of the gate passed every question. The durable fix is better passages, not luckier numerics. Further reading: Thinking Machines, "Defeating Nondeterminism in LLM Inference" (2025).
Lesson: on shared inference servers, caching is part of your correctness surface.
The opposite failure: a draft inside an unfinished thought
The model reasons inside <think> tags. When a long trace hit the token limit, the reply ended inside an unclosed <think> that sometimes held a draft answer with "found": true, and the parser showed it. Now a reply cut off mid-reasoning counts as no answer, which becomes a refusal. Simplified from the real code:
def parse_llm_json(content):
text = _THINK.sub("", content or "").strip() # _THINK strips closed <think> blocks
if "<think>" in text: # cut off mid-reasoning: not an answer
return None
...
Corpus does not yet check finish_reason == "length"; a stricter guard would, and would also handle chat templates that open <think> in the prompt.
Lesson: never let a truncated generation look like a success.
How do you debug RAG refusals like these?
Unit tests caught none of them. Live runs did: first the demo itself, then the small live gate built with the bug 1 fix. It asks real questions against the running server and checks that answers cite the right document, refuse when they should, and keep the restricted project invisible to the wrong user. A repeat mode runs every question N times; that showed bug 2 was not random. It now runs 18 checks, and one of our recent full runs passed every check.
It is a regression gate, not a retrieval benchmark: a fully green gate said nothing about pages the index had never seen. I have not A/B-tested hybrid against vector-only retrieval on this archive; hybrid is a design choice for literal clause IDs and prices, not a measured win.

Round 2: questions the gate had never seen
For round 2, a separate agent that could not see the gate wrote 35 questions over 26 of the 57 PDFs: 30 answerable from one page, 5 whose answer is not in the archive's text. Tracing every wrong or missing answer found four more pipeline bugs, and two that were the model's (see What is left). One caution: each before/after number below is measured on the same questions that exposed the bug, so it is not an independent test.
4. A capped build had quietly left pages out
A question about page 150 of a structural specification was refused: page 150 was not in the index. An early build had a 120-page cap, and later runs skipped any PDF already listed. On my local index, 8 PDFs had 2,281 pages left out, with no error anywhere. The indexer now continues from each PDF's last indexed page, and the live demo searches 7,107 of 7,225 pages, up from about 4,900.
Lesson: compare indexed pages to each PDF's page count.
5. Fixing bug 4 made vector search too shallow
With the missing pages in, the index grew from 13,939 to 20,463 chunks, and two gate checks started failing. The new pages were not the problem. The vector index is HNSW, an approximate search, and at the library's default search depth it found only 33 of the 60 true nearest chunks for that question (recall 0.55), and none from the answer's document. A deeper search took the worst recall over 10 test questions from 0.55 to 0.98, and the gate back to 16 of 16. Ten questions is a small sample.
A script compares vector search with exact search, with and without the permission filter; I run it after each rebuild. At about 20,000 chunks, exact search took about 6 ms per query, a fair alternative; I keep approximate search for firm-scale archives.
Lesson: approximate search is part of your retrieval pipeline. Measure its recall every time the index grows.
6. Fusion buried a strong single-search hit
For a question about a soil-sample table, the answer page was the second vector hit and absent from the keyword results. Eleven chunks of the same 194-page report sat in both lists at middling ranks, each outscored it, and it never reached the model. The fix is the guaranteed places described above. The answer page was among the passages sent to the model for 28 of 30 known-answer questions, up from 26. No question lost its answer page.
Lesson: agreement between retrievers is a good default, not a veto.
7. The right page, the wrong chunk
For several questions, retrieval returned the right page, but as the chunk next to the answer. Corpus refused, or took the neighboring chunk's value. In one memo, the question asked about a sign structure with groundwater at 15 feet; Corpus answered 10 feet, the figure for the next structure. Now the first few retrieved pages go to the model whole, with a length cap: the small-to-big, or parent-document, pattern. On 27 questions with a checkable quote, the answer text was among the passages for 25, up from 23, for a prompt about 21% longer at the median. No question lost its answer text.
Lesson: the unit you rank and the unit you read do not have to be the same.
What is left
Not every miss is retrieval. The model sometimes read the wrong column of a flattened table row. The fix now live on the demo shows a value only when two readings of the passages, worded differently, give the same one, and refuses otherwise; if the second reading times out, the first is shown. It now reads the right column on the question that exposed the bug. On the round 2 questions it removed the last wrong answer and turned one correct answer into a refusal; those are the questions I tuned on, so that is not an independent result. It roughly doubles answer time, to about 8 seconds at the median. It is not a cure: on a larger 552-PDF test archive, when a question does not say which of two valid unit-price columns it means, it still picks one instead of showing both.
The model also sometimes refuses with the evidence in hand when the question names a project by an identifier the page never prints. And a diagnosis on that second question set, graded automatically, with the misses confirmed by hand, points where bugs 6 and 7 did: the right document among the passages, but not the answer page. My read is that ranking pages inside long documents is the next lever.
Checklist for RAG with page citations
- Make refusal a first-class, visible outcome, and treat every wrong refusal as a retrieval bug report.
- Look at what the model received. Truncation, dedup keys and chunk boundaries caused more wrong refusals than the model did.
- Check coverage and recall: every page indexed, and approximate search checked against exact search as the index grows.
- Apply access control before retrieval.
- Treat truncated generations as failures.
- Keep a live gate with a repeat mode, and test with questions the gate has never seen.
- When a wrong number costs more than a refusal, require two readings to agree, and measure what it does to answer time.
FAQ
In plain terms, what does Corpus by SalemWise do?
It searches an engineering firm's own project archive and answers with the document and page it used. When it finds nothing to support an answer, it says "Nothing on file.": a refusal, not a hallucination. The cited page lets a person check the answer.
Why not just ask ChatGPT, Claude or Copilot?
You can, for a few files: given the right PDF, Claude answered all 30 of my test questions. The hard part is knowing which file to give it in an archive of thousands. Corpus searches the whole archive on the firm's own hardware or ours, never sends documents to a public AI API, hides restricted projects before the search runs, and shows the page behind each answer.
Why is "Nothing on file" better than a best guess?
A wrong number looks exactly like a right one, so it gets copied into the estimate or the bid and surfaces months later as a lost bid or a change order. A refusal is visible at once and costs a few minutes of looking, which is what a firm does today without the tool. So Corpus trades some answers for fewer wrong ones: the fix now live added about one refusal per test set and cut the wrong answers, with the same number of correct answers.
Why does a RAG system with page citations refuse questions it should answer?
In the cases I traced, mostly the pipeline: table-of-contents chunks, trimmed passages, prompt-cache drift, missing pages, shallow vector search, fusion and chunk boundaries. Read the passages the model received before you change the prompt.
About the demo archive: the 57 documents (RFQs, geotechnical reports, specifications, bid tabulations, state DOT forms) come from public agencies and are filed under a fictional sample firm, Cedar Run Civil + Water, which did not author them; any resemblance to a real firm is coincidental. Six public contracts are filed as comparables and marked as such to the model. None of it is a customer's archive. Code excerpts are illustrative and simplified.
I'm Yermek Ibray, founder and CEO of SalemWise Solutions; 16+ years in software, formerly Senior Software Engineer at Google and Meta. We build Corpus by SalemWise, which runs on the firm's hardware or ours with an open-weight model.
If someone at your firm owns the project archive, forward this to them. Try it on the live demo. If a principal should see it, the Corpus page describes a $5,000 fixed pilot on one real RFQ.