Maintenance technician reviewing equipment manuals in an illustrative workshop scene; not a field test
AI and Machine Learning

Offline RAG for Equipment Manuals: Versioned Evidence Before Answers

A technician asks how to reset an E42 alarm. A local model produces a convincing procedure, but cannot identify the machine variant, manual edition, or page it used.

RAGDocument IndexingRetrieval Evaluation
AI workflow bridge

Turning this AI workflow into a production system?

We help define knowledge sources, permissions, model routing, API tools, monitoring, and handoff rules so the workflow can survive real operations.

Explore AI Application Development

A technician asks how to reset an E42 alarm. A local model produces a convincing procedure, but cannot identify the machine variant, manual edition, or page it used. That is not a usable maintenance answer. With no network connection, there may be no online manual to catch the mistake before someone follows the wrong instruction.

The useful product is not “chat on an edge box.” It is a version-aware, access-controlled route to verifiable manual passages. Local generation can make those passages easier to read, but only after the system has selected the right document and established that the user may see it. If the model, manual edition, authorization, or safety prerequisites are uncertain, the interface should expose the gap rather than turn a plausible sentence into an instruction. The architecture below is based on public interfaces and security guidance; no field accuracy or maintenance outcome is claimed.

The first failure is usually document fit, not model fluency

An alarm code can appear in more than one equipment family. Even within a model line, a firmware update may change the condition under which a reset is safe. Splitting all PDFs into chunks and searching for the closest sentence treats semantic similarity as if it were equipment compatibility. Those are different questions.

Before retrieval, the application needs as much of the asset context as the site can establish: model and variant, hardware revision, firmware version, site or asset identifier, and the user's role. Missing identifiers do not have to prevent a document search. They do, however, prevent the system from presenting a candidate procedure as definitely applicable to the machine in front of the technician. A useful response may instead ask the technician to check a nameplate or firmware screen and show the candidate manuals with their editions.

This leads to a firm separation of responsibility. The authorization layer decides which material the user is allowed to see. Retrieval decides which passages are relevant to the identified machine and question. The language model may summarize only those authorized passages. It must not invent a document edition, grant itself access through a prompt, or use a similar machine's procedure to fill a missing step.

Technical architecture diagram
Technical architecture diagram

A knowledge package needs more than extracted text

The index should preserve a document identifier, edition, supported models, section and page or stable anchor, language, effective date, access classification, and a checksum of the extracted passage. A citation that names only a PDF is weak: two editions can share a title while disagreeing on a step. A technician should be able to open the cited passage in the exact document edition used to form the answer.

Structure the source before choosing a chunk size

A manual can place an alarm definition in one table, a prerequisite in a warning box, and the procedure several pages later. Fixed-length chunks can separate the warning from the instruction. Parse the document hierarchy, tables, safety notices, and procedure boundaries before deciding what a retrievable unit is. Where one action depends on a nearby safety condition, keep an explicit link between them. For scanned manuals, record extraction uncertainty and retain a link to the page image. An OCR guess about a voltage or torque value must not silently become a definite answer.

Edition relationships are not simply “newer replaces older.” Older equipment can still require an older manual, and a service bulletin may supersede only one procedure. Package metadata should say where each passage applies and how a bulletin relates to a base edition. When an update is installed, switch the source documents, extracted text, indexes, access labels, and citation targets as one coherent version. Rolling back only the PDF while leaving newer vectors or page anchors active could present an old answer with a new citation. This is a proposed application contract, not a feature guaranteed by any retrieval engine.

Two maintenance-manual editions side by side, illustrating why provenance and page location matter

Exact identifiers and natural-language symptoms need different retrieval paths

An error code such as E42, a connector such as CN7, and a version string should remain exact search terms. A symptom such as “the compressor starts and then stops repeatedly” may use different words from the manual, so semantic retrieval can help discover relevant sections. Official SQLite FTS5 documentation describes full-text querying and BM25 ranking. Qdrant's hybrid-query documentation shows how multiple retrieval paths can be fused. These interfaces make a hybrid design possible; they do not establish that it is more accurate for a particular manual collection.

For a pilot, keep exact asset and document filters ahead of ranking, then compare exact-term, semantic, and fused results on labeled questions. Do not translate every retriever score into a single “confidence” percentage. A BM25 score and a vector similarity are not automatically comparable probabilities. If a semantically similar passage for the wrong machine outranks the exact code in the correct manual, a fluent answer will hide the retrieval error. Log which candidate came from which path, what filters were applied, and which ranking version selected it.

Whether to add a retrieval orchestration framework is a separate choice. Our LlamaIndex RAG knowledge-base guide helps assess that choice, but a framework does not replace equipment-edition and authorization filters.

Authorization belongs inside retrieval, not just on the answer screen

An end-user guide, a certified-service procedure, and a customer-specific maintenance bulletin may all describe the same machine. Searching all of them and removing restricted text only after generation is too late: the restricted passage has already entered the model context. Candidate generation must be constrained by an application-verified authorization scope, and opening a source citation or reading a cached answer must enforce that scope again.

Each indexed passage should carry organization, asset scope, document classification, and allowed-role metadata. When disconnected, the device also needs a defined offline authorization policy: how credentials are validated, when they expire, how revocation is handled after reconnection, and how local keys are protected. If a restricted entitlement has expired and cannot be revalidated, the system should narrow retrieval to public material or refuse access. “The site is offline” is not a reason to expose service-only instructions. Qdrant's filtering documentation establishes that payload filters can be part of a search query; identity, credential issuance, and revocation remain application responsibilities.

Answer caching must follow the same boundary. A second user must not receive the first user's restricted answer because they asked the same question. The cache key and validity policy should include authorization scope and knowledge-package edition. If an administrator withdraws a document, the product must define what happens to its cached answers, offline copies, and old citation links. A prompt saying “respect permissions” does not implement any of these controls.

Make abstention a designed outcome

Generation is useful when an answer needs to combine passages across a manual, but it must not create a repair sequence absent from the source. A minimum answer identifies the applicable machine and edition, states the action or diagnostic clue, links each consequential claim to a passage, and names unresolved prerequisites. When the manual describes a symptom but not a remedy, the system should stop at the symptom instead of borrowing a procedure from a related model.

Local inference is not proof of trustworthiness. The llama.cpp server documentation describes local serving and embedding endpoints, making it one possible runtime component. Whether a suitable model fits a target device or handles the site's manual language well must be measured on that device. Retrieved content may also contain obsolete directions or malicious instructions aimed at the model. OWASP's LLM Top 10 warns that RAG does not eliminate prompt-injection risk. A manual passage is evidence to quote, not a new system instruction.

The interface should withhold an operational answer when the equipment edition is unknown, the user lacks access, source passages conflict, or essential safety prerequisites are missing. It can still identify what must be checked next, show an accessible manual table of contents, or route the issue to a qualified person. Reconnecting to a cloud service can be a separate escalation path, but it must not silently upload restricted documents or use a cloud answer to bypass an unresolved local evidence gap.

Treat package updates as a version switch, not a file sync

A deployable knowledge package contains linked objects: source files, extracted passages, chunk identifiers, full-text or vector indexes, access metadata, citation targets, and the model or prompt-policy version used by the answer layer. Replacing only the PDF while keeping an old index can retrieve text that no longer exists. Replacing only the index can leave a citation pointing to the wrong page. Both failures are hard to spot if the UI shows only a polished answer.

One auditable design downloads a candidate package to an inactive area, verifies its signature and checksums, checks machine compatibility, imports or builds its indexes, and runs a small known-question smoke test before switching the active package identifier. If import fails, the old package stays active and its edition remains visible to users. If a serious content error appears after activation, rollback restores the previous complete package, including index and citation metadata. Device-clock errors, insufficient disk space, and interrupted transfers are part of this design review; a valid signature alone does not prove that activation finished correctly.

Fleet rollout of firmware and model artifacts has its own failure modes, covered in our Edge AI OTA rollout and rollback guide. That release path does not replace the knowledge package's requirement to keep passages, access labels, indexes, and citation targets in sync.

Whether incremental updates are worth their complexity depends on package size, connectivity, and index format. A small pilot may be safer with full-package replacement because it avoids a mixed old/new state. Delta updates become interesting only when full downloads or downtime are genuinely too costly. Full-package rollback also needs spare storage. A resource-constrained box cannot promise both minimal storage and an always-ready previous package without making that tradeoff explicit.

Test retrieval before measuring the language model

Without a de-identified manual corpus, labeled questions, and correct source passages, there is no basis for an accuracy claim. The first pilot asset should therefore be a question set spanning actual model variants, old editions, contradictory bulletins, access levels, and questions for which the manual has no answer. For each question, record the expected passage, who is allowed to view it, and the conditions that require abstention. A few easy demonstration questions are not enough to exercise the riskiest failure paths.

Run retrieval tests without generation first. Does the correct edition enter the candidate set? Are wrong models excluded? Does the citation open at the right page? Are restrictions enforced across every search route? Then test answer behavior: does it stay within authorized evidence, preserve safety prerequisites, and abstain on conflict? Only after these checks should the team measure memory, latency, concurrency, and update rollback on the intended hardware under actual offline conditions. Passing one stage cannot stand in for the others.

Report failures by question type rather than hiding them in one aggregate hit rate. Exact alarm codes, symptom descriptions, multi-section questions, conflicting editions, and unanswerable queries have different failure causes. Save the query, asset context, authorization scope, candidates, citation targets, package edition, and model version for a failed case. Those records show whether the problem came from extraction, chunking, retrieval, access control, or generation; without them, prompt tuning becomes guesswork.

When local question answering is the wrong next step

An offline assistant becomes less reliable when manuals change frequently but field devices cannot receive verified packages. Some service actions require a live work-order approval or a safety permit from a central system; a local model cannot replace that check. Small edge devices may also be unable to store the corpus, index, and generation model together. In those situations, offline full-text search with edition filtering and page-level navigation may serve the technician better than a constrained chat model.

For a workflow dominated by alarm codes and page lookups, start with exact retrieval and accessible PDFs. Generation adds model distribution, compute cost, prompt-injection exposure, and answer auditing. It is justified when users truly need cross-passage explanations and the team can carry those obligations. Even then, the product promise should be to help a person find and verify applicable evidence sooner, not to transfer maintenance responsibility to a model.

Official references