A completion certificate can be an excellent search result and poor compliance evidence. It names a learner, a course, and a passing result. Those terms make it relevant to training. But the result does not prove that the learner’s needs were analysed before training. The certificate says nothing about who assessed the needs, when the assessment happened, or what action followed.
That small mismatch contains most of the engineering problem. Retrieval relevance asks whether a passage is in the right neighborhood. Evidence sufficiency asks whether it proves the claim for this organization, at the relevant time, through a source another reviewer can inspect. The approach starts there: determine applicability, preserve provenance, extract bounded facts, qualify support, surface contradictions, and open the verdict gate only when its evidence conditions are met.
That distinction keeps a review useful: it connects every claim to a source and every gap to a next question.
The certificate that answers the wrong question
Imagine a reviewer asking: “Show that needs were analysed before the training plan was finalized.” A semantic retriever sees learner, training, and result in a completion certificate and ranks it highly. That ranking is defensible. The certificate records a later event, though; it never identifies who assessed the need, when that happened, or what action followed. The failure begins when a later stage turns topical proximity into proof of an earlier process.
The mistake happens at the hand-off: “retrieve a few chunks, then ask a model for yes or no.” The retrieved passages are candidates. Each one still needs a qualification record: which requirement element it supports, which element is absent, whether another source contradicts it, and who reviewed that judgment. A high similarity score is useful for recall; it is not a compliance status.
The error message matters. “No” sounds like a negative finding about the organization. “I found a related certificate, but it does not establish the actor, timing, or prior analysis” is a narrower claim. It tells a reviewer what the system saw, why it stopped, and what evidence could resolve the gap.
Seven criteria, and the boundary of a pre-audit
The French decree establishing the national quality reference framework organizes the framework around seven quality criteria. In engineering shorthand - not official quotations - they cover:
- public information about the service, access, and results;
- identification of objectives and adaptation to beneficiaries;
- adaptation of welcome, pedagogical support, monitoring, and evaluation;
- adequacy of pedagogical, technical, and supervisory resources;
- staff competence and continuing development;
- engagement with the provider’s professional environment; and
- collection and handling of feedback and complaints.
Seven criteria do not collapse into seven document searches. Applicability can depend on the provider, action category, audience, and new-entrant situation. A requirement can call for a document, an observed practice, an interview answer, or a judgment across several sources. The Ministry’s Qualiopi reading guide v9 describes an audit context that can include interviews, documentary verification, observation, sampling, and particular treatment for new entrants.
A document-grounded pre-audit is most useful when it organizes the file, exposes gaps, and prepares sharper questions. Some claims still need evidence beyond uploaded PDFs.
That is a more useful starting point than a pile of polished summaries. A reviewer should be able to see which page supports a claim, which detail is still absent, and what would resolve the gap. The work is practical: preserve the trail, make uncertainty legible, and turn an ambiguous document request into the next concrete question.
Applicability comes before retrieval
Searching the entire framework before establishing scope creates false gaps. Suppose a requirement does not apply to the organization’s action category. Returning “no evidence found” makes absence look like non-conformity when the correct status may be “not applicable.” The retrieval system never had the right question.
The applicability record should be explicit and reviewable. At minimum it carries the organization profile, action category, applicable criteria, new-entrant status, the rule or reviewer decision that produced the scope, and an uncertainty state. The important order is: establish or review applicability before searching for evidence. If the available context cannot resolve scope, mark it unknown and send it to human review instead of converting uncertainty into a negative verdict.
Scope fixes a later denominator, too. Evidence coverage divides requirements with reviewed evidence by all applicable requirements, rather than every requirement in an undifferentiated catalogue. A scope change creates a new applicability decision so an old score keeps its meaning.
Now repeat that discipline at the document level. A website crawl, uploaded procedure, certificate, and custom reference document have different origins. The pre-audit preserves those origins while applying one scoped requirement set; a source added for context must not drift into evidence for a requirement that was never declared applicable.
Related is not sufficient
To make the separation testable, use a synthetic fixture called ILL-NEEDS-01. The fixture asks whether a document supports an illustrative requirement: evidence of needs analysis before training. Evidence sufficiency checks three required elements: actor, timing before training, and needs-analysis action.
Document A names a training coordinator, an interview, recorded needs, and adaptation before delivery. It is relevant and sufficient for the synthetic requirement. Document B is a completion certificate. It is relevant but insufficient because it proves a later result, not prior analysis. Document C concerns a room and facilities. It is unrelated and insufficient.
Figure 1. A synthetic example of the distinction: retrieval can nominate a candidate; topical relation still needs evidence for the requirement.
The two decisions are separate but not statistically independent: sufficiency presupposes relevance. The point of two fields is operational. Retrieval failures and qualification failures need different debugging, evaluation, and remediation.
Fixture record: SYN-DOC-A
- Page identifier:
SYN-DOC-A#page-002 - Exact excerpt: “Before each training session, the training coordinator interviews the learner, records the identified needs, and adapts the learning plan before delivery.”
- Excerpt encoding: exact UTF-8, 154 bytes
- Excerpt SHA-256:
60ed0957ce5d647adb020cfaaa6adae694677c787c0f1a1bcdeed2f916a0cff8 - Structured fact: “A training coordinator records learner needs and adapts the plan before delivery.”
- Fact encoding: exact UTF-8, 81 bytes
- Fact SHA-256:
919a9a2c7e952cfd5aedc3ee9616d81fd70464fe8d3e822a22119bcc4ab9da55 - Qualification: relevant and sufficient
The excerpt supports all three declared elements, making it eligible evidence within the synthetic exercise.
Fixture record: SYN-DOC-B
- Page identifier:
SYN-DOC-B#page-001 - Exact excerpt: “Certificate of completion: the learner completed the Data Analysis workshop and received a passing result.”
- Excerpt encoding: exact UTF-8, 106 bytes
- Excerpt SHA-256:
a37380948cc1dcebc1e7f552795394ab207e2b48646d2c8cf36a6cd2e2b2c5c5 - Structured fact: “A learner completed training and received a passing result.”
- Fact encoding: exact UTF-8, 59 bytes
- Fact SHA-256:
56591d463030fb093fda7e6620131dc93895054a22c0d61d3e75924812dc2a60 - Qualification: relevant but insufficient
B is the hard negative. Its vocabulary makes it a plausible retrieval result, while its event and timing fail the support test.
Fixture record: SYN-DOC-C
- Page identifier:
SYN-DOC-C#page-004 - Exact excerpt: “Room checklist: projector tested, seating arranged, and emergency exits kept clear.”
- Excerpt encoding: exact UTF-8, 83 bytes
- Excerpt SHA-256:
b1ca93e2c20198c4b5a6eeb26d1052440fff9c6d1c32c54a2c0afa3abd371f0e - Structured fact: “A room and its facilities were checked.”
- Fact encoding: exact UTF-8, 39 bytes
- Fact SHA-256:
8beecd2ae11d7b247aaf04e04cdb86956324a14b66f2fd99836ff6b9976b8774 - Qualification: unrelated and insufficient
C checks the easier rejection path and takes less reviewer attention than the related certificate.
Lineage before language
A citation is useful only if another person can recover the evidence that was judged. For each candidate, I would preserve a source identifier, stable page or section identifier, exact excerpt, UTF-8 byte count, content hash, chunk identifier, and derived fact identifier. Downstream records then point from fact to requirement, evidence decision, verdict, and remediation.
The hash protects identity, not meaning. It can show that the reviewed excerpt has not changed; a hash is not evidence that the passage is true or sufficient. Qualification remains a reviewed semantic decision.
The hash input is the text inside the typographic quotation marks; the quotation marks, Markdown emphasis, and list punctuation are presentation delimiters and are excluded.
This compact standard-library snippet computes the frozen certificate source record.
from hashlib import sha256
excerpt = "Certificate of completion: the learner completed the Data Analysis workshop and received a passing result."
encoded = excerpt.encode("utf-8")
record = {
"source_id": "SYN-DOC-B",
"page_id": "SYN-DOC-B#page-001",
"excerpt_bytes": len(encoded),
"excerpt_sha256": sha256(encoded).hexdigest(),
}
assert record["excerpt_bytes"] == 106
assert record["excerpt_sha256"] == "a37380948cc1dcebc1e7f552795394ab207e2b48646d2c8cf36a6cd2e2b2c5c5"
Figure 2. A reference design binds SYN-DOC-B to its provenance and to a bounded decision.
That lineage survives re-indexing. A vector-store position is an implementation detail; source + page + hash is the durable review handle. If extraction changes, the system can show that a new fact came from a new excerpt version rather than overwriting the basis of an earlier verdict.
Facts and contradictions
Chunks are retrieval units, not ideal reasoning units. The model keeps each structured fact with the actor, event or action, object, time, and source span when those fields exist. Missing fields stay missing. The certificate fact says a learner completed training and received a passing result; extraction cannot invent a coordinator or a pre-training date.
Once facts have fields, contradictions become inspectable. A procedure may say the coordinator records needs before every session while a dated case file shows the assessment happened afterward. Both passages can be relevant. The contradiction cannot be averaged away into “probably conforming” or resolved by keeping whichever source scores higher. Preserve the pair, identify the conflicting field, and ask a reviewer to resolve it.
The schema also separates a candidate’s missing fields from a reviewed corpus-level gap. One candidate that fails to name an actor triggers abstention, not incomplete. After a reviewer examines the declared evidence set, “no reviewed source names the actor” may support incomplete; a reviewed “Source X says before delivery; source Y says after delivery” contradiction may support non-conforming. The findings call for different questions and different remediation.
Where AuditMalin fits
AuditMalin is built around this pre-audit problem: turning scattered evidence into indicator-level, cited guidance. The public experience centers on grounded review and clear next steps.
The AuditMalin homepage presents an AI-assisted Qualiopi pre-audit workflow: collect organization context and evidence, analyze at indicator level, provide cited guidance, and prioritize remediation. The public application surface exposes organization profile inputs, action categories, applicable criteria, new-entrant status, website acquisition, document uploads, an analysis tier, custom reference document upload, and additional instructions.
That workflow connects evidence, gaps, and next questions. The lineage, contradiction model, and state machine in this article show the kind of record that makes the hand-off easier to inspect.
A verdict is a gated state transition
A verdict is a state change with prerequisites, not the next token after retrieval. Start at unverified and record both the evidence scope and its review status before choosing a branch. A single candidate with insufficient or uncertain support - including SYN-DOC-B on its own - routes to abstain; it does not establish incomplete. The incomplete state is available only after a reviewer has examined the declared scoped evidence set and confirmed a requirement-level gap. Sufficient reviewed support can support conforming. A reviewed contradiction can support non-conforming after review.
For SYN-DOC-B, the declared support is insufficient, so the path is SYN-DOC-B → abstain → human review, with its citation and missing elements preserved. This path says exactly what was judged: one retrieved candidate. It leaves the corpus-level outcome open until the declared evidence set has actually been examined.
Here is a second fixture-only example. It keeps retrieval relevance distinct from the sufficiency gate.
from dataclasses import dataclass
REQUIRED = frozenset({"actor", "timing before training", "needs-analysis action"})
@dataclass(frozen=True)
class Candidate:
relevant: bool
support: frozenset[str]
def gate(candidate: Candidate) -> str:
if not candidate.relevant:
return "unrelated"
if not REQUIRED <= candidate.support:
return "abstain"
return "eligible-for-human-reviewed-verdict"
assert gate(Candidate(True, frozenset())) == "abstain"
assert gate(Candidate(False, frozenset())) == "unrelated"
assert gate(Candidate(True, REQUIRED)) == "eligible-for-human-reviewed-verdict"
Figure 3. The discriminator keeps candidate-level insufficient or uncertain support on the abstain path and reserves corpus-level states for reviewed evidence.
The gate controls when the system may present a documentary judgment and what provenance must accompany it. The reviewer still decides whether to accept the procedure, reject the certificate, or request another sample.
Abstention needs a useful exit
For SYN-DOC-B, preserve SYN-DOC-B#page-001, record the missing actor, pre-training timing, and needs-analysis action, then request a dated needs-analysis record naming the actor and pre-training action. That is the useful exit: claim, inspected candidate, missing evidence, and next action travel together.
The stored remediation text is deliberately terse: Request a dated needs-analysis record naming the actor and pre-training action.
Compare it with “more documents.” The precise request tells the organization which record would close the gap and why. If a suitable record does not exist, the same instruction becomes a process improvement: define who performs the analysis, capture it before delivery, and retain the dated result.
The output record should therefore contain applicability, status, claim, source and page, excerpt or structured fact, sufficiency rationale, missing elements, remediation, and review state. A user can then work from incomplete to review-ready without confusing a recommended next action with an audit finding.
Remediation also closes the provenance loop. When a new record arrives, it receives its own source identifier and hash, and the requirement is reevaluated against the new evidence set. The earlier abstention remains explainable rather than disappearing from history.
A test that could prove the design wrong
A useful evaluation begins with a blinded fixture, clear denominators, and a record of every reviewer decision. Cover applicable and non-applicable requirements, sufficient evidence, related hard negatives, unrelated negatives, missing documents, and seeded contradictions. Preserve the source bytes and expected lineage before evaluation. Qualified reviewers independently label applicability, citation support, sufficiency, contradictions, and verdict state; adjudication is recorded separately rather than replacing the original ratings.
That makes the review process clearer and gives every later result useful operational context for readers.
Report these measures with their denominators:
| Measure | Formula |
|---|---|
| Citation precision | qualified supported citations / all returned citations |
| Unsupported-verdict rate | non-abstained verdicts without sufficient reviewed support / all non-abstained verdicts |
| Evidence coverage | applicable requirements with reviewed evidence / all applicable requirements |
| Selective risk | errors among non-abstained cases / all non-abstained cases |
| Selective coverage | cases receiving a verdict / all evaluated cases |
| Contradiction precision | correctly flagged reviewed contradiction pairs / all flagged contradiction pairs |
| Contradiction recall | correctly flagged seeded contradiction pairs / all seeded contradiction pairs |
| Krippendorff’s alpha | 1 - observed disagreement / expected disagreement |
Plot selective risk against selective coverage as the abstention threshold changes. A system that issues fewer verdicts may reduce error by declining hard cases; reporting only risk would reward silence, while reporting only coverage would reward unsupported confidence. The curve makes that trade-off visible without pretending one threshold suits every deployment.
For contradiction evaluation, declare the number of seeded contradictions before testing. Precision without recall can look strong if the system flags one obvious pair and misses the rest. For Krippendorff’s alpha, publish the reviewer count, item count, missing ratings, and adjudication count, and note when expected disagreement makes the statistic undefined.
The design is falsifiable. It fails if relevant hard negatives repeatedly cross the sufficiency gate, if returned citations do not support their claims, if contradictory facts are flattened into a confident verdict, or if reviewers cannot reproduce lineage from the public record. It also fails operationally if abstentions omit a resolvable evidence request. Those outcomes should change the gate or evidence model, not be hidden inside an aggregate assistant score.
An implementation order
Start with the audit trail rather than the prose generator:
- Define an applicability record and require a reviewed scope or explicit unknown state.
- Give every source and page a stable identifier; hash the exact bytes used for review.
- Extract structured facts without filling absent actors, actions, or dates.
- Retrieve candidates per applicable requirement and keep relevance separate from sufficiency.
- Store support, missing elements, and contradiction links as inspectable fields.
- Gate public statuses on declared evidence; route uncertainty to abstention and human review.
- Produce a precise remediation request tied to the missing elements.
- Evaluate citations, unsupported verdicts, coverage, selective risk, contradictions, and reviewer agreement with declared denominators.
Generation comes after those controls. Its job is to explain a bounded record in useful language, not to invent the record. That ordering makes a wrong answer diagnosable: scope, retrieval, extraction, qualification, contradiction handling, state gating, or explanation.
The completion certificate still establishes that training occurred and a result was recorded; it cannot carry the earlier needs-analysis claim by itself. The design fails if a related completion certificate can cross the support gate without evidence of the actor, pre-training timing, and needs-analysis action.
Primary sources
- French decree establishing the national quality reference framework, for the seven-criteria framework.
- Ministry Qualiopi reading guide v9, for the official reading and audit-context boundary.
- AuditMalin public site, for the product’s observable pre-audit positioning and workflow.
- AuditMalin application surface, for the visible applicability, acquisition, and instruction inputs described above.