Retrieval recall
Whether the system retrieves every adjudicated passage required to answer the task. Reported at passage and document level.
Resources / Benchmark methodology
Sanctum Lex evaluates retrieval, citations, supported conclusions and deliberate refusals separately. This published methodology defines the protocol; it does not claim a score that has not been produced by a versioned run.
Methodology v1.0 · August 2026
A system can retrieve the correct document and still cite the wrong passage; it can cite correctly and still overstate the conclusion. Each failure mode is therefore scored independently.
Whether the system retrieves every adjudicated passage required to answer the task. Reported at passage and document level.
Whether each cited coordinate directly supports the proposition attached to it, checked against a lawyer-authored gold set.
The percentage of material factual propositions supported by the permitted record, excluding headings and procedural language.
Whether the system refuses or qualifies an answer when decisive support is absent, without refusing questions the record can answer.
Precision, recall and F1 for clauses, dates, parties, obligations and defined terms. Exact and lawyer-accepted matches are shown separately.
Latency, throughput, failure rate and resource use recorded independently from answer quality and reported with the hardware configuration.
Evaluation protocol
Matters are de-identified or synthetic, versioned and kept separate from prompt development.
Qualified lawyers identify expected passages, bounded conclusions, conflicts and questions that must be refused.
Model, prompts, retrieval parameters, hardware, integration state and policy version are recorded before execution.
The evaluated system receives only the permitted corpus and task. Human graders do not alter outputs after the run.
Two reviewers score independently; disputed labels are resolved and inter-rater agreement is published.
Results identify dataset composition, confidence intervals, exclusions, failures and whether any result uses demonstration data.
REPORTING RULE
Demonstration values illustrate interface behaviour. Benchmark values require this protocol and a versioned result file. Deployment values require telemetry from the named client environment and its stated measurement window. Sanctum Lex does not relabel one as another.
We can construct an evaluation set with your lawyers, freeze the configuration and return a reviewable evidence pack.