Skip to content

Resources / Benchmark methodology

A result is useful only when the test can be inspected

Sanctum Lex evaluates retrieval, citations, supported conclusions and deliberate refusals separately. This published methodology defines the protocol; it does not claim a score that has not been produced by a versioned run.

Responsible AI

Methodology v1.0 · August 2026

Six measures, reported without collapsing them into one marketing number

A system can retrieve the correct document and still cite the wrong passage; it can cite correctly and still overstate the conclusion. Each failure mode is therefore scored independently.

MEASURE 01

Retrieval recall

Whether the system retrieves every adjudicated passage required to answer the task. Reported at passage and document level.

MEASURE 02

Citation precision

Whether each cited coordinate directly supports the proposition attached to it, checked against a lawyer-authored gold set.

MEASURE 03

Grounded-answer rate

The percentage of material factual propositions supported by the permitted record, excluding headings and procedural language.

MEASURE 04

Refusal accuracy

Whether the system refuses or qualifies an answer when decisive support is absent, without refusing questions the record can answer.

MEASURE 05

Extraction quality

Precision, recall and F1 for clauses, dates, parties, obligations and defined terms. Exact and lawyer-accepted matches are shown separately.

MEASURE 06

Operational performance

Latency, throughput, failure rate and resource use recorded independently from answer quality and reported with the hardware configuration.

Evaluation protocol

Reproducible by design

  1. 01Freeze the test set

    Matters are de-identified or synthetic, versioned and kept separate from prompt development.

  2. 02Create the gold set

    Qualified lawyers identify expected passages, bounded conclusions, conflicts and questions that must be refused.

  3. 03Lock the configuration

    Model, prompts, retrieval parameters, hardware, integration state and policy version are recorded before execution.

  4. 04Run blind

    The evaluated system receives only the permitted corpus and task. Human graders do not alter outputs after the run.

  5. 05Adjudicate disagreements

    Two reviewers score independently; disputed labels are resolved and inter-rater agreement is published.

  6. 06Publish limitations

    Results identify dataset composition, confidence intervals, exclusions, failures and whether any result uses demonstration data.

REPORTING RULE

Demonstration, benchmark and deployment metrics remain separate.

Demonstration values illustrate interface behaviour. Benchmark values require this protocol and a versioned result file. Deployment values require telemetry from the named client environment and its stated measurement window. Sanctum Lex does not relabel one as another.

Test the claim against your own work.

We can construct an evaluation set with your lawyers, freeze the configuration and return a reviewable evidence pack.