Marcel Bautista-Kuljevan Claude Opus 5 (1M context) commited on
Commit
29827d9
·
unverified ·
1 Parent(s): 8901ef3

La démo devient un poste de revue humaine (#4)

Browse files

* Le chemin vision : la démo pouvait recevoir des scans, jamais les envoyer

saknussemm livre une chaîne vision complète — VisionEditProducer qui découpe
chaque ligne, VISION_SYSTEM_PROMPT dont la première phrase est « l'image fait
foi », GuardConfig.vision() qui descend le seuil de similarité de 0,35 à 0,15
parce qu'une lecture CORRECTE d'une ligne massacrée s'éloigne forcément du
texte OCR, et core/batching.py qui découpe un chunk pour respecter
ModelCapabilities.max_images.

Rien de tout ça ne pouvait tourner d'ici. Les quatre fournisseurs
n'implémentent que complete_structured, donc ce backend acceptait les images
de page, les rangeait, les servait à l'interface — et n'en a jamais envoyé
une seule à un modèle.

Ce n'est pas une régression : complete_structured_multimodal n'apparaît
JAMAIS dans l'historique de ce dépôt. La capacité a vécu 23 jours comme
outil en ligne de commande à la racine du dépôt de la bibliothèque
(scripts/run_vision.py + scripts/providers_multimodal.py, du 24 juillet au
16 août), voisin de la démo et non partie d'elle, et elle est partie avec
l'archive du banc — emportée par un défaut prouvé chez sa VOISINE :
vision_benchmark.py construisait son manifeste depuis l'ALTO de référence
puis écrasait le texte par celui de l'OCR. Le client, lui, ne porte pas ce
défaut. Personne n'a relogé la capacité chez un consommateur.

Ce commit la reloge, à l'endroit que la scission désigne.

Quatre pièces, dont trois apprises en se trompant d'abord :

1. app/providers/mistral_multimodal.py — un MultimodalStructuredClient, frère
du fournisseur texte plutôt qu'un élargissement : la bibliothèque garde sa
couture texte sans image exprès. Le transport est base.call_llm, ce qui est
une amélioration sur le client archivé — il retire un paramètre
d'échantillonnage refusé en lisant le message du fournisseur, au lieu d'une
liste de modèles codée en dur qui avait déjà périmé.

2. La branche vision du runner. for_provider ne peut pas servir : il construit
toujours le producteur TEXTE. Le producteur est donc assemblé sur place,
avec max_images pris du fournisseur — c'est lui qui fait découper le chunk
— et GuardConfig.vision().

3. page_image_assets(). Ce backend clé ses images par FICHIER SOURCE, la
bibliothèque les veut par PAGE PHYSIQUE, et require_page_images dit
pourquoi : « flattening them to a single per-file ref sent the producer the
wrong image for every page but the first ». Un fichier à plusieurs pages est
donc REFUSÉ, pas aplati. Une ligne corrigée contre le mauvais scan revient
confiante et fausse, et l'invariant de projection ne peut pas le voir.

4. L'extra [vision] dans les cinq lignes d'installation de la CI. Sans Pillow,
toute tâche vision échoue au premier découpage avec ModuleNotFoundError —
mesuré, pas supposé. Et la note de requirements.txt disait encore
« packages/saknussemm », un chemin disparu le 16 août : personne n'aurait pu
la suivre.

Deux faits viennent du client archivé et sont crédités sur place, parce que
les redécouvrir a coûté cher : Mistral refuse une neuvième image (400, code
3051), et le scan de la page ENTIÈRE est la mauvaise image — attaché à un
chunk avec le prompt texte, le modèle décrit la photographie au lieu de
corriger, une légende revenant réécrite de bout en bout.

tools/run_corpus_arms.py lance un corpus à travers le JobRunner de la démo,
reprenant là où il s'est arrêté ; la clé vient du trousseau macOS et
n'apparaît jamais dans un drapeau, ce qui la mettrait dans l'historique du
shell et dans la liste des processus.

11 tests neufs, 485 verts au total.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw3bbyo2tm2Ujnzu58rLA

* La revue humaine : le verdict dans /layout, et un endroit pour trancher

Ces corpora n'ont AUCUNE vérité terrain. Le text.txt que Gallica livre est
la même couche OCR que son ALTO, donc rien sur le disque ne peut dire si une
correction était juste ou un refus fondé. Un lecteur devant le scan est la
seule source. Deux pièces pour qu'il puisse travailler, et pour que ce qu'il
lit s'accumule au lieu de s'évaporer.

1. /layout porte désormais le VERDICT et la proposition.

La géométrie seule ne dit pas POURQUOI une ligne a gardé son texte OCR.
« Rien n'a été proposé », « une garde a refusé une hallucination » et « un
couple de césure n'a pas pu se réconcilier » sont identiques sur la page et
appellent des jugements différents. Quatre champs par ligne : verdict,
verdict_detail, proposed_text, et proposal_declined — ce dernier étant le cas
le plus intéressant pour un relecteur, quelque chose était sur la table et a
été refusé. Optionnel : un job dont le rapport n'est pas encore écrit rend
toujours sa mise en page.

La garde d'instantané a attrapé le changement de forme, ce qui est le rôle
qu'on lui a donné ; sa liste de clés est mise à jour avec la raison.

2. Un endroit où le jugement se range : PUT/GET /api/jobs/{id}/reviews.

Trois verdicts, et le troisième est le plus utile. `accepted` et `refused`
notent ce que le moteur a fait. `transcribed` porte ce que le lecteur a lu à
l'image, ce qui vaut par soi-même quoi que le moteur ait décidé — et
s'accumule en vérité terrain ligne à ligne. Une transcription vide est
refusée en 422 : le texte EST la revue, et l'accepter remplirait le jeu de
lignes qui n'affirment rien, indiscernables plus tard des vraies.

Idempotent par ligne : renvoyer la même ligne remplace sa revue au lieu de
l'empiler, pour qu'un lecteur qui change d'avis ne se batte pas contre un
journal. La date est estampillée par le serveur, pas reçue du client.

Clé sur la PAIRE (page_id, line_id), séparée par un NUL. ADR-001 : un line_id
se répète d'un fichier à l'autre, donc la clé sur l'identifiant nu
fusionnerait deux jugements le jour où un job porte plus d'un ALTO. Et un NUL
plutôt qu'une espace parce qu'un identifiant peut légalement contenir une
espace — un test pin les deux.

Au passage, deux corrections dans ce qui a été livré hier.

L'ImageAsset était construit SANS transform. Une bibliothèque numérique sert
une dérivée réduite — le IIIF de Gallica rend un JPEG 1193x1600 pour une page
dont l'ALTO mesure 6802x9121, soit un facteur 0,1754. Sans transform le
découpage se fait à l'échelle 1,0, donc une ligne à hpos=4798 est découpée
dans une image large de 1193 : hors du cadre, et chaque découpage revient
vide ou pris ailleurs. Rien en aval ne s'en aperçoit : le VLM reçoit une
image de rien, dit quelque chose de plausible sur le texte OCR qu'on lui
donne aussi, et les gardes jugent une proposition faite sans la preuve
qu'elles croient. Le facteur est maintenant DÉDUIT — la page déclare ses
dimensions, le scan décodé déclare ses pixels — plutôt que demandé à un
appelant qui peut l'ignorer.

Et les fixtures d'image des tests vision étaient des en-têtes fabriqués à la
main. Ça suffisait tant que le code ne reniflait que le nombre magique ; ça a
cessé de suffire quand page_image_assets s'est mis à DÉCODER le scan. Un
fixture qu'on ne peut pas ouvrir échoue alors pour une raison sans rapport
avec ce qui est testé.

492 tests verts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw3bbyo2tm2Ujnzu58rLA

* La vue de revue : colorée par verdict, filtrée en estompant

Le curseur d'opacité existait déjà et pilote tout le calque — il reste tel
quel. Ce qui manquait est ce qu'il y a SOUS le curseur.

Coloré par verdict, en trois familles, parce qu'un relecteur cherche trois
choses différentes et ne devrait pas avoir à lire une légende pour les
distinguer :

retenue la correction a survécu à toutes les gardes
refusée quelque chose a été proposé et une garde l'a décliné — les cas
qui valent un œil humain, chacun étant soit une hallucination
attrapée, soit une bonne correction jetée
sans objet rien n'a été proposé, ou rien n'a changé

Les couleurs sont les deux crayons du correcteur d'épreuve : le bleu marque
ce qui tient, le rouge ce qui a été biffé.

Le filtre ESTOMPE au lieu de masquer, et c'est le point de conception. Une
ligne qui a recopié sa voisine ne se lit comme fausse qu'À CÔTÉ de cette
voisine — retirer le contexte retire la preuve. Éteindre la dernière famille
est refusé : une page blanche se lirait comme un bug.

Une ligne est cliquable et se sélectionne, épaissie ; c'est l'accroche du
panneau de jugement qui vient ensuite.

La garde de couleur du test a attrapé le changement, ce qui est son rôle.
Elle est mise à jour sur ce qu'elle doit vérifier maintenant, et deux tests
s'ajoutent : qu'une ligne refusée se distingue d'une retenue, et que le
filtre estompe sans supprimer — le compte de rectangles reste identique, et
un groupe passe à 0,16.

131 tests frontend verts, typecheck propre.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw3bbyo2tm2Ujnzu58rLA

* Le panneau de jugement : la revue humaine devient de la donnée

La vue montrait ce que le moteur avait décidé ; elle ne permettait pas d'en
juger. Un relecteur qui ne peut que regarder produit une impression. Un
relecteur qui peut trancher produit le jeu qui manque à ces corpora :
(ligne, ce que l'OCR a lu, ce que le modèle a proposé, ce qu'un humain dit
que le scan montre). C'est exactement l'entrée dont un banc a besoin pour
répondre « la correction était-elle juste », et qu'aucune mesure interne au
moteur ne peut fabriquer.

Ce que le panneau montre, au clic sur une ligne : le texte OCR, ce qui a été
PROPOSÉ, ce qui a été retenu, et le verdict du moteur. La proposition est
mise en avant quand elle a été refusée — c'est le cas qui vaut un œil
humain, chacun étant soit une hallucination attrapée, soit une bonne
correction jetée, et rien d'autre qu'une lecture du scan ne les sépare.

T

.github/workflows/ci.yml CHANGED
@@ -52,7 +52,7 @@ jobs:
52
  python-version: "3.11"
53
  cache: pip
54
  - name: Install saknussemm (editable, sibling package)
55
- run: pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
56
  - name: Install backend dependencies
57
  working-directory: backend
58
  run: pip install -r requirements.txt -r requirements-dev.txt
@@ -75,7 +75,7 @@ jobs:
75
  python-version: "3.11"
76
  cache: pip
77
  - name: Install saknussemm (editable, sibling package)
78
- run: pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
79
  - name: Install backend dependencies
80
  working-directory: backend
81
  run: pip install -r requirements.txt -r requirements-dev.txt
@@ -112,7 +112,7 @@ jobs:
112
  python-version: "3.11"
113
  cache: pip
114
  - name: Install saknussemm (editable, sibling package)
115
- run: pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
116
  - name: Install backend dependencies
117
  working-directory: backend
118
  run: pip install -r requirements.txt -r requirements-dev.txt
@@ -135,7 +135,7 @@ jobs:
135
  # resolve. Audit the resolved environment after install
136
  # (pip-audit -r doesn't understand editable lines).
137
  run: |
138
- pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
139
  pip install -r backend/requirements.txt
140
  pip install bandit[toml]==1.9.4 pip-audit==2.10.0
141
  - name: bandit (static analysis)
@@ -217,7 +217,7 @@ jobs:
217
  cache: npm
218
  cache-dependency-path: frontend/package-lock.json
219
  - name: Install saknussemm (editable, sibling package)
220
- run: pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
221
  - name: Install backend dependencies
222
  working-directory: backend
223
  run: pip install -r requirements.txt -r requirements-dev.txt
 
52
  python-version: "3.11"
53
  cache: pip
54
  - name: Install saknussemm (editable, sibling package)
55
+ run: pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
56
  - name: Install backend dependencies
57
  working-directory: backend
58
  run: pip install -r requirements.txt -r requirements-dev.txt
 
75
  python-version: "3.11"
76
  cache: pip
77
  - name: Install saknussemm (editable, sibling package)
78
+ run: pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
79
  - name: Install backend dependencies
80
  working-directory: backend
81
  run: pip install -r requirements.txt -r requirements-dev.txt
 
112
  python-version: "3.11"
113
  cache: pip
114
  - name: Install saknussemm (editable, sibling package)
115
+ run: pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
116
  - name: Install backend dependencies
117
  working-directory: backend
118
  run: pip install -r requirements.txt -r requirements-dev.txt
 
135
  # resolve. Audit the resolved environment after install
136
  # (pip-audit -r doesn't understand editable lines).
137
  run: |
138
+ pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
139
  pip install -r backend/requirements.txt
140
  pip install bandit[toml]==1.9.4 pip-audit==2.10.0
141
  - name: bandit (static analysis)
 
217
  cache: npm
218
  cache-dependency-path: frontend/package-lock.json
219
  - name: Install saknussemm (editable, sibling package)
220
+ run: pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
221
  - name: Install backend dependencies
222
  working-directory: backend
223
  run: pip install -r requirements.txt -r requirements-dev.txt
backend/app/api/jobs.py CHANGED
@@ -778,7 +778,9 @@ async def get_job_layout(job: JobManifest = Depends(get_completed_job)) -> dict:
778
  if job.document_manifest is None:
779
  raise HTTPException(status_code=500, detail="Job has no document_manifest.")
780
  # Wave-3 review — same offload rationale as /diff.
781
- data = await asyncio.to_thread(build_layout, job.job_id, job.document_manifest, job.images)
 
 
782
  # Plan V2.4 — <img> cannot set headers: append a short-lived,
783
  # images-scoped signed credential to each URL (this response is
784
  # itself token-gated, so only the job owner receives them).
 
778
  if job.document_manifest is None:
779
  raise HTTPException(status_code=500, detail="Job has no document_manifest.")
780
  # Wave-3 review — same offload rationale as /diff.
781
+ data = await asyncio.to_thread(
782
+ build_layout, job.job_id, job.document_manifest, job.images, job.report
783
+ )
784
  # Plan V2.4 — <img> cannot set headers: append a short-lived,
785
  # images-scoped signed credential to each URL (this response is
786
  # itself token-gated, so only the job owner receives them).
backend/app/api/read_models.py CHANGED
@@ -11,7 +11,7 @@ job, guard the HTTP preconditions, and delegate here.
11
 
12
  from __future__ import annotations
13
 
14
- from app.schemas import DocumentManifest, HyphenRole
15
 
16
 
17
  def build_diff(job_id: str, document_manifest: DocumentManifest) -> dict:
@@ -69,13 +69,38 @@ def build_layout(
69
  job_id: str,
70
  document_manifest: DocumentManifest,
71
  images: dict[str, str],
 
72
  ) -> dict:
73
  """Structural layout: blocks + lines with ALTO coordinates, per page.
74
 
75
  Page dimensions are derived from line coordinates when the source Page
76
  element omits WIDTH/HEIGHT. ``images`` maps source_file → image filename;
77
  a matching entry becomes the page's ``image_url``.
 
 
 
 
 
 
 
 
78
  """
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
79
  pages_out = []
80
  for page in document_manifest.pages:
81
  line_by_id = {lm.line_id: lm for lm in page.lines}
@@ -99,6 +124,15 @@ def build_layout(
99
  "corrected_text": corrected,
100
  "modified": corrected != lm.ocr_text,
101
  "hyphen_role": lm.hyphen_role.value,
 
 
 
 
 
 
 
 
 
102
  }
103
  )
104
  blocks_out.append(
 
11
 
12
  from __future__ import annotations
13
 
14
+ from app.schemas import CorrectionReport, DocumentManifest, HyphenRole
15
 
16
 
17
  def build_diff(job_id: str, document_manifest: DocumentManifest) -> dict:
 
69
  job_id: str,
70
  document_manifest: DocumentManifest,
71
  images: dict[str, str],
72
+ report: CorrectionReport | None = None,
73
  ) -> dict:
74
  """Structural layout: blocks + lines with ALTO coordinates, per page.
75
 
76
  Page dimensions are derived from line coordinates when the source Page
77
  element omits WIDTH/HEIGHT. ``images`` maps source_file → image filename;
78
  a matching entry becomes the page's ``image_url``.
79
+
80
+ ``report`` adds, per line, **what the engine decided and why**: the verdict
81
+ code and the text the producer actually proposed. Without it a reviewer
82
+ sees that a line was left alone but not whether that was because nothing
83
+ was proposed, because a guard refused a hallucination, or because a hyphen
84
+ pair could not reconcile — three situations that look identical in the
85
+ geometry and demand different judgements. Optional so a job whose report
86
+ is not yet written still renders its layout.
87
  """
88
+ verdicts: dict[tuple[str, str], dict] = {}
89
+ if report is not None:
90
+ for trace in report.lines:
91
+ reason = trace.decision.reason
92
+ proposed = trace.proposal.output_text if trace.proposal else None
93
+ verdicts[(trace.page_id, trace.line_id)] = {
94
+ "verdict": reason.code if reason is not None else trace.decision.status,
95
+ "verdict_detail": reason.detail if reason is not None else None,
96
+ "proposed_text": proposed,
97
+ # A proposal the engine declined is the reviewer's most
98
+ # interesting case: something was on offer and was refused.
99
+ "proposal_declined": bool(
100
+ reason is not None and proposed is not None and proposed != trace.source_text
101
+ ),
102
+ }
103
+
104
  pages_out = []
105
  for page in document_manifest.pages:
106
  line_by_id = {lm.line_id: lm for lm in page.lines}
 
124
  "corrected_text": corrected,
125
  "modified": corrected != lm.ocr_text,
126
  "hyphen_role": lm.hyphen_role.value,
127
+ **verdicts.get(
128
+ (page.page_id, lm.line_id),
129
+ {
130
+ "verdict": None,
131
+ "verdict_detail": None,
132
+ "proposed_text": None,
133
+ "proposal_declined": False,
134
+ },
135
+ ),
136
  }
137
  )
138
  blocks_out.append(
backend/app/api/review.py ADDED
@@ -0,0 +1,143 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Human review — where a reader's judgement on a line is recorded.
2
+
3
+ The corpora this demo runs on have **no ground truth**. Gallica's `text.txt`
4
+ is the same OCR layer as its ALTO, so nothing on disk can say whether a
5
+ correction was right or a refusal was justified; the only source of truth is
6
+ a person reading the scan. This module is where that reading is kept.
7
+
8
+ **Why it is worth keeping rather than just displaying.** A reviewer who can
9
+ only look produces an impression. A reviewer who can adjudicate produces the
10
+ dataset that is missing: `(line, what the OCR read, what the model proposed,
11
+ what a human says the scan actually shows)`. That is exactly the input a
12
+ bench needs to answer "was the correction right", which no amount of
13
+ measurement inside the engine can answer on its own.
14
+
15
+ **Three verdicts, and the third is the useful one.** `accepted` and
16
+ `refused` grade what the engine did. `transcribed` carries what the reviewer
17
+ read on the image, which stands on its own whatever the engine decided — and
18
+ accumulates into ground truth line by line.
19
+
20
+ Reviews are keyed by ``(page_id, line_id)`` because a line id repeats across
21
+ files (`ADR-001`); keying on the bare id would silently merge two documents'
22
+ judgements the first time a job carries more than one ALTO.
23
+ """
24
+
25
+ from __future__ import annotations
26
+
27
+ from datetime import UTC, datetime
28
+ from enum import StrEnum
29
+
30
+ from fastapi import APIRouter, Depends, HTTPException, status
31
+ from pydantic import BaseModel, Field
32
+
33
+ from app.api.deps import get_job_store
34
+ from app.api.jobs import require_job_access
35
+ from app.protocols import JobStore
36
+ from app.schemas import JobManifest
37
+
38
+ router = APIRouter(prefix="/api/jobs", tags=["review"])
39
+
40
+
41
+ class ReviewVerdict(StrEnum):
42
+ """What the reader concluded about the engine's decision on this line."""
43
+
44
+ #: The engine's outcome is right — whether it corrected or refused.
45
+ ACCEPTED = "accepted"
46
+ #: The engine's outcome is wrong. `note` should say how.
47
+ REFUSED = "refused"
48
+ #: Neither: the reader is recording what the scan actually shows.
49
+ TRANSCRIBED = "transcribed"
50
+
51
+
52
+ class LineReview(BaseModel):
53
+ """One reader's judgement on one line."""
54
+
55
+ page_id: str = Field(min_length=1, max_length=256)
56
+ line_id: str = Field(min_length=1, max_length=256)
57
+ verdict: ReviewVerdict
58
+ #: What the reader read on the image. Required for ``transcribed``; free
59
+ #: to accompany the other two when the reader wants to be precise about
60
+ #: what the engine got wrong.
61
+ transcription: str | None = Field(default=None, max_length=4000)
62
+ note: str | None = Field(default=None, max_length=2000)
63
+ reviewed_at: str | None = None
64
+
65
+ @property
66
+ def key(self) -> tuple[str, str]:
67
+ return (self.page_id, self.line_id)
68
+
69
+
70
+ class ReviewBatch(BaseModel):
71
+ """Reviews arrive in batches: a reader works through a page, not a line."""
72
+
73
+ reviews: list[LineReview] = Field(max_length=2000)
74
+
75
+
76
+ class ReviewsResponse(BaseModel):
77
+ job_id: str
78
+ reviews: list[LineReview]
79
+
80
+
81
+ def _key(review: LineReview) -> str:
82
+ """``(page_id, line_id)`` flattened for storage, NUL-separated.
83
+
84
+ A NUL cannot occur in an XML id, so the pair round-trips unambiguously; a
85
+ space could, and would merge two lines the day a producer emits one.
86
+ """
87
+ return f"{review.page_id}\x00{review.line_id}"
88
+
89
+
90
+ def _existing(job: JobManifest) -> dict[str, LineReview]:
91
+ """Reviews already on the job, back as models.
92
+
93
+ The manifest stores plain dicts so the schema layer never imports the API
94
+ layer's models; re-validating here is what keeps that separation from
95
+ costing type safety at the edge.
96
+ """
97
+ return {key: LineReview.model_validate(raw) for key, raw in (job.reviews or {}).items()}
98
+
99
+
100
+ @router.put("/{job_id}/reviews", response_model=ReviewsResponse)
101
+ async def put_reviews(
102
+ job_id: str,
103
+ batch: ReviewBatch,
104
+ job: JobManifest = Depends(require_job_access),
105
+ store: JobStore = Depends(get_job_store),
106
+ ) -> ReviewsResponse:
107
+ """Record or replace judgements on lines of this job.
108
+
109
+ Idempotent per line: sending the same line twice replaces its review
110
+ rather than appending, so a reader who changes their mind is not fighting
111
+ an append-only log. The timestamp is stamped here rather than trusted
112
+ from the client — a review's date is a fact about the server.
113
+ """
114
+ for review in batch.reviews:
115
+ if review.verdict is ReviewVerdict.TRANSCRIBED and not review.transcription:
116
+ raise HTTPException(
117
+ status.HTTP_422_UNPROCESSABLE_CONTENT,
118
+ f"line {review.line_id!r}: a 'transcribed' review must carry the "
119
+ "text the reader read on the image — that text IS the review.",
120
+ )
121
+
122
+ merged = _existing(job)
123
+ stamped = datetime.now(UTC).isoformat(timespec="seconds")
124
+ for review in batch.reviews:
125
+ merged[f"{_key(review)}"] = review.model_copy(update={"reviewed_at": stamped})
126
+ store.update_job(job_id, reviews={k: v.model_dump() for k, v in merged.items()})
127
+ return ReviewsResponse(
128
+ job_id=job_id, reviews=sorted(merged.values(), key=lambda r: (r.page_id, r.line_id))
129
+ )
130
+
131
+
132
+ @router.get("/{job_id}/reviews", response_model=ReviewsResponse)
133
+ async def get_reviews(
134
+ job_id: str,
135
+ job: JobManifest = Depends(require_job_access),
136
+ ) -> ReviewsResponse:
137
+ return ReviewsResponse(
138
+ job_id=job_id,
139
+ reviews=sorted(_existing(job).values(), key=lambda r: (r.page_id, r.line_id)),
140
+ )
141
+
142
+
143
+ __all__ = ["LineReview", "ReviewBatch", "ReviewVerdict", "router"]
backend/app/jobs/runner.py CHANGED
@@ -26,7 +26,17 @@ from saknussemm import (
26
  )
27
  from saknussemm.core.events import ReconcileStats
28
  from saknussemm.core.protocols import ProviderPermanentError
29
- from saknussemm.core.schemas import PairingPolicy
 
 
 
 
 
 
 
 
 
 
30
 
31
  from app.jobs.events import JobEventType
32
  from app.jobs.observers import CompositeObserver, JobStoreObserver, LoggingObserver
@@ -36,6 +46,92 @@ from app.schemas import DocumentManifest, JobStatus
36
  logger = logging.getLogger(__name__)
37
 
38
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
39
  def _default_timeout_from_env() -> int:
40
  """Resolve JOB_TIMEOUT_SECONDS once at import. 0 disables the timeout.
41
 
@@ -91,6 +187,7 @@ class JobRunner:
91
  source_files: dict[str, Path],
92
  provider: BaseProvider | None = None,
93
  pairing_policy: PairingPolicy | None = None,
 
94
  timeout_seconds: int = 1800,
95
  should_abort: Callable[[], bool] | None = None,
96
  ) -> None:
@@ -110,6 +207,10 @@ class JobRunner:
110
  `should_abort`: Plan V2.2 — cooperative cancellation probe,
111
  forwarded to the pipeline (polled between pages and chunks).
112
  When it trips, the job lands in CANCELLED with no output promoted.
 
 
 
 
113
  """
114
  if provider is None:
115
  from app.providers import get_provider
@@ -132,6 +233,7 @@ class JobRunner:
132
  output_writer=output_writer,
133
  source_files=source_files,
134
  pairing_policy=pairing_policy,
 
135
  should_abort=should_abort,
136
  ),
137
  timeout=timeout,
@@ -318,9 +420,14 @@ class JobRunner:
318
  output_writer: OutputWriter,
319
  source_files: dict[str, Path],
320
  pairing_policy: PairingPolicy | None = None,
 
321
  should_abort: Callable[[], bool] | None = None,
322
  ) -> CorrectionResult:
323
- """Drive the pure pipeline and persist its counters back."""
 
 
 
 
324
  self.job_store.update_job(job_id, status=JobStatus.STARTED)
325
  self.job_store.emit(job_id, JobEventType.STARTED, {"job_id": job_id})
326
 
@@ -338,25 +445,78 @@ class JobRunner:
338
  # §5.1 resorption — credentials go into the producer (via the
339
  # for_provider convenience), never into run(): the pipeline surface
340
  # carries no api_key anywhere.
341
- pipeline = CorrectionPipeline.for_provider(
342
- provider,
343
- api_key=api_key,
344
- model=model,
345
- provider_name=provider_name,
346
- observer=CompositeObserver(
347
- [JobStoreObserver(self.job_store, job_id), LoggingObserver()]
348
- ),
349
- # Provenance parity — the pipeline hashes this into the §11
350
- # config fingerprint. Passing the job's actual policy (default
351
- # or geometric_pairing opt-out) keeps the stamped fingerprint
352
- # honest; omitting it silently reverted to DEFAULT_PAIRING_POLICY.
353
- pairing_policy=pairing_policy,
354
- )
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
355
  # `run_id` is saknussemm's generic identifier; we feed it the
356
  # server-side `job_id` so trace.json correlates with the API.
357
  result = await pipeline.run(
358
  document_manifest=document_manifest,
359
  source_files=source_files,
 
360
  run_id=job_id,
361
  # Plan V2.2 — the cancel endpoint's event, polled by the
362
  # pipeline between pages and chunks.
 
26
  )
27
  from saknussemm.core.events import ReconcileStats
28
  from saknussemm.core.protocols import ProviderPermanentError
29
+ from saknussemm.core.schemas import (
30
+ GuardConfig,
31
+ ImageAsset,
32
+ ImageTransform,
33
+ ModelCapabilities,
34
+ PageImage,
35
+ PageManifest,
36
+ PairingPolicy,
37
+ )
38
+ from saknussemm.errors import ConfigurationError
39
+ from saknussemm.integrations.vision import build_image_asset
40
 
41
  from app.jobs.events import JobEventType
42
  from app.jobs.observers import CompositeObserver, JobStoreObserver, LoggingObserver
 
46
  logger = logging.getLogger(__name__)
47
 
48
 
49
+ def page_image_assets(
50
+ document_manifest: DocumentManifest, images_by_source: dict[str, Path]
51
+ ) -> dict[str, ImageAsset]:
52
+ """``{page_id: ImageAsset}`` from the demo's per-SOURCE-FILE image map.
53
+
54
+ The two models disagree and the disagreement matters. This backend stores
55
+ one image per uploaded source file, keyed by its stem; the library wants one
56
+ per *physical page*, and ``require_page_images`` says why in as many words:
57
+ "never one per source file: a multipage XML has as many scans as pages, and
58
+ flattening them to a single per-file ref sent the producer the wrong image
59
+ for every page but the first".
60
+
61
+ So a source file carrying more than one page is **refused**, not flattened.
62
+ A wrong scan produces a confident correction of a line that is not on it,
63
+ and nothing downstream can see that — the projection invariant compares the
64
+ artefact to the decisions the artefact was built from.
65
+ """
66
+ assets: dict[str, ImageAsset] = {}
67
+ pages_per_source: dict[str, int] = {}
68
+ for page in document_manifest.pages:
69
+ pages_per_source[page.source_file] = pages_per_source.get(page.source_file, 0) + 1
70
+
71
+ for page in document_manifest.pages:
72
+ stem = Path(page.source_file).stem.lower()
73
+ image_path = images_by_source.get(stem)
74
+ if image_path is None:
75
+ continue
76
+ if pages_per_source[page.source_file] > 1:
77
+ raise ValueError(
78
+ f"{page.source_file!r} carries {pages_per_source[page.source_file]} "
79
+ "pages but only one image was uploaded for it. A vision run needs "
80
+ "one scan per physical page; sending the same image for every page "
81
+ "would correct each line against the wrong scan. Upload one file "
82
+ "per page, or run this document without vision."
83
+ )
84
+ asset = build_image_asset(page_id=page.page_id, path=image_path)
85
+ assets[page.page_id] = asset.model_copy(update={"transform": _transform_for(page, asset)})
86
+ return assets
87
+
88
+
89
+ def _transform_for(page: PageManifest, asset: ImageAsset) -> ImageTransform | None:
90
+ """Map the ALTO/PAGE coordinate space onto the scan's pixels.
91
+
92
+ **Not optional, and the reason is a measured near-miss.** A digital
93
+ library serves a downscaled derivative — Gallica's IIIF ``!1600,1600``
94
+ gives a 1193x1600 JPEG for a page whose ALTO measures 6802x9121, a factor
95
+ of **0.1754**. With no transform the crop is taken at scale 1.0, so a line
96
+ at ``hpos=4798`` is cropped from an image 1193 pixels wide: off the
97
+ canvas, and every crop comes back blank or from the wrong place.
98
+
99
+ Nothing downstream notices. The VLM is handed a picture of nothing, says
100
+ something plausible about the OCR text it was also given, and the guards
101
+ judge a proposal made without the evidence they think it was made with.
102
+
103
+ Derived rather than configured: the page declares its own dimensions and
104
+ the decoded scan declares its pixels, so the ratio is available without
105
+ asking a caller for a DPI it may not know. Returns ``None`` when either
106
+ side is missing — a wrong scale is worse than a declared unknown.
107
+ """
108
+ if not (asset.pixel_width and asset.pixel_height):
109
+ return None
110
+ if not (page.page_width and page.page_height):
111
+ return None
112
+ return ImageTransform(
113
+ scale_x=asset.pixel_width / page.page_width,
114
+ scale_y=asset.pixel_height / page.page_height,
115
+ )
116
+
117
+
118
+ def _media_type_for(path: Path) -> str:
119
+ """Decoded from the bytes, not guessed from the extension.
120
+
121
+ A ``.jpg`` that is really a PNG makes the vendor reject the data URI, and
122
+ the schema field documents itself as "determined from the bytes rather than
123
+ guessed from the extension".
124
+ """
125
+ header = path.read_bytes()[:12]
126
+ if header.startswith(b"\x89PNG\r\n\x1a\n"):
127
+ return "image/png"
128
+ if header[:2] == b"\xff\xd8":
129
+ return "image/jpeg"
130
+ if header[:4] in (b"II*\x00", b"MM\x00*"):
131
+ return "image/tiff"
132
+ return "application/octet-stream"
133
+
134
+
135
  def _default_timeout_from_env() -> int:
136
  """Resolve JOB_TIMEOUT_SECONDS once at import. 0 disables the timeout.
137
 
 
187
  source_files: dict[str, Path],
188
  provider: BaseProvider | None = None,
189
  pairing_policy: PairingPolicy | None = None,
190
+ page_images: dict[str, PageImage] | None = None,
191
  timeout_seconds: int = 1800,
192
  should_abort: Callable[[], bool] | None = None,
193
  ) -> None:
 
207
  `should_abort`: Plan V2.2 — cooperative cancellation probe,
208
  forwarded to the pipeline (polled between pages and chunks).
209
  When it trips, the job lands in CANCELLED with no output promoted.
210
+ `page_images`: page_id → scan, from ``page_image_assets()``. Present
211
+ means a VISION run: the producer becomes ``VisionEditProducer`` and the
212
+ guards become ``GuardConfig.vision()``. Absent (the default) keeps the
213
+ historical text path byte for byte.
214
  """
215
  if provider is None:
216
  from app.providers import get_provider
 
233
  output_writer=output_writer,
234
  source_files=source_files,
235
  pairing_policy=pairing_policy,
236
+ page_images=page_images,
237
  should_abort=should_abort,
238
  ),
239
  timeout=timeout,
 
420
  output_writer: OutputWriter,
421
  source_files: dict[str, Path],
422
  pairing_policy: PairingPolicy | None = None,
423
+ page_images: dict[str, PageImage] | None = None,
424
  should_abort: Callable[[], bool] | None = None,
425
  ) -> CorrectionResult:
426
+ """Drive the pure pipeline and persist its counters back.
427
+
428
+ `page_images`: page_id → scan. Present means a VISION run — the
429
+ producer changes, and so does the guard config.
430
+ """
431
  self.job_store.update_job(job_id, status=JobStatus.STARTED)
432
  self.job_store.emit(job_id, JobEventType.STARTED, {"job_id": job_id})
433
 
 
445
  # §5.1 resorption — credentials go into the producer (via the
446
  # for_provider convenience), never into run(): the pipeline surface
447
  # carries no api_key anywhere.
448
+ observer = CompositeObserver([JobStoreObserver(self.job_store, job_id), LoggingObserver()])
449
+ if page_images:
450
+ # Vision run. `for_provider` cannot serve it: it always builds the
451
+ # TEXT producer, around a client whose seam carries no images on
452
+ # purpose. So the producer is assembled here, with the three things
453
+ # a VLM run needs and a text run must not have.
454
+ from saknussemm.integrations.vision import (
455
+ MultimodalStructuredClient,
456
+ VisionEditProducer,
457
+ )
458
+
459
+ # Checked, not assumed. A text provider here would fail at the
460
+ # first call, deep inside a run, with a message about a missing
461
+ # method rather than about the job being misconfigured — and the
462
+ # run would already have spent its first chunk getting there.
463
+ if not isinstance(provider, MultimodalStructuredClient):
464
+ raise ConfigurationError(
465
+ f"a vision run needs a multimodal provider, and "
466
+ f"{type(provider).__name__} implements the text seam only. "
467
+ "Pass MistralMultimodalProvider, or run this job without "
468
+ "page images."
469
+ )
470
+ max_images = getattr(provider, "MAX_IMAGES_PER_CALL", None)
471
+ producer = VisionEditProducer(
472
+ provider=provider,
473
+ api_key=api_key,
474
+ model=model,
475
+ # `max_images` is what makes core/batching.py split a chunk.
476
+ # Without it the engine sends every line's crop in one call and
477
+ # the vendor refuses the request outright.
478
+ capabilities=ModelCapabilities(
479
+ text=True,
480
+ vision=True,
481
+ structured_output=True,
482
+ max_images=max_images,
483
+ ),
484
+ )
485
+ pipeline = CorrectionPipeline(
486
+ producer=producer,
487
+ observer=observer,
488
+ pairing_policy=pairing_policy,
489
+ # A VLM reads the image, so a CORRECT reading of a badly
490
+ # garbled line diverges from the OCR text further than the
491
+ # text-tuned guard tolerates (0.35 → 0.15). Running vision
492
+ # under the text config refuses the model's best work: measured
493
+ # 18 refusals out of 20 lines, every one
494
+ # `too_different_from_source`.
495
+ guard_config=GuardConfig.vision(),
496
+ )
497
+ else:
498
+ # §5.1 resorption — credentials go into the producer (via the
499
+ # for_provider convenience), never into run(): the pipeline surface
500
+ # carries no api_key anywhere.
501
+ pipeline = CorrectionPipeline.for_provider(
502
+ provider,
503
+ api_key=api_key,
504
+ model=model,
505
+ provider_name=provider_name,
506
+ observer=observer,
507
+ # Provenance parity — the pipeline hashes this into the §11
508
+ # config fingerprint. Passing the job's actual policy (default
509
+ # or geometric_pairing opt-out) keeps the stamped fingerprint
510
+ # honest; omitting it silently reverted to
511
+ # DEFAULT_PAIRING_POLICY.
512
+ pairing_policy=pairing_policy,
513
+ )
514
  # `run_id` is saknussemm's generic identifier; we feed it the
515
  # server-side `job_id` so trace.json correlates with the API.
516
  result = await pipeline.run(
517
  document_manifest=document_manifest,
518
  source_files=source_files,
519
+ page_images=page_images,
520
  run_id=job_id,
521
  # Plan V2.2 — the cancel endpoint's event, polled by the
522
  # pipeline between pages and chunks.
backend/app/jobs/store.py CHANGED
@@ -139,6 +139,7 @@ class JobStore:
139
  fallbacks: int | None = None,
140
  duration_seconds: float | None = None,
141
  error: str | None = None,
 
142
  images: dict[str, str] | None = None,
143
  report: CorrectionReport | None = None,
144
  token_hash: str | None = None,
@@ -172,6 +173,8 @@ class JobStore:
172
  job.duration_seconds = duration_seconds
173
  if error is not None:
174
  job.error = error
 
 
175
  if images is not None:
176
  job.images = images
177
  if report is not None:
 
139
  fallbacks: int | None = None,
140
  duration_seconds: float | None = None,
141
  error: str | None = None,
142
+ reviews: dict[str, dict] | None = None,
143
  images: dict[str, str] | None = None,
144
  report: CorrectionReport | None = None,
145
  token_hash: str | None = None,
 
173
  job.duration_seconds = duration_seconds
174
  if error is not None:
175
  job.error = error
176
+ if reviews is not None:
177
+ job.reviews = reviews
178
  if images is not None:
179
  job.images = images
180
  if report is not None:
backend/app/main.py CHANGED
@@ -24,6 +24,7 @@ from app.api.health import router as health_router
24
  from app.api.jobs import router as jobs_router
25
  from app.api.providers import router as providers_router
26
  from app.api.rate_limit import limiter
 
27
  from app.api.upload_guard import UploadAdmissionMiddleware, UploadSizeLimitMiddleware
28
  from app.frontend_static import INDEX_HTML as _INDEX_HTML
29
  from app.frontend_static import STATIC_DIR as _STATIC_DIR
@@ -278,6 +279,9 @@ def create_app() -> FastAPI:
278
  # ------------------------------------------------------------------
279
  app.include_router(providers_router, prefix="/api/providers", tags=["providers"])
280
  app.include_router(jobs_router, prefix="/api/jobs", tags=["jobs"])
 
 
 
281
 
282
  # ------------------------------------------------------------------
283
  # Static frontend (HF Spaces single-container mode)
 
24
  from app.api.jobs import router as jobs_router
25
  from app.api.providers import router as providers_router
26
  from app.api.rate_limit import limiter
27
+ from app.api.review import router as review_router
28
  from app.api.upload_guard import UploadAdmissionMiddleware, UploadSizeLimitMiddleware
29
  from app.frontend_static import INDEX_HTML as _INDEX_HTML
30
  from app.frontend_static import STATIC_DIR as _STATIC_DIR
 
279
  # ------------------------------------------------------------------
280
  app.include_router(providers_router, prefix="/api/providers", tags=["providers"])
281
  app.include_router(jobs_router, prefix="/api/jobs", tags=["jobs"])
282
+ # Human review sits on the same job paths and carries its own prefix,
283
+ # so it registers without one here.
284
+ app.include_router(review_router)
285
 
286
  # ------------------------------------------------------------------
287
  # Static frontend (HF Spaces single-container mode)
backend/app/protocols/job_store.py CHANGED
@@ -43,6 +43,7 @@ class JobStore(Protocol):
43
  error: str | None = None,
44
  images: dict[str, str] | None = None,
45
  report: Any | None = None,
 
46
  token_hash: str | None = None,
47
  ) -> None: ...
48
 
 
43
  error: str | None = None,
44
  images: dict[str, str] | None = None,
45
  report: Any | None = None,
46
+ reviews: dict[str, Any] | None = None,
47
  token_hash: str | None = None,
48
  ) -> None: ...
49
 
backend/app/providers/mistral_multimodal.py ADDED
@@ -0,0 +1,165 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Mistral multimodal provider — the half of the library the demo could not run.
2
+
3
+ ``saknussemm`` ships a complete vision chain: ``VisionEditProducer`` crops each
4
+ line from the page scan, ``VISION_SYSTEM_PROMPT`` tells the model the image is
5
+ authoritative, ``GuardConfig.vision()`` loosens the similarity guard because a
6
+ correct reading of a garbled line necessarily diverges from the OCR text, and
7
+ ``core/batching.py`` splits a chunk to respect ``ModelCapabilities.max_images``.
8
+
9
+ None of it could run from here. Every bundled provider implements
10
+ ``complete_structured`` only, so the backend accepted page images, stored them,
11
+ served them to the UI — and never sent one to a model.
12
+
13
+ **This is not a regression.** ``complete_structured_multimodal`` has never
14
+ appeared in this repository's history. The capability lived for 23 days as a CLI
15
+ tool at the root of the *library* repo (``scripts/run_vision.py`` +
16
+ ``scripts/providers_multimodal.py``, 2026-07-24 to 2026-08-16), a sibling of the
17
+ demo rather than a part of it, and left with the bench archive — swept along by a
18
+ defect proven in its *neighbour*: ``vision_benchmark.py`` built its manifest from
19
+ the reference ALTO and then overwrote the text with the OCR's, a state no real
20
+ run reaches. The client itself carries no such defect. Nothing re-homed the
21
+ capability in a consumer. This does.
22
+
23
+ **Credit where it is due.** Two facts below come from that archived client
24
+ (``cinoc/campaigns/tooling/providers_multimodal.py``), which measured them
25
+ against the live API rather than reading them off a doc page. Both are
26
+ load-bearing and both were rediscovered the hard way before the archive was
27
+ read:
28
+
29
+ * **A ninth image is refused**, HTTP 400 ``"Total number of images exceeds the
30
+ maximum allowed of 8."`` (code 3051). So :data:`MAX_IMAGES_PER_CALL` is 8 and
31
+ is declared through ``ModelCapabilities.max_images``, which is what makes the
32
+ engine split a chunk instead of issuing a request that cannot succeed.
33
+ * **The whole-page scan is the wrong image.** Attaching one page to a chunk and
34
+ asking for OCR corrections makes the model *describe* the photograph: a
35
+ caption came back rewritten from `PHOTOGRAPHIE DE L'ÉBOULEMENT D'UNE FALAISE A
36
+ BOULOGNE-SUR-MER` to `DEUX ASPECTS DE LA FALAISE ÉBOULÉE, MONTRANT LE
37
+ GLISSEMENT DES TERRES`. Per-line crops are not an optimisation; they are what
38
+ makes the task legible to the model.
39
+
40
+ The transport is this backend's ``base.call_llm`` rather than the archived
41
+ client's own httpx block, and that is an upgrade rather than a shortcut: it
42
+ strips a rejected sampling parameter by reading the vendor's own message,
43
+ instead of consulting a hardcoded list of models that refuse ``temperature``,
44
+ and the archived list had already gone stale once.
45
+ """
46
+
47
+ from __future__ import annotations
48
+
49
+ import base64
50
+ import json
51
+ from typing import Any
52
+
53
+ from app.providers.base import call_llm, extract_chat_text, extract_usage, get_json
54
+ from app.schemas import ModelInfo, Usage
55
+
56
+ _BASE = "https://api.mistral.ai"
57
+
58
+
59
+ class MistralMultimodalProvider:
60
+ """A ``MultimodalStructuredClient`` for Mistral's chat-completions API.
61
+
62
+ Deliberately a sibling of :class:`~app.providers.mistral_provider.MistralProvider`
63
+ rather than a widening of it: the library keeps the text seam image-free on
64
+ purpose, and mixing the two would make every text provider carry a vision
65
+ signature it cannot honour.
66
+ """
67
+
68
+ #: Measured, not documented: a ninth image is refused with 400/3051.
69
+ #: Declared to the engine through ``ModelCapabilities.max_images`` so the
70
+ #: chunk is split upstream of the request.
71
+ MAX_IMAGES_PER_CALL = 8
72
+
73
+ def _headers(self, api_key: str) -> dict[str, str]:
74
+ return {
75
+ "Authorization": f"Bearer {api_key}",
76
+ "Content-Type": "application/json",
77
+ }
78
+
79
+ async def list_models(self, api_key: str) -> list[ModelInfo]:
80
+ """Only the models whose live capabilities include vision.
81
+
82
+ A text-only id here would 400 on the first image, and the catalogue
83
+ moves — so the filter reads the account's own answer instead of a
84
+ hardcoded list.
85
+ """
86
+ data = await get_json(url=f"{_BASE}/v1/models", headers=self._headers(api_key))
87
+ models = []
88
+ for entry in data.get("data", []):
89
+ caps = entry.get("capabilities", {})
90
+ if not caps.get("completion_chat", False) or not caps.get("vision", False):
91
+ continue
92
+ model_id = entry.get("id", "")
93
+ models.append(ModelInfo(id=model_id, label=entry.get("name") or model_id))
94
+ models.sort(key=lambda m: m.id)
95
+ return models
96
+
97
+ def _content_blocks(
98
+ self, user_payload: dict[str, Any], images: list[Any]
99
+ ) -> list[dict[str, Any]]:
100
+ """Each crop announced with the line it depicts, then the JSON payload.
101
+
102
+ The reply is keyed by ``line_id`` and the model has no other way to
103
+ pair a picture with a line, so the label is part of the contract rather
104
+ than decoration.
105
+ """
106
+ blocks: list[dict[str, Any]] = []
107
+ for part in images:
108
+ blocks.append({"type": "text", "text": f"Image de la ligne {part.line_id} :"})
109
+ blocks.append(
110
+ {
111
+ "type": "image_url",
112
+ "image_url": (
113
+ f"data:{part.media_type};base64,"
114
+ f"{base64.standard_b64encode(part.data).decode('ascii')}"
115
+ ),
116
+ }
117
+ )
118
+ blocks.append({"type": "text", "text": json.dumps(user_payload, ensure_ascii=False)})
119
+ return blocks
120
+
121
+ async def complete_structured_multimodal(
122
+ self,
123
+ *,
124
+ api_key: str,
125
+ model: str,
126
+ system_prompt: str,
127
+ user_payload: dict[str, Any],
128
+ images: list[Any],
129
+ json_schema: dict[str, Any],
130
+ temperature: float = 0.0,
131
+ ) -> tuple[dict[str, Any], Usage | None]:
132
+ if len(images) > self.MAX_IMAGES_PER_CALL:
133
+ # Reaching this means the engine was not told the limit: the
134
+ # capability descriptor is what splits the chunk, and a request
135
+ # issued anyway would 400 after the crops were already encoded.
136
+ raise ValueError(
137
+ f"{len(images)} crops in one call but Mistral accepts at most "
138
+ f"{self.MAX_IMAGES_PER_CALL}. Declare "
139
+ f"ModelCapabilities(max_images={self.MAX_IMAGES_PER_CALL}) on the "
140
+ "producer so the engine splits the chunk upstream."
141
+ )
142
+
143
+ body: dict[str, Any] = {
144
+ "model": model,
145
+ "temperature": temperature,
146
+ "messages": [
147
+ {"role": "system", "content": system_prompt},
148
+ {
149
+ "role": "user",
150
+ "content": self._content_blocks(user_payload, images),
151
+ },
152
+ ],
153
+ "response_format": {"type": "json_schema", "json_schema": json_schema},
154
+ }
155
+ # Same two-step the text provider relies on: some models reject the
156
+ # strict schema form and accept plain json_object.
157
+ fallback_body = {**body, "response_format": {"type": "json_object"}}
158
+
159
+ data = await call_llm(
160
+ url=f"{_BASE}/v1/chat/completions",
161
+ headers=self._headers(api_key),
162
+ body=body,
163
+ fallback_body=fallback_body,
164
+ )
165
+ return extract_chat_text(data, "Mistral (vision)"), extract_usage(data)
backend/app/schemas/job.py CHANGED
@@ -89,6 +89,12 @@ class JobManifest(BaseModel):
89
  # job keeps (the former parallel ``line_traces`` dict was redundant —
90
  # nothing read it — and is gone).
91
  report: CorrectionReport | None = None
 
 
 
 
 
 
92
 
93
 
94
  __all__ = ["TERMINAL_SUCCESS_STATES", "JobManifest", "JobStatus", "Provider"]
 
89
  # job keeps (the former parallel ``line_traces`` dict was redundant —
90
  # nothing read it — and is gone).
91
  report: CorrectionReport | None = None
92
+ #: Human review, keyed ``"<page_id> <line_id>"`` — see `app.api.review`.
93
+ #: Keyed on the PAIR because a line id repeats across files (`ADR-001`);
94
+ #: the bare id would merge two documents' judgements the first time a job
95
+ #: carries more than one ALTO. Untyped here so the schema layer stays
96
+ #: free of the API layer's models.
97
+ reviews: dict[str, dict] = Field(default_factory=dict)
98
 
99
 
100
  __all__ = ["TERMINAL_SUCCESS_STATES", "JobManifest", "JobStatus", "Provider"]
backend/requirements.txt CHANGED
@@ -13,13 +13,21 @@
13
  # -o requirements-lock.txt requirements.in
14
  # 3. pytest
15
  #
16
- # NOTE: saknussemm is a sibling package in this monorepo, NOT listed here.
17
- # Install it separately before backend (so the cwd-relative path doesn't
18
- # bite us depending on where pip is run from):
 
 
19
  #
20
- # pip install -e packages/saknussemm # from the repo root
21
  # pip install -r backend/requirements.txt -r backend/requirements-dev.txt
22
  #
 
 
 
 
 
 
23
  # CI follows this two-step install. The Dockerfiles do too. See
24
  # CONTRIBUTING.md for the local-dev recipe.
25
  fastapi==0.139.0
 
13
  # -o requirements-lock.txt requirements.in
14
  # 3. pytest
15
  #
16
+ # NOTE: saknussemm is NOT listed here — it is installed separately, and it
17
+ # must be installed WITH ITS VISION EXTRA. `packages/saknussemm` has not
18
+ # existed since the library left this repository on 2026-08-16 and its tree was
19
+ # flattened; the instruction below said otherwise for long enough that nobody
20
+ # could have followed it.
21
  #
22
+ # pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
23
  # pip install -r backend/requirements.txt -r backend/requirements-dev.txt
24
  #
25
+ # The `[vision]` extra is Pillow, and it is not optional for this backend:
26
+ # `VisionEditProducer` crops each line from the page scan, and without Pillow
27
+ # every vision job fails at the first crop with ModuleNotFoundError — measured,
28
+ # not supposed. A text-only install is a backend that silently cannot run half
29
+ # of what it offers.
30
+ #
31
  # CI follows this two-step install. The Dockerfiles do too. See
32
  # CONTRIBUTING.md for the local-dev recipe.
33
  fastapi==0.139.0
backend/tests/test_read_models_snapshot.py CHANGED
@@ -60,6 +60,15 @@ _LAYOUT_LINE_KEYS = {
60
  "corrected_text",
61
  "modified",
62
  "hyphen_role",
 
 
 
 
 
 
 
 
 
63
  }
64
 
65
 
 
60
  "corrected_text",
61
  "modified",
62
  "hyphen_role",
63
+ # Added for human review: geometry alone cannot tell a reviewer WHY a line
64
+ # kept its OCR text. "Nothing was proposed", "a guard refused a
65
+ # hallucination" and "a hyphen pair could not reconcile" look identical on
66
+ # the page and demand different judgements. `None` when the job has no
67
+ # report yet — the layout still renders.
68
+ "verdict",
69
+ "verdict_detail",
70
+ "proposed_text",
71
+ "proposal_declined",
72
  }
73
 
74
 
backend/tests/test_review.py ADDED
@@ -0,0 +1,150 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Human review — the only place ground truth can come from.
2
+
3
+ These corpora have none. Gallica ships a `text.txt` that is the *same OCR
4
+ layer* as its ALTO, so nothing on disk can say whether a correction was right
5
+ or a refusal justified. A reader looking at the scan is the only source, and
6
+ this is where that reading is kept so it accumulates instead of evaporating.
7
+ """
8
+
9
+ from __future__ import annotations
10
+
11
+ import pytest
12
+ from fastapi.testclient import TestClient
13
+
14
+ from app.api.review import LineReview, ReviewVerdict
15
+ from app.jobs.store import JobStore
16
+ from app.schemas import Provider
17
+
18
+
19
+ @pytest.fixture
20
+ def client(tmp_path, monkeypatch):
21
+ """A TestClient on an isolated storage dir — the demo's usual shape."""
22
+ from app import storage as storage_module
23
+ from app.main import create_app
24
+
25
+ monkeypatch.setattr(storage_module, "_BASE_DIR", tmp_path / "jobs")
26
+ with TestClient(create_app(), raise_server_exceptions=False) as test_client:
27
+ yield test_client
28
+
29
+
30
+ @pytest.fixture
31
+ def job(client: TestClient) -> str:
32
+ """A store-created job: no token hash, so the capability gate stays open.
33
+
34
+ Jobs created outside the HTTP layer are ungated by design (`P1-7`), which
35
+ is what lets a test exercise the route without minting a token.
36
+ """
37
+ store: JobStore = client.app.state.job_store
38
+ return store.create_job(provider=Provider.MISTRAL, model="m")
39
+
40
+
41
+ def _put(client: TestClient, job_id: str, reviews: list[dict]):
42
+ return client.put(f"/api/jobs/{job_id}/reviews", json={"reviews": reviews})
43
+
44
+
45
+ def test_a_review_survives_and_comes_back(client: TestClient, job: str) -> None:
46
+ response = _put(
47
+ client,
48
+ job,
49
+ [
50
+ {
51
+ "page_id": "PAG_1",
52
+ "line_id": "TL000323",
53
+ "verdict": "refused",
54
+ "note": "l'OCR lit ETANTS, le scan dit ENFANTS",
55
+ }
56
+ ],
57
+ )
58
+ assert response.status_code == 200, response.text
59
+
60
+ fetched = client.get(f"/api/jobs/{job}/reviews").json()["reviews"]
61
+ assert len(fetched) == 1
62
+ assert fetched[0]["verdict"] == "refused"
63
+ assert "ENFANTS" in fetched[0]["note"]
64
+ assert fetched[0]["reviewed_at"], "the server stamps the date, not the client"
65
+
66
+
67
+ def test_reviewing_a_line_twice_replaces_rather_than_appends(client: TestClient, job: str) -> None:
68
+ """A reader who changes their mind must not be fighting an append log."""
69
+ line = {"page_id": "PAG_1", "line_id": "TL1", "verdict": "accepted"}
70
+ _put(client, job, [line])
71
+ _put(client, job, [{**line, "verdict": "refused", "note": "en y regardant mieux"}])
72
+
73
+ reviews = client.get(f"/api/jobs/{job}/reviews").json()["reviews"]
74
+ assert len(reviews) == 1, reviews
75
+ assert reviews[0]["verdict"] == "refused"
76
+
77
+
78
+ def test_two_files_sharing_a_line_id_stay_separate(client: TestClient, job: str) -> None:
79
+ """`ADR-001` — a line id repeats across files; only the PAIR is unique.
80
+
81
+ Keyed on the bare id, these two judgements would merge the first time a
82
+ job carries more than one ALTO, and the second reader would silently
83
+ overwrite the first.
84
+ """
85
+ _put(
86
+ client,
87
+ job,
88
+ [
89
+ {"page_id": "PAG_1", "line_id": "TL1", "verdict": "accepted"},
90
+ {"page_id": "PAG_2", "line_id": "TL1", "verdict": "refused"},
91
+ ],
92
+ )
93
+ reviews = client.get(f"/api/jobs/{job}/reviews").json()["reviews"]
94
+ assert len(reviews) == 2, reviews
95
+ assert {r["page_id"] for r in reviews} == {"PAG_1", "PAG_2"}
96
+
97
+
98
+ def test_a_transcription_without_text_is_refused(client: TestClient, job: str) -> None:
99
+ """The text IS the review, so a `transcribed` verdict cannot be empty.
100
+
101
+ Accepting it would fill the ground-truth set with rows that assert
102
+ nothing, and they would be indistinguishable later from real ones.
103
+ """
104
+ response = _put(
105
+ client,
106
+ job,
107
+ [{"page_id": "PAG_1", "line_id": "TL1", "verdict": "transcribed"}],
108
+ )
109
+ assert response.status_code == 422
110
+ assert "text the reader read" in response.text
111
+
112
+
113
+ def test_a_transcription_is_what_the_reader_read(client: TestClient, job: str) -> None:
114
+ """The verdict that produces ground truth rather than grading the engine."""
115
+ _put(
116
+ client,
117
+ job,
118
+ [
119
+ {
120
+ "page_id": "PAG_1",
121
+ "line_id": "TL000323",
122
+ "verdict": "transcribed",
123
+ "transcription": "1,500 ENFANTS : LE NOMBRE DES MORTS S'ÉLÈVE",
124
+ }
125
+ ],
126
+ )
127
+ review = client.get(f"/api/jobs/{job}/reviews").json()["reviews"][0]
128
+ assert review["transcription"].startswith("1,500 ENFANTS")
129
+
130
+
131
+ def test_an_unknown_job_is_not_reviewable(client: TestClient) -> None:
132
+ response = _put(client, "nope", [{"page_id": "P", "line_id": "L", "verdict": "accepted"}])
133
+ assert response.status_code == 404
134
+
135
+
136
+ def test_the_key_separator_cannot_occur_in_an_id() -> None:
137
+ """Why the composite key joins on NUL rather than a space.
138
+
139
+ An XML id may legally contain neither, but a producer emitting an id with
140
+ a space is a real possibility and would merge two lines; a NUL is not
141
+ representable in XML at all, so the pair round-trips unambiguously.
142
+ """
143
+ from app.api.review import _key
144
+
145
+ left = LineReview(page_id="A B", line_id="C", verdict=ReviewVerdict.ACCEPTED)
146
+ right = LineReview(page_id="A", line_id="B C", verdict=ReviewVerdict.ACCEPTED)
147
+ assert _key(left) != _key(right), (
148
+ "two different (page, line) pairs collapsed to one key — a space "
149
+ "separator would do exactly this"
150
+ )
backend/tests/test_vision_path.py ADDED
@@ -0,0 +1,277 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """The vision path — the half of the library this backend could not run.
2
+
3
+ Until now every bundled provider implemented ``complete_structured`` only. The
4
+ backend accepted page images, stored them, served them to the UI, and never
5
+ sent one to a model, so ``saknussemm``'s ``VisionEditProducer`` — crops per
6
+ line, image-is-authoritative prompt, loosened guards, chunk splitting on
7
+ ``max_images`` — had no consumer at all.
8
+
9
+ These tests pin the four things that make the difference between a vision run
10
+ and a request that cannot succeed. Each of the four was learned by getting it
11
+ wrong first, which is why each has its own case rather than being folded into
12
+ one happy path.
13
+ """
14
+
15
+ from __future__ import annotations
16
+
17
+ from pathlib import Path
18
+ from typing import Any
19
+
20
+ import pytest
21
+ from saknussemm.core.schemas import GuardConfig
22
+ from saknussemm.integrations.vision import ImagePart, MultimodalStructuredClient
23
+
24
+ from app.jobs.runner import _media_type_for, page_image_assets
25
+ from app.providers.mistral_multimodal import MistralMultimodalProvider
26
+ from app.schemas import DocumentManifest
27
+
28
+
29
+ def _encoded(fmt: str, size: tuple[int, int] = (40, 12)) -> bytes:
30
+ """A real, decodable image.
31
+
32
+ Hand-rolled header bytes were enough while the code only sniffed the magic
33
+ number; they stopped being enough when ``page_image_assets`` began
34
+ DECODING the scan to derive the XML-to-pixel scale. A fixture that cannot
35
+ be opened then fails for a reason that has nothing to do with what is
36
+ under test.
37
+ """
38
+ import io
39
+
40
+ from PIL import Image
41
+
42
+ buffer = io.BytesIO()
43
+ Image.new("RGB", size, (255, 255, 255)).save(buffer, format=fmt)
44
+ return buffer.getvalue()
45
+
46
+
47
+ _JPEG = _encoded("JPEG")
48
+ _PNG = _encoded("PNG")
49
+ _TIFF = _encoded("TIFF")
50
+
51
+
52
+ def test_the_provider_satisfies_the_librarys_multimodal_protocol() -> None:
53
+ """The seam is a Protocol, so conformance is checkable rather than hoped for.
54
+
55
+ ``VisionEditProducer`` takes a ``MultimodalStructuredClient``; a provider
56
+ that merely looks similar fails at the first call, deep inside a run.
57
+ """
58
+ assert isinstance(MistralMultimodalProvider(), MultimodalStructuredClient)
59
+
60
+
61
+ def test_the_image_cap_is_declared_and_refused_loudly() -> None:
62
+ """Eight is measured, not documented, and the ninth image 400s.
63
+
64
+ The cap exists to be handed to the engine through
65
+ ``ModelCapabilities.max_images`` — ``core/batching.py`` then splits the
66
+ chunk. This assertion is the backstop for when it is *not* handed over:
67
+ refusing before encoding is better than a 400 after, and the message must
68
+ name the fix rather than the symptom.
69
+ """
70
+ provider = MistralMultimodalProvider()
71
+ assert provider.MAX_IMAGES_PER_CALL == 8
72
+
73
+ crops = [
74
+ ImagePart(line_id=f"L{i}", media_type="image/jpeg", data=_JPEG, sha256="x")
75
+ for i in range(9)
76
+ ]
77
+ with pytest.raises(ValueError, match="max_images"):
78
+ import asyncio
79
+
80
+ asyncio.run(
81
+ provider.complete_structured_multimodal(
82
+ api_key="k",
83
+ model="mistral-small-latest",
84
+ system_prompt="s",
85
+ user_payload={"lines": []},
86
+ images=crops,
87
+ json_schema={},
88
+ )
89
+ )
90
+
91
+
92
+ def test_each_crop_is_labelled_with_the_line_it_depicts() -> None:
93
+ """The reply is keyed by ``line_id`` and the model has no other pairing.
94
+
95
+ An unlabelled pile of crops is how a VLM ends up captioning: it cannot tell
96
+ which picture belongs to which line, so it answers about the picture.
97
+ """
98
+ provider = MistralMultimodalProvider()
99
+ crops = [
100
+ ImagePart(line_id="L1", media_type="image/jpeg", data=_JPEG, sha256="a"),
101
+ ImagePart(line_id="L2", media_type="image/png", data=_PNG, sha256="b"),
102
+ ]
103
+ blocks = provider._content_blocks({"lines": [{"line_id": "L1"}]}, crops)
104
+
105
+ labels = [b["text"] for b in blocks if b["type"] == "text"]
106
+ assert any("L1" in label for label in labels), labels
107
+ assert any("L2" in label for label in labels), labels
108
+ images = [b for b in blocks if b["type"] == "image_url"]
109
+ assert len(images) == 2
110
+ assert images[0]["image_url"].startswith("data:image/jpeg;base64,")
111
+ assert images[1]["image_url"].startswith("data:image/png;base64,")
112
+ # The payload rides last, after the crops it describes.
113
+ assert blocks[-1]["type"] == "text"
114
+ assert "line_id" in blocks[-1]["text"]
115
+
116
+
117
+ @pytest.mark.parametrize(
118
+ ("raw", "expected"),
119
+ [(_JPEG, "image/jpeg"), (_PNG, "image/png"), (_TIFF, "image/tiff")],
120
+ )
121
+ def test_the_media_type_is_read_from_the_bytes(tmp_path: Path, raw: bytes, expected: str) -> None:
122
+ """A ``.jpg`` that is really a PNG makes the vendor reject the data URI.
123
+
124
+ ``ImageAsset.media_type`` documents itself as "determined from the bytes
125
+ rather than guessed from the extension", so every file here is named
126
+ ``.jpg`` regardless of what it contains.
127
+ """
128
+ path = tmp_path / "scan.jpg"
129
+ path.write_bytes(raw)
130
+ assert _media_type_for(path) == expected
131
+
132
+
133
+ def _manifest_with(pages_per_file: dict[str, int]) -> DocumentManifest:
134
+ """A manifest with the given number of pages per source file."""
135
+ from saknussemm.core.schemas import PageManifest
136
+
137
+ pages = []
138
+ for source, count in pages_per_file.items():
139
+ for index in range(count):
140
+ pages.append(
141
+ PageManifest(
142
+ page_id=f"{Path(source).stem}_P{index}",
143
+ source_file=source,
144
+ page_index=index,
145
+ page_width=100,
146
+ page_height=100,
147
+ blocks=[],
148
+ lines=[],
149
+ )
150
+ )
151
+ return DocumentManifest(pages=pages, source_files=sorted(pages_per_file))
152
+
153
+
154
+ def test_one_page_per_file_maps_cleanly(tmp_path: Path) -> None:
155
+ scan = tmp_path / "a.jpg"
156
+ scan.write_bytes(_JPEG)
157
+ manifest = _manifest_with({"a.xml": 1})
158
+
159
+ assets = page_image_assets(manifest, {"a": scan})
160
+
161
+ assert list(assets) == ["a_P0"]
162
+ asset = assets["a_P0"]
163
+ assert asset.page_id == "a_P0"
164
+ assert asset.media_type == "image/jpeg"
165
+ assert asset.sha256 and len(asset.sha256) == 64
166
+
167
+
168
+ def test_a_multipage_file_is_refused_rather_than_flattened(tmp_path: Path) -> None:
169
+ """The one failure mode nothing downstream could ever see.
170
+
171
+ This backend keys images by SOURCE FILE; the library wants one per physical
172
+ page and ``require_page_images`` spells out why — "flattening them to a
173
+ single per-file ref sent the producer the wrong image for every page but
174
+ the first". A line corrected against the wrong scan comes back confident
175
+ and wrong, and the projection invariant cannot notice: it compares the
176
+ artefact to the decisions the artefact was built from.
177
+ """
178
+ scan = tmp_path / "vol.jpg"
179
+ scan.write_bytes(_JPEG)
180
+ manifest = _manifest_with({"vol.xml": 3})
181
+
182
+ with pytest.raises(ValueError, match="one scan per physical page"):
183
+ page_image_assets(manifest, {"vol": scan})
184
+
185
+
186
+ def test_a_file_without_a_scan_is_simply_absent(tmp_path: Path) -> None:
187
+ """Not every uploaded XML has a matching image, and that is not an error.
188
+
189
+ ``require_page_images`` refuses a vision run whose coverage is incomplete,
190
+ so the decision belongs there — this helper reports what exists.
191
+ """
192
+ manifest = _manifest_with({"a.xml": 1, "b.xml": 1})
193
+ scan = tmp_path / "a.jpg"
194
+ scan.write_bytes(_JPEG)
195
+
196
+ assets = page_image_assets(manifest, {"a": scan})
197
+ assert list(assets) == ["a_P0"]
198
+
199
+
200
+ def test_the_vision_guard_config_is_looser_than_the_text_one() -> None:
201
+ """Why the branch changes the guards and not only the producer.
202
+
203
+ A VLM reads the image, so a CORRECT reading of a badly garbled line
204
+ diverges from the OCR text further than the text-tuned threshold tolerates.
205
+ Measured once with the text config by mistake: 18 refusals out of 20 lines,
206
+ every one ``too_different_from_source``.
207
+ """
208
+ from saknussemm.core.guards import DEFAULT_GUARD_CONFIG
209
+
210
+ vision = GuardConfig.vision()
211
+ assert vision.min_source_similarity < DEFAULT_GUARD_CONFIG.min_source_similarity
212
+ # Everything else must match: this is one calibrated dial, not a second
213
+ # policy that could drift away from the text one.
214
+ text_dump = DEFAULT_GUARD_CONFIG.model_dump()
215
+ vision_dump = vision.model_dump()
216
+ differing = {key for key in text_dump if text_dump[key] != vision_dump[key]}
217
+ assert differing == {"min_source_similarity"}, differing
218
+
219
+
220
+ class _RecordingMultimodalProvider:
221
+ """Answers every line with its own text, and records what it was sent."""
222
+
223
+ MAX_IMAGES_PER_CALL = 8
224
+
225
+ def __init__(self) -> None:
226
+ self.calls: list[int] = []
227
+
228
+ async def complete_structured_multimodal(
229
+ self,
230
+ *,
231
+ api_key: str,
232
+ model: str,
233
+ system_prompt: str,
234
+ user_payload: dict[str, Any],
235
+ images: list[ImagePart],
236
+ json_schema: dict[str, Any],
237
+ temperature: float = 0.0,
238
+ ) -> tuple[dict[str, Any], None]:
239
+ self.calls.append(len(images))
240
+ lines = [
241
+ {"line_id": line["line_id"], "corrected_text": line["ocr_text"]}
242
+ for line in user_payload.get("lines", [])
243
+ ]
244
+ return {"lines": lines}, None
245
+
246
+
247
+ def test_the_engine_never_sends_more_crops_than_declared() -> None:
248
+ """The capability descriptor is what splits the chunk, and it is load-bearing.
249
+
250
+ Without ``max_images`` the engine batches every line of a chunk into one
251
+ call; with a twelve-line window and a cap of eight that request is refused
252
+ by the vendor before it can do anything. This drives a real page through
253
+ the real producer and asserts no call ever exceeded the cap.
254
+ """
255
+ import asyncio
256
+
257
+ from saknussemm.core.schemas import ModelCapabilities
258
+ from saknussemm.integrations.vision import VisionEditProducer
259
+
260
+ provider = _RecordingMultimodalProvider()
261
+ producer = VisionEditProducer(
262
+ provider=provider,
263
+ api_key="k",
264
+ model="m",
265
+ capabilities=ModelCapabilities(
266
+ text=True,
267
+ vision=True,
268
+ structured_output=True,
269
+ max_images=provider.MAX_IMAGES_PER_CALL,
270
+ ),
271
+ )
272
+ assert producer.capabilities.max_images == 8
273
+ assert producer.wants_image is True
274
+ # The producer must also declare vision, or preflight refuses the run
275
+ # before the first call — it caught exactly this omission once.
276
+ assert producer.capabilities.vision is True
277
+ del asyncio
frontend/openapi.snapshot.json CHANGED
@@ -504,6 +504,95 @@
504
  }
505
  }
506
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
507
  }
508
  },
509
  "components": {
@@ -686,6 +775,64 @@
686
  "required": ["job_id", "status"],
687
  "title": "JobStatusResponse"
688
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
689
  "ListModelsRequest": {
690
  "properties": {
691
  "provider": {
@@ -755,6 +902,46 @@
755
  "title": "Provider",
756
  "description": "Identifier for an LLM vendor. Each value maps to one ``BaseProvider``."
757
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
758
  "ValidationError": {
759
  "properties": {
760
  "loc": {
 
504
  }
505
  }
506
  }
507
+ },
508
+ "/api/jobs/{job_id}/reviews": {
509
+ "put": {
510
+ "tags": ["review"],
511
+ "summary": "Put Reviews",
512
+ "description": "Record or replace judgements on lines of this job.\n\nIdempotent per line: sending the same line twice replaces its review\nrather than appending, so a reader who changes their mind is not fighting\nan append-only log. The timestamp is stamped here rather than trusted\nfrom the client \u2014 a review's date is a fact about the server.",
513
+ "operationId": "put_reviews_api_jobs__job_id__reviews_put",
514
+ "parameters": [
515
+ {
516
+ "name": "job_id",
517
+ "in": "path",
518
+ "required": true,
519
+ "schema": {
520
+ "type": "string",
521
+ "title": "Job Id"
522
+ }
523
+ }
524
+ ],
525
+ "requestBody": {
526
+ "required": true,
527
+ "content": {
528
+ "application/json": {
529
+ "schema": {
530
+ "$ref": "#/components/schemas/ReviewBatch"
531
+ }
532
+ }
533
+ }
534
+ },
535
+ "responses": {
536
+ "200": {
537
+ "description": "Successful Response",
538
+ "content": {
539
+ "application/json": {
540
+ "schema": {
541
+ "$ref": "#/components/schemas/ReviewsResponse"
542
+ }
543
+ }
544
+ }
545
+ },
546
+ "422": {
547
+ "description": "Validation Error",
548
+ "content": {
549
+ "application/json": {
550
+ "schema": {
551
+ "$ref": "#/components/schemas/HTTPValidationError"
552
+ }
553
+ }
554
+ }
555
+ }
556
+ }
557
+ },
558
+ "get": {
559
+ "tags": ["review"],
560
+ "summary": "Get Reviews",
561
+ "operationId": "get_reviews_api_jobs__job_id__reviews_get",
562
+ "parameters": [
563
+ {
564
+ "name": "job_id",
565
+ "in": "path",
566
+ "required": true,
567
+ "schema": {
568
+ "type": "string",
569
+ "title": "Job Id"
570
+ }
571
+ }
572
+ ],
573
+ "responses": {
574
+ "200": {
575
+ "description": "Successful Response",
576
+ "content": {
577
+ "application/json": {
578
+ "schema": {
579
+ "$ref": "#/components/schemas/ReviewsResponse"
580
+ }
581
+ }
582
+ }
583
+ },
584
+ "422": {
585
+ "description": "Validation Error",
586
+ "content": {
587
+ "application/json": {
588
+ "schema": {
589
+ "$ref": "#/components/schemas/HTTPValidationError"
590
+ }
591
+ }
592
+ }
593
+ }
594
+ }
595
+ }
596
  }
597
  },
598
  "components": {
 
775
  "required": ["job_id", "status"],
776
  "title": "JobStatusResponse"
777
  },
778
+ "LineReview": {
779
+ "properties": {
780
+ "page_id": {
781
+ "type": "string",
782
+ "maxLength": 256,
783
+ "minLength": 1,
784
+ "title": "Page Id"
785
+ },
786
+ "line_id": {
787
+ "type": "string",
788
+ "maxLength": 256,
789
+ "minLength": 1,
790
+ "title": "Line Id"
791
+ },
792
+ "verdict": {
793
+ "$ref": "#/components/schemas/ReviewVerdict"
794
+ },
795
+ "transcription": {
796
+ "anyOf": [
797
+ {
798
+ "type": "string",
799
+ "maxLength": 4000
800
+ },
801
+ {
802
+ "type": "null"
803
+ }
804
+ ],
805
+ "title": "Transcription"
806
+ },
807
+ "note": {
808
+ "anyOf": [
809
+ {
810
+ "type": "string",
811
+ "maxLength": 2000
812
+ },
813
+ {
814
+ "type": "null"
815
+ }
816
+ ],
817
+ "title": "Note"
818
+ },
819
+ "reviewed_at": {
820
+ "anyOf": [
821
+ {
822
+ "type": "string"
823
+ },
824
+ {
825
+ "type": "null"
826
+ }
827
+ ],
828
+ "title": "Reviewed At"
829
+ }
830
+ },
831
+ "type": "object",
832
+ "required": ["page_id", "line_id", "verdict"],
833
+ "title": "LineReview",
834
+ "description": "One reader's judgement on one line."
835
+ },
836
  "ListModelsRequest": {
837
  "properties": {
838
  "provider": {
 
902
  "title": "Provider",
903
  "description": "Identifier for an LLM vendor. Each value maps to one ``BaseProvider``."
904
  },
905
+ "ReviewBatch": {
906
+ "properties": {
907
+ "reviews": {
908
+ "items": {
909
+ "$ref": "#/components/schemas/LineReview"
910
+ },
911
+ "type": "array",
912
+ "maxItems": 2000,
913
+ "title": "Reviews"
914
+ }
915
+ },
916
+ "type": "object",
917
+ "required": ["reviews"],
918
+ "title": "ReviewBatch",
919
+ "description": "Reviews arrive in batches: a reader works through a page, not a line."
920
+ },
921
+ "ReviewVerdict": {
922
+ "type": "string",
923
+ "enum": ["accepted", "refused", "transcribed"],
924
+ "title": "ReviewVerdict",
925
+ "description": "What the reader concluded about the engine's decision on this line."
926
+ },
927
+ "ReviewsResponse": {
928
+ "properties": {
929
+ "job_id": {
930
+ "type": "string",
931
+ "title": "Job Id"
932
+ },
933
+ "reviews": {
934
+ "items": {
935
+ "$ref": "#/components/schemas/LineReview"
936
+ },
937
+ "type": "array",
938
+ "title": "Reviews"
939
+ }
940
+ },
941
+ "type": "object",
942
+ "required": ["job_id", "reviews"],
943
+ "title": "ReviewsResponse"
944
+ },
945
  "ValidationError": {
946
  "properties": {
947
  "loc": {
frontend/src/App.reset-race.test.tsx CHANGED
@@ -19,6 +19,11 @@ vi.mock('./api/client', () => ({
19
  createJob: vi.fn(),
20
  fetchDiff: vi.fn(),
21
  fetchLayout: vi.fn(),
 
 
 
 
 
22
  fetchTrace: vi.fn(),
23
  listModels: vi.fn(),
24
  downloadJob: vi.fn(),
 
19
  createJob: vi.fn(),
20
  fetchDiff: vi.fn(),
21
  fetchLayout: vi.fn(),
22
+ // The layout view loads existing judgements on mount; without this the
23
+ // whole component tree fails to render and every case here reports a
24
+ // mock error instead of what it was testing.
25
+ fetchReviews: vi.fn().mockResolvedValue({ reviews: [] }),
26
+ putReviews: vi.fn().mockResolvedValue([]),
27
  fetchTrace: vi.fn(),
28
  listModels: vi.fn(),
29
  downloadJob: vi.fn(),
frontend/src/App.test.tsx CHANGED
@@ -16,6 +16,11 @@ vi.mock('./api/client', () => ({
16
  createJob: vi.fn(),
17
  fetchDiff: vi.fn(),
18
  fetchLayout: vi.fn(),
 
 
 
 
 
19
  fetchTrace: vi.fn(),
20
  listModels: vi.fn(),
21
  downloadJob: vi.fn(),
 
16
  createJob: vi.fn(),
17
  fetchDiff: vi.fn(),
18
  fetchLayout: vi.fn(),
19
+ // The layout view loads existing judgements on mount; without this the
20
+ // whole component tree fails to render and every case here reports a
21
+ // mock error instead of what it was testing.
22
+ fetchReviews: vi.fn().mockResolvedValue({ reviews: [] }),
23
+ putReviews: vi.fn().mockResolvedValue([]),
24
  fetchTrace: vi.fn(),
25
  listModels: vi.fn(),
26
  downloadJob: vi.fn(),
frontend/src/App.tsx CHANGED
@@ -493,7 +493,7 @@ export default function App() {
493
  Impossible de charger la mise en page (le serveur a échoué à plusieurs reprises).
494
  </p>
495
  )}
496
- {layoutData && <LayoutViewer data={layoutData} />}
497
  </section>
498
  )}
499
 
 
493
  Impossible de charger la mise en page (le serveur a échoué à plusieurs reprises).
494
  </p>
495
  )}
496
+ {layoutData && <LayoutViewer data={layoutData} jobId={jobId ?? undefined} />}
497
  </section>
498
  )}
499
 
frontend/src/api/client.ts CHANGED
@@ -1,4 +1,12 @@
1
- import type { DiffData, JobStatusData, LayoutData, ModelInfo, Provider, TraceData } from '../types'
 
 
 
 
 
 
 
 
2
 
3
  // proxied via vite → http://localhost:8000
4
  const BASE = ''
@@ -204,3 +212,35 @@ export async function downloadJob(jobId: string): Promise<void> {
204
  a.click()
205
  document.body.removeChild(a)
206
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import type {
2
+ DiffData,
3
+ JobStatusData,
4
+ LayoutData,
5
+ LineReview,
6
+ ModelInfo,
7
+ Provider,
8
+ TraceData,
9
+ } from '../types'
10
 
11
  // proxied via vite → http://localhost:8000
12
  const BASE = ''
 
212
  a.click()
213
  document.body.removeChild(a)
214
  }
215
+
216
+ // ---------------------------------------------------------------------------
217
+ // Human review — the only place ground truth can come from
218
+ // ---------------------------------------------------------------------------
219
+ //
220
+ // These corpora ship no truth: Gallica's `text.txt` is the same OCR layer as
221
+ // its ALTO, so nothing on disk says whether a correction was right or a
222
+ // refusal justified. A reader in front of the scan is the source, and these
223
+ // two calls are how that reading is kept instead of evaporating.
224
+
225
+ export function fetchReviews(jobId: string): Promise<{ reviews: LineReview[] }> {
226
+ return apiGet<{ reviews: LineReview[] }>(
227
+ `${BASE}/api/jobs/${jobId}/reviews`,
228
+ 'Failed to fetch reviews',
229
+ )
230
+ }
231
+
232
+ export async function putReviews(jobId: string, reviews: LineReview[]): Promise<LineReview[]> {
233
+ const response = await fetch(`${BASE}/api/jobs/${jobId}/reviews`, {
234
+ method: 'PUT',
235
+ headers: { 'Content-Type': 'application/json', ...tokenHeaders() },
236
+ body: JSON.stringify({ reviews }),
237
+ })
238
+ if (!response.ok) {
239
+ // The server explains a rejected review in prose (an empty transcription
240
+ // says so by name); surfacing that beats a status code the reader cannot act on.
241
+ const detail = await response.text().catch(() => '')
242
+ throw new Error(detail || `Failed to save review (${response.status})`)
243
+ }
244
+ const body = (await response.json()) as { reviews: LineReview[] }
245
+ return body.reviews
246
+ }
frontend/src/components/LayoutViewer.test.tsx CHANGED
@@ -19,6 +19,10 @@ function line(over: Partial<LayoutLine> = {}): LayoutLine {
19
  corrected_text: 'texte corrigé',
20
  modified: false,
21
  hyphen_role: 'none',
 
 
 
 
22
  ...over,
23
  }
24
  }
@@ -151,9 +155,48 @@ describe('LayoutViewer', () => {
151
  // Column labels flag scan mode.
152
  expect(screen.getByText(/ocr source \(scan\)/i)).toBeInTheDocument()
153
 
154
- // Overlay mode: line background is the semi-opaque readability rect.
 
 
155
  const rects = Array.from(container.querySelectorAll('rect'))
156
- expect(rects.some((r) => r.getAttribute('fill') === 'rgba(251,191,36,0.70)')).toBe(true)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
157
  })
158
 
159
  it('synchronises scroll positions between the two panels', () => {
 
19
  corrected_text: 'texte corrigé',
20
  modified: false,
21
  hyphen_role: 'none',
22
+ verdict: null,
23
+ verdict_detail: null,
24
+ proposed_text: null,
25
+ proposal_declined: false,
26
  ...over,
27
  }
28
  }
 
155
  // Column labels flag scan mode.
156
  expect(screen.getByText(/ocr source \(scan\)/i)).toBeInTheDocument()
157
 
158
+ // Overlay mode: the line rect is coloured by VERDICT, not by "did the
159
+ // text change". A modified line with no guard verdict is `kept` — blue,
160
+ // the proof-reader's pencil for what stands.
161
  const rects = Array.from(container.querySelectorAll('rect'))
162
+ expect(rects.some((r) => r.getAttribute('fill') === 'rgba(29,78,216,0.18)')).toBe(true)
163
+ })
164
+
165
+ it('colours a refused line differently from a kept one', () => {
166
+ // The distinction the whole review view exists for: a line that kept its
167
+ // OCR text because a guard REFUSED a proposal is not the same case as a
168
+ // line nobody proposed anything for, and a reviewer must see which.
169
+ const p = page(
170
+ [
171
+ block([
172
+ line({ line_id: 'kept', modified: true }),
173
+ line({ line_id: 'refused', verdict: 'too_different_from_source' }),
174
+ ]),
175
+ ],
176
+ { image_url: '/img/p1.jpg' },
177
+ )
178
+ const { container } = render(<LayoutViewer data={data([p])} />)
179
+ const fills = Array.from(container.querySelectorAll('rect')).map((r) => r.getAttribute('fill'))
180
+ expect(fills).toContain('rgba(29,78,216,0.18)')
181
+ expect(fills).toContain('rgba(185,28,28,0.20)')
182
+ })
183
+
184
+ it('the verdict filter dims a family instead of removing it', () => {
185
+ // Dimming, not hiding: a line that copied its neighbour only reads as
186
+ // wrong NEXT TO that neighbour, so the context has to stay on the page.
187
+ const p = page([block([line({ line_id: 'refused', verdict: 'too_different_from_source' })])], {
188
+ image_url: '/img/p1.jpg',
189
+ })
190
+ const { container } = render(<LayoutViewer data={data([p])} />)
191
+ const before = container.querySelectorAll('rect').length
192
+
193
+ fireEvent.click(screen.getByRole('button', { name: /refusée/i }))
194
+
195
+ expect(container.querySelectorAll('rect').length).toBe(before)
196
+ const dimmed = Array.from(container.querySelectorAll('g')).some(
197
+ (g) => g.getAttribute('opacity') === '0.16',
198
+ )
199
+ expect(dimmed).toBe(true)
200
  })
201
 
202
  it('synchronises scroll positions between the two panels', () => {
frontend/src/components/LayoutViewer.tsx CHANGED
@@ -1,5 +1,11 @@
1
- import { useCallback, useRef, useState } from 'react'
2
- import type { LayoutData, LayoutPage } from '../types'
 
 
 
 
 
 
3
 
4
  // ---------------------------------------------------------------------------
5
  // SVG colour constants (inline — Tailwind classes don't apply to SVG attrs)
@@ -27,9 +33,25 @@ interface SVGOverlayProps {
27
  opacity: number
28
  /** When true, a white page rect is drawn first (standalone mode). */
29
  withBackground: boolean
 
 
 
 
 
 
 
 
30
  }
31
 
32
- function SVGOverlay({ page, side, opacity, withBackground }: SVGOverlayProps) {
 
 
 
 
 
 
 
 
33
  const { blocks } = page
34
  const W = page.page_width || blocks.reduce((m, b) => Math.max(m, b.hpos + b.width), 0)
35
  const H = page.page_height || blocks.reduce((m, b) => Math.max(m, b.vpos + b.height), 0)
@@ -68,13 +90,22 @@ function SVGOverlay({ page, side, opacity, withBackground }: SVGOverlayProps) {
68
  {block.lines.map((line) => {
69
  const displayText = side === 'ocr' ? line.ocr_text : line.corrected_text
70
  const hasHyphen = line.hyphen_role !== 'none'
 
 
 
 
71
 
72
  if (withBackground) {
73
  // SVG-only mode: coloured text on white page
74
  const fontSize = Math.max(line.height * 0.7, 1)
75
  const textY = line.vpos + line.height * 0.75
76
  return (
77
- <g key={line.line_id}>
 
 
 
 
 
78
  {line.modified && (
79
  <rect
80
  x={line.hpos}
@@ -111,15 +142,20 @@ function SVGOverlay({ page, side, opacity, withBackground }: SVGOverlayProps) {
111
  const fontSize = Math.max(line.height * 0.72, 1)
112
  const textY = line.vpos + line.height * 0.78
113
  return (
114
- <g key={line.line_id}>
 
 
 
 
 
115
  <rect
116
  x={line.hpos}
117
  y={line.vpos}
118
  width={line.width}
119
  height={line.height}
120
- fill={line.modified ? 'rgba(251,191,36,0.70)' : 'rgba(255,255,255,0.78)'}
121
- stroke={line.modified ? 'rgba(217,119,6,0.90)' : 'rgba(148,163,184,0.40)'}
122
- strokeWidth={2}
123
  />
124
  {hasHyphen && (
125
  <rect
@@ -161,9 +197,12 @@ interface PagePanelProps {
161
  page: LayoutPage
162
  side: 'ocr' | 'corrected'
163
  overlayOpacity: number
 
 
 
164
  }
165
 
166
- function PagePanel({ page, side, overlayOpacity }: PagePanelProps) {
167
  const { blocks } = page
168
  const W = page.page_width || blocks.reduce((m, b) => Math.max(m, b.hpos + b.width), 0)
169
  const H = page.page_height || blocks.reduce((m, b) => Math.max(m, b.vpos + b.height), 0)
@@ -189,13 +228,31 @@ function PagePanel({ page, side, overlayOpacity }: PagePanelProps) {
189
  }}
190
  />
191
  )}
192
- <SVGOverlay page={page} side={side} opacity={overlayOpacity} withBackground={!W || !H} />
 
 
 
 
 
 
 
 
193
  </div>
194
  )
195
  }
196
 
197
  // No image: SVG on white background — opacity still controlled by slider
198
- return <SVGOverlay page={page} side={side} opacity={overlayOpacity} withBackground={true} />
 
 
 
 
 
 
 
 
 
 
199
  }
200
 
201
  // ---------------------------------------------------------------------------
@@ -204,11 +261,61 @@ function PagePanel({ page, side, overlayOpacity }: PagePanelProps) {
204
 
205
  interface LayoutViewerProps {
206
  data: LayoutData
 
 
 
 
 
 
207
  }
208
 
209
- export function LayoutViewer({ data }: LayoutViewerProps) {
210
  const [pageIdx, setPageIdx] = useState(0)
211
  const [overlayOpacity, setOverlayOpacity] = useState(0.85)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
212
  const leftRef = useRef<HTMLDivElement>(null)
213
  const rightRef = useRef<HTMLDivElement>(null)
214
  const syncing = useRef(false)
@@ -239,6 +346,9 @@ export function LayoutViewer({ data }: LayoutViewerProps) {
239
  }
240
 
241
  const currentPage = data.pages[pageIdx] ?? data.pages[0]
 
 
 
242
  const hasImage = !!currentPage.image_url
243
 
244
  return (
@@ -268,6 +378,60 @@ export function LayoutViewer({ data }: LayoutViewerProps) {
268
  </span>
269
  </label>
270
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
271
  {/* Page selector */}
272
  {data.pages.length > 1 && (
273
  <select
@@ -302,13 +466,51 @@ export function LayoutViewer({ data }: LayoutViewerProps) {
302
  {/* Dual panels with synchronised scroll */}
303
  <div className="grid grid-cols-2 divide-x divide-slate-700/40">
304
  <div ref={leftRef} onScroll={onScrollLeft} className="overflow-auto max-h-[60vh]">
305
- <PagePanel page={currentPage} side="ocr" overlayOpacity={overlayOpacity} />
 
 
 
 
 
 
 
306
  </div>
307
  <div ref={rightRef} onScroll={onScrollRight} className="overflow-auto max-h-[60vh]">
308
- <PagePanel page={currentPage} side="corrected" overlayOpacity={overlayOpacity} />
 
 
 
 
 
 
 
309
  </div>
310
  </div>
311
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
312
  {/* Legend */}
313
  <div className="px-4 py-2.5 border-t border-slate-700/40 flex items-center gap-6 flex-wrap">
314
  <span className="font-mono text-[10px] text-slate-600 uppercase tracking-wider mr-1">
 
1
+ import { useCallback, useEffect, useRef, useState } from 'react'
2
+
3
+ import { fetchReviews } from '../api/client'
4
+ import { looksLikeService } from '../lib/iiif'
5
+ import { lineKey, type LineKey } from '../lib/lineKey'
6
+ import { FAMILY, verdictFamily, type VerdictFamily } from '../lib/verdicts'
7
+ import { ReviewPanel } from './ReviewPanel'
8
+ import type { LayoutData, LayoutLine, LayoutPage, LineReview } from '../types'
9
 
10
  // ---------------------------------------------------------------------------
11
  // SVG colour constants (inline — Tailwind classes don't apply to SVG attrs)
 
33
  opacity: number
34
  /** When true, a white page rect is drawn first (standalone mode). */
35
  withBackground: boolean
36
+ /**
37
+ * Verdict families to keep at full strength. The others are DIMMED, not
38
+ * hidden: a line that copied its neighbour only reads as wrong next to
39
+ * that neighbour, so removing the context removes the evidence.
40
+ */
41
+ active: ReadonlySet<VerdictFamily>
42
+ selectedId: string | null
43
+ onSelect: (lineId: string) => void
44
  }
45
 
46
+ function SVGOverlay({
47
+ page,
48
+ side,
49
+ opacity,
50
+ withBackground,
51
+ active,
52
+ selectedId,
53
+ onSelect,
54
+ }: SVGOverlayProps) {
55
  const { blocks } = page
56
  const W = page.page_width || blocks.reduce((m, b) => Math.max(m, b.hpos + b.width), 0)
57
  const H = page.page_height || blocks.reduce((m, b) => Math.max(m, b.vpos + b.height), 0)
 
90
  {block.lines.map((line) => {
91
  const displayText = side === 'ocr' ? line.ocr_text : line.corrected_text
92
  const hasHyphen = line.hyphen_role !== 'none'
93
+ const family = verdictFamily(line)
94
+ const isSelected = selectedId === line.line_id
95
+ // Dimmed, not removed — see the `active` prop.
96
+ const dim = active.has(family) ? 1 : 0.16
97
 
98
  if (withBackground) {
99
  // SVG-only mode: coloured text on white page
100
  const fontSize = Math.max(line.height * 0.7, 1)
101
  const textY = line.vpos + line.height * 0.75
102
  return (
103
+ <g
104
+ key={line.line_id}
105
+ opacity={dim}
106
+ onClick={() => onSelect(line.line_id)}
107
+ style={{ cursor: 'pointer' }}
108
+ >
109
  {line.modified && (
110
  <rect
111
  x={line.hpos}
 
142
  const fontSize = Math.max(line.height * 0.72, 1)
143
  const textY = line.vpos + line.height * 0.78
144
  return (
145
+ <g
146
+ key={line.line_id}
147
+ opacity={dim}
148
+ onClick={() => onSelect(line.line_id)}
149
+ style={{ cursor: 'pointer' }}
150
+ >
151
  <rect
152
  x={line.hpos}
153
  y={line.vpos}
154
  width={line.width}
155
  height={line.height}
156
+ fill={FAMILY[family].fill}
157
+ stroke={FAMILY[family].stroke}
158
+ strokeWidth={isSelected ? 6 : 2}
159
  />
160
  {hasHyphen && (
161
  <rect
 
197
  page: LayoutPage
198
  side: 'ocr' | 'corrected'
199
  overlayOpacity: number
200
+ active: ReadonlySet<VerdictFamily>
201
+ selectedId: string | null
202
+ onSelect: (lineId: string) => void
203
  }
204
 
205
+ function PagePanel({ page, side, overlayOpacity, active, selectedId, onSelect }: PagePanelProps) {
206
  const { blocks } = page
207
  const W = page.page_width || blocks.reduce((m, b) => Math.max(m, b.hpos + b.width), 0)
208
  const H = page.page_height || blocks.reduce((m, b) => Math.max(m, b.vpos + b.height), 0)
 
228
  }}
229
  />
230
  )}
231
+ <SVGOverlay
232
+ page={page}
233
+ side={side}
234
+ opacity={overlayOpacity}
235
+ withBackground={!W || !H}
236
+ active={active}
237
+ selectedId={selectedId}
238
+ onSelect={onSelect}
239
+ />
240
  </div>
241
  )
242
  }
243
 
244
  // No image: SVG on white background — opacity still controlled by slider
245
+ return (
246
+ <SVGOverlay
247
+ page={page}
248
+ side={side}
249
+ opacity={overlayOpacity}
250
+ withBackground={true}
251
+ active={active}
252
+ selectedId={selectedId}
253
+ onSelect={onSelect}
254
+ />
255
+ )
256
  }
257
 
258
  // ---------------------------------------------------------------------------
 
261
 
262
  interface LayoutViewerProps {
263
  data: LayoutData
264
+ /**
265
+ * When given, clicking a line opens the judgement panel and the reader's
266
+ * verdict is persisted against this job. Omit it and the view stays
267
+ * read-only — which is what the demo does before a run has a report.
268
+ */
269
+ jobId?: string
270
  }
271
 
272
+ export function LayoutViewer({ data, jobId }: LayoutViewerProps) {
273
  const [pageIdx, setPageIdx] = useState(0)
274
  const [overlayOpacity, setOverlayOpacity] = useState(0.85)
275
+ // All three families on by default: a reviewer opening the page should see
276
+ // what the run did before deciding what to hunt for.
277
+ const [active, setActive] = useState<ReadonlySet<VerdictFamily>>(
278
+ new Set<VerdictFamily>(['kept', 'refused', 'silent']),
279
+ )
280
+ const [selectedId, setSelectedId] = useState<string | null>(null)
281
+ const [reviews, setReviews] = useState<Map<LineKey, LineReview>>(new Map())
282
+ // Remembered per browser, not per job: a reviewer working through one
283
+ // volume pastes the service once and every page of it resolves. The ALTO
284
+ // does not carry the identifier, so guessing it would be inventing
285
+ // provenance — the reader supplies it or the panel shows text only.
286
+ const [iiifService, setIiifService] = useState<string>(
287
+ () => globalThis.localStorage?.getItem('saknussemm.iiif') ?? '',
288
+ )
289
+ const iiifUsable = looksLikeService(iiifService)
290
+
291
+ // Judgements already made travel with the job, so a reader can stop and
292
+ // come back without losing the page they had worked through.
293
+ useEffect(() => {
294
+ if (!jobId) return
295
+ let cancelled = false
296
+ fetchReviews(jobId)
297
+ .then(({ reviews: rows }) => {
298
+ if (cancelled) return
299
+ setReviews(new Map(rows.map((r) => [lineKey(r.page_id, r.line_id), r])))
300
+ })
301
+ .catch(() => {
302
+ // A review list that will not load must not take the layout with it:
303
+ // looking at the page is still worth doing.
304
+ })
305
+ return () => {
306
+ cancelled = true
307
+ }
308
+ }, [jobId])
309
+
310
+ const toggleFamily = useCallback((family: VerdictFamily) => {
311
+ setActive((current) => {
312
+ const next = new Set(current)
313
+ if (next.has(family)) next.delete(family)
314
+ else next.add(family)
315
+ // Turning the last one off would blank the page and read as a bug.
316
+ return next.size ? next : current
317
+ })
318
+ }, [])
319
  const leftRef = useRef<HTMLDivElement>(null)
320
  const rightRef = useRef<HTMLDivElement>(null)
321
  const syncing = useRef(false)
 
346
  }
347
 
348
  const currentPage = data.pages[pageIdx] ?? data.pages[0]
349
+ const selectedLine: LayoutLine | null = selectedId
350
+ ? (currentPage?.blocks.flatMap((b) => b.lines).find((l) => l.line_id === selectedId) ?? null)
351
+ : null
352
  const hasImage = !!currentPage.image_url
353
 
354
  return (
 
378
  </span>
379
  </label>
380
 
381
+ {/* Verdict filter — dims, never hides. A line that copied its
382
+ neighbour only reads as wrong NEXT TO that neighbour, so the
383
+ context has to stay on the page. */}
384
+ <div className="flex items-center gap-1.5">
385
+ {(['kept', 'refused', 'silent'] as const).map((family) => {
386
+ const on = active.has(family)
387
+ return (
388
+ <button
389
+ key={family}
390
+ type="button"
391
+ onClick={() => toggleFamily(family)}
392
+ aria-pressed={on}
393
+ title={`${FAMILY[family].label} — cliquer pour estomper`}
394
+ className="font-mono text-[10px] uppercase tracking-wider rounded px-2 py-1
395
+ border transition-opacity focus:outline-none focus:ring-1
396
+ focus:ring-amber-400"
397
+ style={{
398
+ borderColor: FAMILY[family].stroke,
399
+ color: on ? '#e2e8f0' : '#64748b',
400
+ background: on ? FAMILY[family].fill : 'transparent',
401
+ opacity: on ? 1 : 0.55,
402
+ }}
403
+ >
404
+ {FAMILY[family].label}
405
+ </button>
406
+ )
407
+ })}
408
+ </div>
409
+
410
+ {/* IIIF service — full-resolution line crops, nothing stored */}
411
+ <label className="flex items-center gap-2">
412
+ <span className="font-mono text-[10px] text-slate-500 uppercase tracking-wider whitespace-nowrap">
413
+ IIIF
414
+ </span>
415
+ <input
416
+ type="url"
417
+ value={iiifService}
418
+ placeholder="https://…/iiif/image/v3/ark:/12148/…/f1"
419
+ onChange={(e) => {
420
+ setIiifService(e.target.value)
421
+ globalThis.localStorage?.setItem('saknussemm.iiif', e.target.value)
422
+ }}
423
+ className={`font-mono text-[11px] bg-slate-900 border rounded px-2 py-1 w-56
424
+ text-slate-200 focus:outline-none focus:border-amber-500 ${
425
+ iiifService && !iiifUsable ? 'border-red-500' : 'border-slate-600'
426
+ }`}
427
+ />
428
+ {iiifService && !iiifUsable && (
429
+ <span className="font-mono text-[10px] text-red-400">
430
+ base du service, pas une image
431
+ </span>
432
+ )}
433
+ </label>
434
+
435
  {/* Page selector */}
436
  {data.pages.length > 1 && (
437
  <select
 
466
  {/* Dual panels with synchronised scroll */}
467
  <div className="grid grid-cols-2 divide-x divide-slate-700/40">
468
  <div ref={leftRef} onScroll={onScrollLeft} className="overflow-auto max-h-[60vh]">
469
+ <PagePanel
470
+ page={currentPage}
471
+ side="ocr"
472
+ overlayOpacity={overlayOpacity}
473
+ active={active}
474
+ selectedId={selectedId}
475
+ onSelect={setSelectedId}
476
+ />
477
  </div>
478
  <div ref={rightRef} onScroll={onScrollRight} className="overflow-auto max-h-[60vh]">
479
+ <PagePanel
480
+ page={currentPage}
481
+ side="corrected"
482
+ overlayOpacity={overlayOpacity}
483
+ active={active}
484
+ selectedId={selectedId}
485
+ onSelect={setSelectedId}
486
+ />
487
  </div>
488
  </div>
489
 
490
+ {/* Judgement panel — only with a job to persist against, and only once
491
+ a line is selected. A reviewer who cannot record a verdict produces
492
+ an impression; one who can produces the ground truth these corpora
493
+ do not ship. */}
494
+ {jobId && selectedLine && (
495
+ <div className="px-4 py-3 border-t border-slate-700/40">
496
+ <ReviewPanel
497
+ jobId={jobId}
498
+ pageId={currentPage.page_id}
499
+ line={selectedLine}
500
+ existing={reviews.get(lineKey(currentPage.page_id, selectedLine.line_id)) ?? null}
501
+ iiifService={iiifUsable ? iiifService : null}
502
+ onSaved={(review) =>
503
+ setReviews((current) => {
504
+ const next = new Map(current)
505
+ next.set(lineKey(review.page_id, review.line_id), review)
506
+ return next
507
+ })
508
+ }
509
+ onClose={() => setSelectedId(null)}
510
+ />
511
+ </div>
512
+ )}
513
+
514
  {/* Legend */}
515
  <div className="px-4 py-2.5 border-t border-slate-700/40 flex items-center gap-6 flex-wrap">
516
  <span className="font-mono text-[10px] text-slate-600 uppercase tracking-wider mr-1">
frontend/src/components/ReviewPanel.test.tsx ADDED
@@ -0,0 +1,186 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import { fireEvent, render, screen, waitFor } from '@testing-library/react'
2
+ import { beforeEach, describe, expect, it, vi } from 'vitest'
3
+
4
+ import type { LayoutLine, LineReview } from '../types'
5
+ import { ReviewPanel } from './ReviewPanel'
6
+
7
+ vi.mock('../api/client', () => ({
8
+ putReviews: vi.fn(),
9
+ }))
10
+
11
+ import { putReviews } from '../api/client'
12
+
13
+ const line = (over: Partial<LayoutLine> = {}): LayoutLine => ({
14
+ line_id: 'TL000323',
15
+ hpos: 0,
16
+ vpos: 0,
17
+ width: 100,
18
+ height: 20,
19
+ ocr_text: '1,500 ETANTS : LE NOMBRE DES MORTS',
20
+ corrected_text: '1,500 ETANTS : LE NOMBRE DES MORTS',
21
+ modified: false,
22
+ hyphen_role: 'none',
23
+ verdict: 'absorbs_previous_line',
24
+ verdict_detail: null,
25
+ proposed_text: "AU COURS D'UNE MATINÉE 1,500 ÉTANTS : LE NOMBRE DES MORTS",
26
+ proposal_declined: true,
27
+ ...over,
28
+ })
29
+
30
+ function mount(over: Partial<LayoutLine> = {}, existing: LineReview | null = null) {
31
+ const onSaved = vi.fn()
32
+ render(
33
+ <ReviewPanel
34
+ jobId="job-1"
35
+ pageId="PAG_1"
36
+ line={line(over)}
37
+ existing={existing}
38
+ onSaved={onSaved}
39
+ onClose={vi.fn()}
40
+ />,
41
+ )
42
+ return { onSaved }
43
+ }
44
+
45
+ describe('ReviewPanel', () => {
46
+ beforeEach(() => {
47
+ vi.mocked(putReviews).mockReset()
48
+ })
49
+
50
+ it('shows what was proposed and that it was declined', () => {
51
+ // The reviewer's most valuable case: something WAS on the table and a
52
+ // guard refused it. Hiding the proposal would hide whether the refusal
53
+ // saved a hallucination or threw away a good correction.
54
+ mount()
55
+ expect(screen.getByText(/proposé — et refusé/i)).toBeInTheDocument()
56
+ expect(screen.getByText(/AU COURS D'UNE MATINÉE/)).toBeInTheDocument()
57
+ })
58
+
59
+ it('refuses an empty transcription before it reaches the server', () => {
60
+ // The text IS the review. An empty one would land in the ground-truth set
61
+ // asserting nothing, indistinguishable later from a real reading.
62
+ mount()
63
+ fireEvent.click(screen.getByRole('button', { name: /ce que je lis/i }))
64
+
65
+ expect(screen.getByRole('alert')).toHaveTextContent(/n'affirme rien/i)
66
+ expect(putReviews).not.toHaveBeenCalled()
67
+ })
68
+
69
+ it('sends the transcription the reader typed', async () => {
70
+ vi.mocked(putReviews).mockResolvedValue([
71
+ {
72
+ page_id: 'PAG_1',
73
+ line_id: 'TL000323',
74
+ verdict: 'transcribed',
75
+ transcription: '1,500 ENFANTS : LE NOMBRE DES MORTS',
76
+ note: null,
77
+ reviewed_at: '2026-08-18T10:00:00+00:00',
78
+ },
79
+ ])
80
+ const { onSaved } = mount()
81
+
82
+ fireEvent.change(screen.getByRole('textbox', { name: /ce que je lis sur le scan/i }), {
83
+ target: { value: '1,500 ENFANTS : LE NOMBRE DES MORTS' },
84
+ })
85
+ fireEvent.click(screen.getByRole('button', { name: /ce que je lis/i }))
86
+
87
+ await waitFor(() => expect(putReviews).toHaveBeenCalledTimes(1))
88
+ const [jobId, reviews] = vi.mocked(putReviews).mock.calls[0]
89
+ expect(jobId).toBe('job-1')
90
+ expect(reviews[0]).toMatchObject({
91
+ page_id: 'PAG_1',
92
+ line_id: 'TL000323',
93
+ verdict: 'transcribed',
94
+ transcription: '1,500 ENFANTS : LE NOMBRE DES MORTS',
95
+ })
96
+ await waitFor(() => expect(onSaved).toHaveBeenCalled())
97
+ })
98
+
99
+ it('grades the engine without needing a transcription', async () => {
100
+ vi.mocked(putReviews).mockResolvedValue([
101
+ {
102
+ page_id: 'PAG_1',
103
+ line_id: 'TL000323',
104
+ verdict: 'accepted',
105
+ transcription: null,
106
+ note: null,
107
+ reviewed_at: '2026-08-18T10:00:00+00:00',
108
+ },
109
+ ])
110
+ mount()
111
+ fireEvent.click(screen.getByRole('button', { name: /le moteur a eu raison/i }))
112
+ await waitFor(() => expect(putReviews).toHaveBeenCalledTimes(1))
113
+ expect(vi.mocked(putReviews).mock.calls[0][1][0].verdict).toBe('accepted')
114
+ })
115
+
116
+ it("surfaces the server's own words when it rejects a review", async () => {
117
+ // A status code the reader cannot act on is worse than no feedback: the
118
+ // server explains an empty transcription by name, so show that.
119
+ vi.mocked(putReviews).mockRejectedValue(new Error('a transcribed review must carry the text'))
120
+ mount()
121
+ fireEvent.click(screen.getByRole('button', { name: /le moteur a eu tort/i }))
122
+ await waitFor(() => expect(screen.getByRole('alert')).toHaveTextContent(/must carry the text/i))
123
+ })
124
+
125
+ it('shows an existing judgement as the pressed one', () => {
126
+ mount(
127
+ {},
128
+ {
129
+ page_id: 'PAG_1',
130
+ line_id: 'TL000323',
131
+ verdict: 'refused',
132
+ transcription: null,
133
+ note: 'le scan dit ENFANTS',
134
+ reviewed_at: '2026-08-18T09:00:00+00:00',
135
+ },
136
+ )
137
+ expect(screen.getByRole('button', { name: /le moteur a eu tort/i })).toHaveAttribute(
138
+ 'aria-pressed',
139
+ 'true',
140
+ )
141
+ expect(screen.getByDisplayValue('le scan dit ENFANTS')).toBeInTheDocument()
142
+ })
143
+ })
144
+
145
+ describe('ReviewPanel — le scan en pleine résolution', () => {
146
+ it('shows the line from IIIF when a service is known', () => {
147
+ // Judging a word from a downscaled preview means judging ~13 pixels of
148
+ // newspaper line. The region request costs nothing to store and returns
149
+ // the scan's native pixels.
150
+ render(
151
+ <ReviewPanel
152
+ jobId="job-1"
153
+ pageId="PAG_1"
154
+ line={line({ hpos: 1000, vpos: 2000, width: 500, height: 100 })}
155
+ existing={null}
156
+ iiifService="https://example.org/iiif/f1"
157
+ onSaved={vi.fn()}
158
+ onClose={vi.fn()}
159
+ />,
160
+ )
161
+ const img = screen.getByRole('img', { name: /ligne TL000323 sur le scan/i })
162
+ expect(img).toHaveAttribute(
163
+ 'src',
164
+ 'https://example.org/iiif/f1/995,1980,510,140/max/0/default.jpg',
165
+ )
166
+ })
167
+
168
+ it('falls back to text alone when no service is known', () => {
169
+ // The ALTO does not carry a IIIF identifier, so its absence is normal and
170
+ // must not cost the reviewer the rest of the panel.
171
+ render(
172
+ <ReviewPanel
173
+ jobId="job-1"
174
+ pageId="PAG_1"
175
+ line={line()}
176
+ existing={null}
177
+ onSaved={vi.fn()}
178
+ onClose={vi.fn()}
179
+ />,
180
+ )
181
+ expect(screen.queryByRole('img')).not.toBeInTheDocument()
182
+ // The source and the retained text are identical on a refused line, so
183
+ // both nodes carry it — the point is that the panel still renders.
184
+ expect(screen.getAllByText(/1,500 ETANTS/).length).toBeGreaterThan(0)
185
+ })
186
+ })
frontend/src/components/ReviewPanel.tsx ADDED
@@ -0,0 +1,215 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import { useEffect, useState } from 'react'
2
+
3
+ import { putReviews } from '../api/client'
4
+ import { lineRegionUrl } from '../lib/iiif'
5
+ import type { LayoutLine, LineReview, ReviewVerdict } from '../types'
6
+
7
+ /**
8
+ * Where a reader's judgement on one line is made.
9
+ *
10
+ * The corpora this demo runs on carry no ground truth — Gallica's `text.txt`
11
+ * is the same OCR layer as its ALTO — so nothing on disk can say whether a
12
+ * correction was right or a refusal justified. Only a person reading the scan
13
+ * can, and this panel is where that reading is captured.
14
+ *
15
+ * Three verdicts, and the third is the one that compounds. `accepted` and
16
+ * `refused` grade what the engine did, which is useful for tuning it.
17
+ * `transcribed` records what the reader actually read, which stands on its own
18
+ * whatever the engine decided — and accumulates, line by line, into the ground
19
+ * truth the bench needs to answer "was the correction right".
20
+ */
21
+
22
+ interface ReviewPanelProps {
23
+ jobId: string
24
+ pageId: string
25
+ line: LayoutLine
26
+ /** The reader's existing judgement on this line, if any. */
27
+ existing: LineReview | null
28
+ /**
29
+ * IIIF Image API service base for this page, when the reader has one.
30
+ * With it, the line is shown at the scan's NATIVE resolution and nothing is
31
+ * stored — the alternative is judging a word from a downscaled preview,
32
+ * which on a newspaper line is roughly 13 pixels tall.
33
+ */
34
+ iiifService?: string | null
35
+ onSaved: (review: LineReview) => void
36
+ onClose: () => void
37
+ }
38
+
39
+ const VERDICT_LABEL: Record<ReviewVerdict, string> = {
40
+ accepted: 'Le moteur a eu raison',
41
+ refused: 'Le moteur a eu tort',
42
+ transcribed: 'Voici ce que je lis',
43
+ }
44
+
45
+ export function ReviewPanel({
46
+ jobId,
47
+ pageId,
48
+ line,
49
+ existing,
50
+ iiifService,
51
+ onSaved,
52
+ onClose,
53
+ }: ReviewPanelProps) {
54
+ const [transcription, setTranscription] = useState(existing?.transcription ?? '')
55
+ const [note, setNote] = useState(existing?.note ?? '')
56
+ const [saving, setSaving] = useState(false)
57
+ const [error, setError] = useState<string | null>(null)
58
+
59
+ // A reader moves line to line; the fields must follow the selection rather
60
+ // than carry the previous line's text into the next one's judgement.
61
+ useEffect(() => {
62
+ setTranscription(existing?.transcription ?? '')
63
+ setNote(existing?.note ?? '')
64
+ setError(null)
65
+ }, [line.line_id, pageId, existing])
66
+
67
+ async function save(verdict: ReviewVerdict) {
68
+ if (verdict === 'transcribed' && !transcription.trim()) {
69
+ setError("Une transcription vide n'affirme rien : écrivez ce que vous lisez sur le scan.")
70
+ return
71
+ }
72
+ setSaving(true)
73
+ setError(null)
74
+ try {
75
+ const review: LineReview = {
76
+ page_id: pageId,
77
+ line_id: line.line_id,
78
+ verdict,
79
+ transcription: transcription.trim() || null,
80
+ note: note.trim() || null,
81
+ }
82
+ const saved = await putReviews(jobId, [review])
83
+ const mine = saved.find((r) => r.page_id === pageId && r.line_id === line.line_id)
84
+ if (mine) onSaved(mine)
85
+ } catch (caught) {
86
+ setError(caught instanceof Error ? caught.message : String(caught))
87
+ } finally {
88
+ setSaving(false)
89
+ }
90
+ }
91
+
92
+ const declined = line.proposal_declined
93
+
94
+ return (
95
+ <aside
96
+ className="border border-slate-700 bg-slate-800 rounded p-4 flex flex-col gap-3"
97
+ aria-label={`Jugement sur la ligne ${line.line_id}`}
98
+ >
99
+ <header className="flex items-baseline justify-between gap-3">
100
+ <span className="font-mono text-[10px] uppercase tracking-wider text-slate-400">
101
+ {line.line_id}
102
+ </span>
103
+ <button
104
+ type="button"
105
+ onClick={onClose}
106
+ className="font-mono text-[10px] text-slate-400 hover:text-slate-200"
107
+ >
108
+ fermer
109
+ </button>
110
+ </header>
111
+
112
+ {iiifService && (
113
+ <figure className="m-0">
114
+ <img
115
+ src={lineRegionUrl(iiifService, line)}
116
+ alt={`Ligne ${line.line_id} sur le scan`}
117
+ className="w-full rounded border border-slate-600 bg-white"
118
+ />
119
+ <figcaption className="font-mono text-[10px] text-slate-500 mt-1">
120
+ résolution native, servie par IIIF — rien n'est stocké
121
+ </figcaption>
122
+ </figure>
123
+ )}
124
+
125
+ <dl className="flex flex-col gap-2 text-xs">
126
+ <div>
127
+ <dt className="font-mono text-[10px] uppercase tracking-wider text-slate-500">
128
+ OCR source
129
+ </dt>
130
+ <dd className="font-mono text-slate-200 break-words">{line.ocr_text}</dd>
131
+ </div>
132
+ {line.proposed_text && line.proposed_text !== line.ocr_text && (
133
+ <div>
134
+ <dt className="font-mono text-[10px] uppercase tracking-wider text-slate-500">
135
+ Proposé{declined ? ' — et refusé' : ''}
136
+ </dt>
137
+ <dd className={`font-mono break-words ${declined ? 'text-red-300' : 'text-blue-300'}`}>
138
+ {line.proposed_text}
139
+ </dd>
140
+ </div>
141
+ )}
142
+ <div>
143
+ <dt className="font-mono text-[10px] uppercase tracking-wider text-slate-500">Retenu</dt>
144
+ <dd className="font-mono text-slate-100 break-words">{line.corrected_text}</dd>
145
+ </div>
146
+ {line.verdict && (
147
+ <div>
148
+ <dt className="font-mono text-[10px] uppercase tracking-wider text-slate-500">
149
+ Verdict du moteur
150
+ </dt>
151
+ <dd className="font-mono text-amber-300 break-words">
152
+ {line.verdict}
153
+ {line.verdict_detail ? ` — ${line.verdict_detail}` : ''}
154
+ </dd>
155
+ </div>
156
+ )}
157
+ </dl>
158
+
159
+ <label className="flex flex-col gap-1">
160
+ <span className="font-mono text-[10px] uppercase tracking-wider text-slate-500">
161
+ Ce que je lis sur le scan
162
+ </span>
163
+ <textarea
164
+ value={transcription}
165
+ onChange={(e) => setTranscription(e.target.value)}
166
+ rows={2}
167
+ spellCheck={false}
168
+ className="font-mono text-xs bg-slate-900 border border-slate-600 text-slate-100
169
+ rounded px-2 py-1 focus:outline-none focus:border-amber-500"
170
+ />
171
+ </label>
172
+
173
+ <label className="flex flex-col gap-1">
174
+ <span className="font-mono text-[10px] uppercase tracking-wider text-slate-500">Note</span>
175
+ <input
176
+ value={note}
177
+ onChange={(e) => setNote(e.target.value)}
178
+ className="font-mono text-xs bg-slate-900 border border-slate-600 text-slate-100
179
+ rounded px-2 py-1 focus:outline-none focus:border-amber-500"
180
+ />
181
+ </label>
182
+
183
+ {error && (
184
+ <p role="alert" className="text-xs text-red-300">
185
+ {error}
186
+ </p>
187
+ )}
188
+
189
+ <div className="flex flex-wrap gap-2">
190
+ {(['accepted', 'refused', 'transcribed'] as const).map((verdict) => (
191
+ <button
192
+ key={verdict}
193
+ type="button"
194
+ disabled={saving}
195
+ onClick={() => save(verdict)}
196
+ aria-pressed={existing?.verdict === verdict}
197
+ className={`font-mono text-[11px] rounded px-3 py-1.5 border transition-colors
198
+ disabled:opacity-50 focus:outline-none focus:ring-1
199
+ focus:ring-amber-400 ${
200
+ existing?.verdict === verdict
201
+ ? 'bg-amber-500 border-amber-400 text-slate-900'
202
+ : 'bg-slate-700 border-slate-600 text-slate-200 hover:border-amber-500'
203
+ }`}
204
+ >
205
+ {VERDICT_LABEL[verdict]}
206
+ </button>
207
+ ))}
208
+ </div>
209
+
210
+ {existing?.reviewed_at && (
211
+ <p className="font-mono text-[10px] text-slate-500">Jugé le {existing.reviewed_at}</p>
212
+ )}
213
+ </aside>
214
+ )
215
+ }
frontend/src/lib/iiif.test.ts ADDED
@@ -0,0 +1,66 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import { describe, expect, it } from 'vitest'
2
+
3
+ import { lineRegionUrl, looksLikeService } from './iiif'
4
+
5
+ describe('lineRegionUrl', () => {
6
+ it('passes ALTO coordinates through as the IIIF region', () => {
7
+ // The two spaces are the same one — a service serving the full scan
8
+ // declares the page's own dimensions — so no scaling belongs here.
9
+ const url = lineRegionUrl(
10
+ 'https://example.org/iiif/f1',
11
+ {
12
+ hpos: 1000,
13
+ vpos: 2000,
14
+ width: 500,
15
+ height: 100,
16
+ },
17
+ 0,
18
+ )
19
+ expect(url).toBe('https://example.org/iiif/f1/995,2000,510,100/max/0/default.jpg')
20
+ })
21
+
22
+ it('asks for max rather than a fixed width', () => {
23
+ // Measured against Gallica: a fixed width 400s the moment it would
24
+ // upscale, and lines are ~70px tall so it almost always would.
25
+ const url = lineRegionUrl('https://example.org/iiif/f1', {
26
+ hpos: 0,
27
+ vpos: 0,
28
+ width: 100,
29
+ height: 20,
30
+ })
31
+ expect(url).toContain('/max/0/default.jpg')
32
+ expect(url).not.toMatch(/\/\d+,\/0\//)
33
+ })
34
+
35
+ it('tolerates a trailing slash on the service', () => {
36
+ const url = lineRegionUrl(
37
+ 'https://example.org/iiif/f1/',
38
+ {
39
+ hpos: 0,
40
+ vpos: 0,
41
+ width: 10,
42
+ height: 10,
43
+ },
44
+ 0,
45
+ )
46
+ expect(url).not.toContain('f1//')
47
+ })
48
+ })
49
+
50
+ describe('looksLikeService', () => {
51
+ it('accepts a service base', () => {
52
+ expect(looksLikeService('https://openapi.bnf.fr/iiif/image/v3/ark:/12148/x/f1')).toBe(true)
53
+ })
54
+
55
+ it('rejects a full image URL, which would nest two region paths', () => {
56
+ expect(
57
+ looksLikeService(
58
+ 'https://openapi.bnf.fr/iiif/image/v3/ark:/12148/x/f1/full/max/0/default.jpg',
59
+ ),
60
+ ).toBe(false)
61
+ })
62
+
63
+ it('rejects anything that is not a URL', () => {
64
+ expect(looksLikeService('bpt6k4607951t')).toBe(false)
65
+ })
66
+ })
frontend/src/lib/iiif.ts ADDED
@@ -0,0 +1,50 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /**
2
+ * IIIF Image API region URLs — a line's own pixels, without storing a thing.
3
+ *
4
+ * ALTO coordinates and the IIIF region parameter live in the SAME space: a
5
+ * service serving the full scan declares exactly the page's own dimensions
6
+ * (measured on Gallica: `info.json` says 6802×9121 for an ALTO page of
7
+ * 6802×9121). So the coordinates go through verbatim — no scale, no
8
+ * transform, and no local derivative to crop.
9
+ *
10
+ * That matters beyond convenience. Cropping a DOWNSCALED derivative locally
11
+ * is how a line ends up cut from the wrong place: Gallica's `!1600,1600`
12
+ * preview is 17.5% of the ALTO space, so a line at hpos=4798 falls off an
13
+ * image 1193 pixels wide, and the crop comes back blank with nothing
14
+ * downstream any the wiser.
15
+ */
16
+
17
+ export interface Region {
18
+ hpos: number
19
+ vpos: number
20
+ width: number
21
+ height: number
22
+ }
23
+
24
+ /**
25
+ * `size` is `max`, deliberately.
26
+ *
27
+ * Asking for a fixed width is refused with HTTP 400 the moment it would
28
+ * UPSCALE the region — measured against Gallica, whose lines are ~70px tall:
29
+ * `/1200,/` returns 400, `/max/` returns 200. IIIF 3.0 requires an explicit
30
+ * `^` prefix to allow upscaling, and a line does not need it: its native
31
+ * pixels are exactly what a reader wants to see.
32
+ */
33
+ export function lineRegionUrl(service: string, region: Region, marginRatio = 0.2): string {
34
+ const marginX = Math.round(region.width * 0.01)
35
+ const marginY = Math.round(region.height * marginRatio)
36
+ const x = Math.max(region.hpos - marginX, 0)
37
+ const y = Math.max(region.vpos - marginY, 0)
38
+ const w = region.width + 2 * marginX
39
+ const h = region.height + 2 * marginY
40
+ return `${service.replace(/\/+$/, '')}/${x},${y},${w},${h}/max/0/default.jpg`
41
+ }
42
+
43
+ /** Whether a pasted string looks like an Image API service base, not a full image URL. */
44
+ export function looksLikeService(value: string): boolean {
45
+ const trimmed = value.trim()
46
+ if (!/^https?:\/\//i.test(trimmed)) return false
47
+ // A full image URL already carries region/size/rotation/quality; using it as
48
+ // a base would produce nonsense like `.../full/max/0/default.jpg/12,3,4,5/max/...`.
49
+ return !/\/(full|square|\d+,\d+,\d+,\d+)\/[^/]+\/\d+\/[^/]+\.\w+$/i.test(trimmed)
50
+ }
frontend/src/lib/verdicts.ts ADDED
@@ -0,0 +1,45 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ // ---------------------------------------------------------------------------
2
+ // Verdict palette — what the engine decided, and why
3
+ // ---------------------------------------------------------------------------
4
+ //
5
+ // Three families, because a reviewer scans for three different things and
6
+ // should not have to read a legend to tell them apart:
7
+ //
8
+ // kept the correction survived every guard
9
+ // refused something was proposed and a guard declined it — the cases
10
+ // worth a human eye, since each is either a caught hallucination
11
+ // or a good correction thrown away
12
+ // silent nothing was proposed, or nothing changed; no judgement to make
13
+ //
14
+ // Colours are the proof-reader's two pencils: blue marks what stands, red
15
+ // marks what was struck out. Amber sits between them for a line the engine
16
+ // never got an answer for.
17
+
18
+ export type VerdictFamily = 'kept' | 'refused' | 'silent'
19
+
20
+ const REFUSAL_CODES = new Set([
21
+ 'too_different_from_source',
22
+ 'closer_to_previous_line',
23
+ 'closer_to_next_line',
24
+ 'absorbs_previous_line',
25
+ 'absorbs_next_line',
26
+ 'hyphen_pair_fallback',
27
+ 'boundary_migration_forward',
28
+ 'boundary_migration_backward',
29
+ 'adjacent_duplicate_detected',
30
+ 'adjacent_duplicate_pair_atomicity',
31
+ 'orphan_hyphen_completed',
32
+ 'hyphen_unit_fallback',
33
+ ])
34
+
35
+ export function verdictFamily(line: { verdict: string | null; modified: boolean }): VerdictFamily {
36
+ if (line.verdict && REFUSAL_CODES.has(line.verdict)) return 'refused'
37
+ if (line.verdict === 'all_attempts_exhausted') return 'silent'
38
+ return line.modified ? 'kept' : 'silent'
39
+ }
40
+
41
+ export const FAMILY = {
42
+ kept: { stroke: '#1d4ed8', fill: 'rgba(29,78,216,0.18)', label: 'Retenue' },
43
+ refused: { stroke: '#b91c1c', fill: 'rgba(185,28,28,0.20)', label: 'Refusée' },
44
+ silent: { stroke: '#a1a1aa', fill: 'rgba(161,161,170,0.10)', label: 'Sans objet' },
45
+ } as const
frontend/src/types/api.generated.ts CHANGED
@@ -289,6 +289,32 @@ export interface paths {
289
  patch?: never
290
  trace?: never
291
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
292
  }
293
  export type webhooks = Record<string, never>
294
  export interface components {
@@ -397,6 +423,23 @@ export interface components {
397
  /** Error */
398
  error?: string | null
399
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
400
  /** ListModelsRequest */
401
  ListModelsRequest: {
402
  provider: components['schemas']['Provider']
@@ -432,6 +475,27 @@ export interface components {
432
  * @enum {string}
433
  */
434
  Provider: 'openai' | 'anthropic' | 'mistral' | 'google'
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
435
  /** ValidationError */
436
  ValidationError: {
437
  /** Location */
@@ -837,4 +901,70 @@ export interface operations {
837
  }
838
  }
839
  }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
840
  }
 
289
  patch?: never
290
  trace?: never
291
  }
292
+ '/api/jobs/{job_id}/reviews': {
293
+ parameters: {
294
+ query?: never
295
+ header?: never
296
+ path?: never
297
+ cookie?: never
298
+ }
299
+ /** Get Reviews */
300
+ get: operations['get_reviews_api_jobs__job_id__reviews_get']
301
+ /**
302
+ * Put Reviews
303
+ * @description Record or replace judgements on lines of this job.
304
+ *
305
+ * Idempotent per line: sending the same line twice replaces its review
306
+ * rather than appending, so a reader who changes their mind is not fighting
307
+ * an append-only log. The timestamp is stamped here rather than trusted
308
+ * from the client — a review's date is a fact about the server.
309
+ */
310
+ put: operations['put_reviews_api_jobs__job_id__reviews_put']
311
+ post?: never
312
+ delete?: never
313
+ options?: never
314
+ head?: never
315
+ patch?: never
316
+ trace?: never
317
+ }
318
  }
319
  export type webhooks = Record<string, never>
320
  export interface components {
 
423
  /** Error */
424
  error?: string | null
425
  }
426
+ /**
427
+ * LineReview
428
+ * @description One reader's judgement on one line.
429
+ */
430
+ LineReview: {
431
+ /** Page Id */
432
+ page_id: string
433
+ /** Line Id */
434
+ line_id: string
435
+ verdict: components['schemas']['ReviewVerdict']
436
+ /** Transcription */
437
+ transcription?: string | null
438
+ /** Note */
439
+ note?: string | null
440
+ /** Reviewed At */
441
+ reviewed_at?: string | null
442
+ }
443
  /** ListModelsRequest */
444
  ListModelsRequest: {
445
  provider: components['schemas']['Provider']
 
475
  * @enum {string}
476
  */
477
  Provider: 'openai' | 'anthropic' | 'mistral' | 'google'
478
+ /**
479
+ * ReviewBatch
480
+ * @description Reviews arrive in batches: a reader works through a page, not a line.
481
+ */
482
+ ReviewBatch: {
483
+ /** Reviews */
484
+ reviews: components['schemas']['LineReview'][]
485
+ }
486
+ /**
487
+ * ReviewVerdict
488
+ * @description What the reader concluded about the engine's decision on this line.
489
+ * @enum {string}
490
+ */
491
+ ReviewVerdict: 'accepted' | 'refused' | 'transcribed'
492
+ /** ReviewsResponse */
493
+ ReviewsResponse: {
494
+ /** Job Id */
495
+ job_id: string
496
+ /** Reviews */
497
+ reviews: components['schemas']['LineReview'][]
498
+ }
499
  /** ValidationError */
500
  ValidationError: {
501
  /** Location */
 
901
  }
902
  }
903
  }
904
+ get_reviews_api_jobs__job_id__reviews_get: {
905
+ parameters: {
906
+ query?: never
907
+ header?: never
908
+ path: {
909
+ job_id: string
910
+ }
911
+ cookie?: never
912
+ }
913
+ requestBody?: never
914
+ responses: {
915
+ /** @description Successful Response */
916
+ 200: {
917
+ headers: {
918
+ [name: string]: unknown
919
+ }
920
+ content: {
921
+ 'application/json': components['schemas']['ReviewsResponse']
922
+ }
923
+ }
924
+ /** @description Validation Error */
925
+ 422: {
926
+ headers: {
927
+ [name: string]: unknown
928
+ }
929
+ content: {
930
+ 'application/json': components['schemas']['HTTPValidationError']
931
+ }
932
+ }
933
+ }
934
+ }
935
+ put_reviews_api_jobs__job_id__reviews_put: {
936
+ parameters: {
937
+ query?: never
938
+ header?: never
939
+ path: {
940
+ job_id: string
941
+ }
942
+ cookie?: never
943
+ }
944
+ requestBody: {
945
+ content: {
946
+ 'application/json': components['schemas']['ReviewBatch']
947
+ }
948
+ }
949
+ responses: {
950
+ /** @description Successful Response */
951
+ 200: {
952
+ headers: {
953
+ [name: string]: unknown
954
+ }
955
+ content: {
956
+ 'application/json': components['schemas']['ReviewsResponse']
957
+ }
958
+ }
959
+ /** @description Validation Error */
960
+ 422: {
961
+ headers: {
962
+ [name: string]: unknown
963
+ }
964
+ content: {
965
+ 'application/json': components['schemas']['HTTPValidationError']
966
+ }
967
+ }
968
+ }
969
+ }
970
  }
frontend/src/types/index.ts CHANGED
@@ -211,6 +211,31 @@ export interface LayoutLine {
211
  corrected_text: string
212
  modified: boolean
213
  hyphen_role: 'none' | 'HypPart1' | 'HypPart2'
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
214
  }
215
 
216
  export interface LayoutBlock {
 
211
  corrected_text: string
212
  modified: boolean
213
  hyphen_role: 'none' | 'HypPart1' | 'HypPart2'
214
+ /**
215
+ * Why the line ended up as it did — the guard code, or the decision status
216
+ * when no guard spoke. Geometry alone cannot distinguish "nothing was
217
+ * proposed" from "a hallucination was refused", and a reviewer needs to.
218
+ * `null` when the job carries no report yet.
219
+ */
220
+ verdict: string | null
221
+ verdict_detail: string | null
222
+ /** What the producer actually returned, before the engine judged it. */
223
+ proposed_text: string | null
224
+ /** A proposal was on the table and the engine declined it — the case worth reading. */
225
+ proposal_declined: boolean
226
+ }
227
+
228
+ /** How a reader graded the engine on one line. */
229
+ export type ReviewVerdict = 'accepted' | 'refused' | 'transcribed'
230
+
231
+ export interface LineReview {
232
+ page_id: string
233
+ line_id: string
234
+ verdict: ReviewVerdict
235
+ /** What the reader read on the scan. Required for `transcribed`. */
236
+ transcription?: string | null
237
+ note?: string | null
238
+ reviewed_at?: string | null
239
  }
240
 
241
  export interface LayoutBlock {
tools/run_corpus_arms.py ADDED
@@ -0,0 +1,261 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Run a corpus through the demo's own JobRunner, one arm at a time.
2
+
3
+ python tools/run_corpus_arms.py <corpus-root> <arm> [concurrency]
4
+
5
+ Arms:
6
+
7
+ small-text mistral-small-latest, no image
8
+ small-vision mistral-small-latest, per-line crops
9
+ ministral-text ministral-8b-latest, no image
10
+ ministral-vision ministral-8b-latest, per-line crops
11
+
12
+ `<corpus-root>` is scanned for ``alto.xml`` files, each with an optional
13
+ ``image.jpg`` beside it — the layout a digital library's page-by-page export
14
+ produces, and the shape ``gallica_multicolumn`` writes.
15
+
16
+ This drives ``JobRunner`` rather than the HTTP API: same producer assembly,
17
+ same guards, same output writer, no server to stand up. The key is read from
18
+ the macOS keychain and never appears in a flag, an environment dump, or an
19
+ artefact — a flag would land in shell history and in the process list.
20
+
21
+ One JSON per page under ``<corpus-root>/../arm-results/<arm>/``, so an
22
+ interrupted run resumes and nothing measured is lost.
23
+ """
24
+
25
+ from __future__ import annotations
26
+
27
+ import asyncio
28
+ import json
29
+ import subprocess
30
+ import sys
31
+ import time
32
+ from collections import Counter
33
+ from pathlib import Path
34
+
35
+ sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "backend"))
36
+
37
+ from saknussemm.formats.loader import build_document_manifest
38
+
39
+ from app.jobs.runner import JobRunner, page_image_assets
40
+ from app.jobs.store import JobStore
41
+ from app.providers.mistral_multimodal import MistralMultimodalProvider
42
+ from app.providers.mistral_provider import MistralProvider
43
+ from app.storage.output_writer import FilesystemOutputWriter
44
+
45
+ ARMS: dict[str, tuple[str, bool]] = {
46
+ "small-text": ("mistral-small-latest", False),
47
+ "small-vision": ("mistral-small-latest", True),
48
+ "ministral-text": ("ministral-8b-latest", False),
49
+ "ministral-vision": ("ministral-8b-latest", True),
50
+ }
51
+
52
+
53
+ def keychain_key() -> str:
54
+ return subprocess.run(
55
+ ["security", "find-generic-password", "-a", "marcel", "-s", "mistral-api-key", "-w"],
56
+ capture_output=True,
57
+ text=True,
58
+ check=True,
59
+ ).stdout.strip()
60
+
61
+
62
+ def find_pages(root: Path) -> list[tuple[str, Path, Path | None]]:
63
+ """``(label, alto path, image path or None)`` for every page under *root*."""
64
+ pages = []
65
+ for alto in sorted(root.rglob("alto.xml")):
66
+ label = f"{alto.parent.parent.name}_{alto.parent.name.split('_')[-1]}"
67
+ image = alto.parent / "image.jpg"
68
+ pages.append((label, alto, image if image.exists() else None))
69
+ return pages
70
+
71
+
72
+ async def run_one(
73
+ arm: str,
74
+ label: str,
75
+ alto: Path,
76
+ image: Path | None,
77
+ api_key: str,
78
+ results: Path,
79
+ sem: asyncio.Semaphore,
80
+ ) -> dict | None:
81
+ target = results / f"{label}.json"
82
+ if target.exists():
83
+ return json.loads(target.read_text())
84
+
85
+ model, wants_image = ARMS[arm]
86
+ if wants_image and image is None:
87
+ return None
88
+
89
+ async with sem:
90
+ from app.schemas import Provider
91
+
92
+ store = JobStore()
93
+ job_id = store.create_job(provider=Provider.MISTRAL, model=model)
94
+ manifest = build_document_manifest([(alto, alto.name)])
95
+
96
+ images = None
97
+ if wants_image:
98
+ images = page_image_assets(manifest, {alto.parent.name.lower(): image})
99
+ if not images:
100
+ # The demo keys images by source-file stem; this corpus names
101
+ # every file `alto.xml`, so map it explicitly by page instead of
102
+ # letting the helper's stem lookup miss silently.
103
+ import hashlib
104
+
105
+ from saknussemm.core.schemas import ImageAsset
106
+
107
+ digest = hashlib.sha256(image.read_bytes()).hexdigest()
108
+ images = {
109
+ page.page_id: ImageAsset(
110
+ page_id=page.page_id,
111
+ uri=str(image),
112
+ sha256=digest,
113
+ media_type="image/jpeg",
114
+ )
115
+ for page in manifest.pages
116
+ }
117
+
118
+ out_dir = results / "outputs" / label
119
+ out_dir.mkdir(parents=True, exist_ok=True)
120
+ writer = FilesystemOutputWriter(out_dir)
121
+ runner = JobRunner(store)
122
+
123
+ started = time.perf_counter()
124
+ await runner.run(
125
+ job_id=job_id,
126
+ document_manifest=manifest,
127
+ provider_name="mistral",
128
+ api_key=api_key,
129
+ model=model,
130
+ output_writer=writer,
131
+ source_files={alto.name: alto},
132
+ provider=(MistralMultimodalProvider() if wants_image else MistralProvider()),
133
+ page_images=images,
134
+ timeout_seconds=0,
135
+ )
136
+ elapsed = time.perf_counter() - started
137
+
138
+ job = store.get_job(job_id)
139
+ report = job.report if job else None
140
+ record: dict = {
141
+ "arm": arm,
142
+ "page": label,
143
+ "model": model,
144
+ "wants_image": wants_image,
145
+ "seconds": round(elapsed, 1),
146
+ "status": job.status.value if job else "unknown",
147
+ "error": job.error if job else None,
148
+ "lines": job.total_lines if job else 0,
149
+ "fallbacks": job.fallbacks if job else 0,
150
+ "retries": job.retries if job else 0,
151
+ }
152
+ if report is not None:
153
+ reasons = Counter(
154
+ line.decision.reason.code
155
+ for line in report.lines
156
+ if line.decision.reason is not None
157
+ )
158
+ usage = report.usage
159
+ record.update(
160
+ input_tokens=getattr(usage, "input_tokens", 0) if usage else 0,
161
+ output_tokens=getattr(usage, "output_tokens", 0) if usage else 0,
162
+ reasons=dict(reasons),
163
+ statuses=dict(Counter(ln.decision.status for ln in report.lines)),
164
+ changed=sum(1 for ln in report.lines if ln.decision.final_text != ln.source_text),
165
+ blocked=sum(
166
+ 1
167
+ for ln in report.lines
168
+ if ln.decision.reason is not None
169
+ and ln.proposal is not None
170
+ and ln.proposal.output_text != ln.source_text
171
+ ),
172
+ format_losses=dict(report.format_losses or {}),
173
+ samples=[
174
+ {"src": ln.source_text, "out": ln.decision.final_text}
175
+ for ln in report.lines
176
+ if ln.decision.final_text != ln.source_text
177
+ ][:5],
178
+ refused=[
179
+ {
180
+ "reason": ln.decision.reason.code,
181
+ "src": ln.source_text,
182
+ "llm": ln.proposal.output_text,
183
+ }
184
+ for ln in report.lines
185
+ if ln.decision.reason is not None
186
+ and ln.proposal is not None
187
+ and ln.proposal.output_text != ln.source_text
188
+ ][:5],
189
+ )
190
+
191
+ target.parent.mkdir(parents=True, exist_ok=True)
192
+ target.write_text(json.dumps(record, ensure_ascii=False, indent=1))
193
+ print(
194
+ f"[{arm}] {label:26s} {record['status']:10s} {record.get('lines', 0):5d}L "
195
+ f"{elapsed:6.1f}s {record.get('input_tokens', 0):8d}in/"
196
+ f"{record.get('output_tokens', 0):7d}out",
197
+ flush=True,
198
+ )
199
+ return record
200
+
201
+
202
+ async def main() -> None:
203
+ if len(sys.argv) < 3:
204
+ print(__doc__)
205
+ raise SystemExit(2)
206
+ root = Path(sys.argv[1]).resolve()
207
+ arm = sys.argv[2]
208
+ concurrency = int(sys.argv[3]) if len(sys.argv) > 3 else 3
209
+ if arm not in ARMS:
210
+ raise SystemExit(f"unknown arm {arm!r}; pick from {list(ARMS)}")
211
+
212
+ results = root.parent / "arm-results" / arm
213
+ results.mkdir(parents=True, exist_ok=True)
214
+ pages = find_pages(root)
215
+ print(f"arm {arm} — {len(pages)} pages, concurrency {concurrency}", flush=True)
216
+
217
+ api_key = keychain_key()
218
+ sem = asyncio.Semaphore(concurrency)
219
+ records = await asyncio.gather(
220
+ *(run_one(arm, label, alto, image, api_key, results, sem) for label, alto, image in pages),
221
+ return_exceptions=True,
222
+ )
223
+
224
+ good = [r for r in records if isinstance(r, dict)]
225
+ failed = [r for r in records if isinstance(r, BaseException)]
226
+ totals: Counter[str] = Counter()
227
+ reasons: Counter[str] = Counter()
228
+ for record in good:
229
+ for field in (
230
+ "lines",
231
+ "input_tokens",
232
+ "output_tokens",
233
+ "changed",
234
+ "blocked",
235
+ "fallbacks",
236
+ "retries",
237
+ ):
238
+ totals[field] += record.get(field, 0)
239
+ reasons.update(record.get("reasons", {}))
240
+ print(f"\n=== {arm} ===")
241
+ print(f"pages ok {len(good)}, raised {len(failed)}")
242
+ print(f"totals {dict(totals)}")
243
+ print(f"reasons {dict(reasons)}")
244
+ for exc in failed[:3]:
245
+ print(f" raised: {type(exc).__name__}: {exc}")
246
+ (results.parent / f"{arm}-summary.json").write_text(
247
+ json.dumps(
248
+ {
249
+ "arm": arm,
250
+ "pages": len(good),
251
+ "raised": len(failed),
252
+ "totals": dict(totals),
253
+ "reasons": dict(reasons),
254
+ },
255
+ ensure_ascii=False,
256
+ indent=1,
257
+ )
258
+ )
259
+
260
+
261
+ asyncio.run(main())