Spaces:
Sleeping
La démo devient un poste de revue humaine (#4)
Browse files* Le chemin vision : la démo pouvait recevoir des scans, jamais les envoyer
saknussemm livre une chaîne vision complète — VisionEditProducer qui découpe
chaque ligne, VISION_SYSTEM_PROMPT dont la première phrase est « l'image fait
foi », GuardConfig.vision() qui descend le seuil de similarité de 0,35 à 0,15
parce qu'une lecture CORRECTE d'une ligne massacrée s'éloigne forcément du
texte OCR, et core/batching.py qui découpe un chunk pour respecter
ModelCapabilities.max_images.
Rien de tout ça ne pouvait tourner d'ici. Les quatre fournisseurs
n'implémentent que complete_structured, donc ce backend acceptait les images
de page, les rangeait, les servait à l'interface — et n'en a jamais envoyé
une seule à un modèle.
Ce n'est pas une régression : complete_structured_multimodal n'apparaît
JAMAIS dans l'historique de ce dépôt. La capacité a vécu 23 jours comme
outil en ligne de commande à la racine du dépôt de la bibliothèque
(scripts/run_vision.py + scripts/providers_multimodal.py, du 24 juillet au
16 août), voisin de la démo et non partie d'elle, et elle est partie avec
l'archive du banc — emportée par un défaut prouvé chez sa VOISINE :
vision_benchmark.py construisait son manifeste depuis l'ALTO de référence
puis écrasait le texte par celui de l'OCR. Le client, lui, ne porte pas ce
défaut. Personne n'a relogé la capacité chez un consommateur.
Ce commit la reloge, à l'endroit que la scission désigne.
Quatre pièces, dont trois apprises en se trompant d'abord :
1. app/providers/mistral_multimodal.py — un MultimodalStructuredClient, frère
du fournisseur texte plutôt qu'un élargissement : la bibliothèque garde sa
couture texte sans image exprès. Le transport est base.call_llm, ce qui est
une amélioration sur le client archivé — il retire un paramètre
d'échantillonnage refusé en lisant le message du fournisseur, au lieu d'une
liste de modèles codée en dur qui avait déjà périmé.
2. La branche vision du runner. for_provider ne peut pas servir : il construit
toujours le producteur TEXTE. Le producteur est donc assemblé sur place,
avec max_images pris du fournisseur — c'est lui qui fait découper le chunk
— et GuardConfig.vision().
3. page_image_assets(). Ce backend clé ses images par FICHIER SOURCE, la
bibliothèque les veut par PAGE PHYSIQUE, et require_page_images dit
pourquoi : « flattening them to a single per-file ref sent the producer the
wrong image for every page but the first ». Un fichier à plusieurs pages est
donc REFUSÉ, pas aplati. Une ligne corrigée contre le mauvais scan revient
confiante et fausse, et l'invariant de projection ne peut pas le voir.
4. L'extra [vision] dans les cinq lignes d'installation de la CI. Sans Pillow,
toute tâche vision échoue au premier découpage avec ModuleNotFoundError —
mesuré, pas supposé. Et la note de requirements.txt disait encore
« packages/saknussemm », un chemin disparu le 16 août : personne n'aurait pu
la suivre.
Deux faits viennent du client archivé et sont crédités sur place, parce que
les redécouvrir a coûté cher : Mistral refuse une neuvième image (400, code
3051), et le scan de la page ENTIÈRE est la mauvaise image — attaché à un
chunk avec le prompt texte, le modèle décrit la photographie au lieu de
corriger, une légende revenant réécrite de bout en bout.
tools/run_corpus_arms.py lance un corpus à travers le JobRunner de la démo,
reprenant là où il s'est arrêté ; la clé vient du trousseau macOS et
n'apparaît jamais dans un drapeau, ce qui la mettrait dans l'historique du
shell et dans la liste des processus.
11 tests neufs, 485 verts au total.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw3bbyo2tm2Ujnzu58rLA
* La revue humaine : le verdict dans /layout, et un endroit pour trancher
Ces corpora n'ont AUCUNE vérité terrain. Le text.txt que Gallica livre est
la même couche OCR que son ALTO, donc rien sur le disque ne peut dire si une
correction était juste ou un refus fondé. Un lecteur devant le scan est la
seule source. Deux pièces pour qu'il puisse travailler, et pour que ce qu'il
lit s'accumule au lieu de s'évaporer.
1. /layout porte désormais le VERDICT et la proposition.
La géométrie seule ne dit pas POURQUOI une ligne a gardé son texte OCR.
« Rien n'a été proposé », « une garde a refusé une hallucination » et « un
couple de césure n'a pas pu se réconcilier » sont identiques sur la page et
appellent des jugements différents. Quatre champs par ligne : verdict,
verdict_detail, proposed_text, et proposal_declined — ce dernier étant le cas
le plus intéressant pour un relecteur, quelque chose était sur la table et a
été refusé. Optionnel : un job dont le rapport n'est pas encore écrit rend
toujours sa mise en page.
La garde d'instantané a attrapé le changement de forme, ce qui est le rôle
qu'on lui a donné ; sa liste de clés est mise à jour avec la raison.
2. Un endroit où le jugement se range : PUT/GET /api/jobs/{id}/reviews.
Trois verdicts, et le troisième est le plus utile. `accepted` et `refused`
notent ce que le moteur a fait. `transcribed` porte ce que le lecteur a lu à
l'image, ce qui vaut par soi-même quoi que le moteur ait décidé — et
s'accumule en vérité terrain ligne à ligne. Une transcription vide est
refusée en 422 : le texte EST la revue, et l'accepter remplirait le jeu de
lignes qui n'affirment rien, indiscernables plus tard des vraies.
Idempotent par ligne : renvoyer la même ligne remplace sa revue au lieu de
l'empiler, pour qu'un lecteur qui change d'avis ne se batte pas contre un
journal. La date est estampillée par le serveur, pas reçue du client.
Clé sur la PAIRE (page_id, line_id), séparée par un NUL. ADR-001 : un line_id
se répète d'un fichier à l'autre, donc la clé sur l'identifiant nu
fusionnerait deux jugements le jour où un job porte plus d'un ALTO. Et un NUL
plutôt qu'une espace parce qu'un identifiant peut légalement contenir une
espace — un test pin les deux.
Au passage, deux corrections dans ce qui a été livré hier.
L'ImageAsset était construit SANS transform. Une bibliothèque numérique sert
une dérivée réduite — le IIIF de Gallica rend un JPEG 1193x1600 pour une page
dont l'ALTO mesure 6802x9121, soit un facteur 0,1754. Sans transform le
découpage se fait à l'échelle 1,0, donc une ligne à hpos=4798 est découpée
dans une image large de 1193 : hors du cadre, et chaque découpage revient
vide ou pris ailleurs. Rien en aval ne s'en aperçoit : le VLM reçoit une
image de rien, dit quelque chose de plausible sur le texte OCR qu'on lui
donne aussi, et les gardes jugent une proposition faite sans la preuve
qu'elles croient. Le facteur est maintenant DÉDUIT — la page déclare ses
dimensions, le scan décodé déclare ses pixels — plutôt que demandé à un
appelant qui peut l'ignorer.
Et les fixtures d'image des tests vision étaient des en-têtes fabriqués à la
main. Ça suffisait tant que le code ne reniflait que le nombre magique ; ça a
cessé de suffire quand page_image_assets s'est mis à DÉCODER le scan. Un
fixture qu'on ne peut pas ouvrir échoue alors pour une raison sans rapport
avec ce qui est testé.
492 tests verts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw3bbyo2tm2Ujnzu58rLA
* La vue de revue : colorée par verdict, filtrée en estompant
Le curseur d'opacité existait déjà et pilote tout le calque — il reste tel
quel. Ce qui manquait est ce qu'il y a SOUS le curseur.
Coloré par verdict, en trois familles, parce qu'un relecteur cherche trois
choses différentes et ne devrait pas avoir à lire une légende pour les
distinguer :
retenue la correction a survécu à toutes les gardes
refusée quelque chose a été proposé et une garde l'a décliné — les cas
qui valent un œil humain, chacun étant soit une hallucination
attrapée, soit une bonne correction jetée
sans objet rien n'a été proposé, ou rien n'a changé
Les couleurs sont les deux crayons du correcteur d'épreuve : le bleu marque
ce qui tient, le rouge ce qui a été biffé.
Le filtre ESTOMPE au lieu de masquer, et c'est le point de conception. Une
ligne qui a recopié sa voisine ne se lit comme fausse qu'À CÔTÉ de cette
voisine — retirer le contexte retire la preuve. Éteindre la dernière famille
est refusé : une page blanche se lirait comme un bug.
Une ligne est cliquable et se sélectionne, épaissie ; c'est l'accroche du
panneau de jugement qui vient ensuite.
La garde de couleur du test a attrapé le changement, ce qui est son rôle.
Elle est mise à jour sur ce qu'elle doit vérifier maintenant, et deux tests
s'ajoutent : qu'une ligne refusée se distingue d'une retenue, et que le
filtre estompe sans supprimer — le compte de rectangles reste identique, et
un groupe passe à 0,16.
131 tests frontend verts, typecheck propre.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Kw3bbyo2tm2Ujnzu58rLA
* Le panneau de jugement : la revue humaine devient de la donnée
La vue montrait ce que le moteur avait décidé ; elle ne permettait pas d'en
juger. Un relecteur qui ne peut que regarder produit une impression. Un
relecteur qui peut trancher produit le jeu qui manque à ces corpora :
(ligne, ce que l'OCR a lu, ce que le modèle a proposé, ce qu'un humain dit
que le scan montre). C'est exactement l'entrée dont un banc a besoin pour
répondre « la correction était-elle juste », et qu'aucune mesure interne au
moteur ne peut fabriquer.
Ce que le panneau montre, au clic sur une ligne : le texte OCR, ce qui a été
PROPOSÉ, ce qui a été retenu, et le verdict du moteur. La proposition est
mise en avant quand elle a été refusée — c'est le cas qui vaut un œil
humain, chacun étant soit une hallucination attrapée, soit une bonne
correction jetée, et rien d'autre qu'une lecture du scan ne les sépare.
T
- .github/workflows/ci.yml +5 -5
- backend/app/api/jobs.py +3 -1
- backend/app/api/read_models.py +35 -1
- backend/app/api/review.py +143 -0
- backend/app/jobs/runner.py +176 -16
- backend/app/jobs/store.py +3 -0
- backend/app/main.py +4 -0
- backend/app/protocols/job_store.py +1 -0
- backend/app/providers/mistral_multimodal.py +165 -0
- backend/app/schemas/job.py +6 -0
- backend/requirements.txt +12 -4
- backend/tests/test_read_models_snapshot.py +9 -0
- backend/tests/test_review.py +150 -0
- backend/tests/test_vision_path.py +277 -0
- frontend/openapi.snapshot.json +187 -0
- frontend/src/App.reset-race.test.tsx +5 -0
- frontend/src/App.test.tsx +5 -0
- frontend/src/App.tsx +1 -1
- frontend/src/api/client.ts +41 -1
- frontend/src/components/LayoutViewer.test.tsx +45 -2
- frontend/src/components/LayoutViewer.tsx +216 -14
- frontend/src/components/ReviewPanel.test.tsx +186 -0
- frontend/src/components/ReviewPanel.tsx +215 -0
- frontend/src/lib/iiif.test.ts +66 -0
- frontend/src/lib/iiif.ts +50 -0
- frontend/src/lib/verdicts.ts +45 -0
- frontend/src/types/api.generated.ts +130 -0
- frontend/src/types/index.ts +25 -0
- tools/run_corpus_arms.py +261 -0
|
@@ -52,7 +52,7 @@ jobs:
|
|
| 52 |
python-version: "3.11"
|
| 53 |
cache: pip
|
| 54 |
- name: Install saknussemm (editable, sibling package)
|
| 55 |
-
run: pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
|
| 56 |
- name: Install backend dependencies
|
| 57 |
working-directory: backend
|
| 58 |
run: pip install -r requirements.txt -r requirements-dev.txt
|
|
@@ -75,7 +75,7 @@ jobs:
|
|
| 75 |
python-version: "3.11"
|
| 76 |
cache: pip
|
| 77 |
- name: Install saknussemm (editable, sibling package)
|
| 78 |
-
run: pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
|
| 79 |
- name: Install backend dependencies
|
| 80 |
working-directory: backend
|
| 81 |
run: pip install -r requirements.txt -r requirements-dev.txt
|
|
@@ -112,7 +112,7 @@ jobs:
|
|
| 112 |
python-version: "3.11"
|
| 113 |
cache: pip
|
| 114 |
- name: Install saknussemm (editable, sibling package)
|
| 115 |
-
run: pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
|
| 116 |
- name: Install backend dependencies
|
| 117 |
working-directory: backend
|
| 118 |
run: pip install -r requirements.txt -r requirements-dev.txt
|
|
@@ -135,7 +135,7 @@ jobs:
|
|
| 135 |
# resolve. Audit the resolved environment after install
|
| 136 |
# (pip-audit -r doesn't understand editable lines).
|
| 137 |
run: |
|
| 138 |
-
pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
|
| 139 |
pip install -r backend/requirements.txt
|
| 140 |
pip install bandit[toml]==1.9.4 pip-audit==2.10.0
|
| 141 |
- name: bandit (static analysis)
|
|
@@ -217,7 +217,7 @@ jobs:
|
|
| 217 |
cache: npm
|
| 218 |
cache-dependency-path: frontend/package-lock.json
|
| 219 |
- name: Install saknussemm (editable, sibling package)
|
| 220 |
-
run: pip install 'saknussemm @ git+https://github.com/maribakulj/saknussemm@main'
|
| 221 |
- name: Install backend dependencies
|
| 222 |
working-directory: backend
|
| 223 |
run: pip install -r requirements.txt -r requirements-dev.txt
|
|
|
|
| 52 |
python-version: "3.11"
|
| 53 |
cache: pip
|
| 54 |
- name: Install saknussemm (editable, sibling package)
|
| 55 |
+
run: pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
|
| 56 |
- name: Install backend dependencies
|
| 57 |
working-directory: backend
|
| 58 |
run: pip install -r requirements.txt -r requirements-dev.txt
|
|
|
|
| 75 |
python-version: "3.11"
|
| 76 |
cache: pip
|
| 77 |
- name: Install saknussemm (editable, sibling package)
|
| 78 |
+
run: pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
|
| 79 |
- name: Install backend dependencies
|
| 80 |
working-directory: backend
|
| 81 |
run: pip install -r requirements.txt -r requirements-dev.txt
|
|
|
|
| 112 |
python-version: "3.11"
|
| 113 |
cache: pip
|
| 114 |
- name: Install saknussemm (editable, sibling package)
|
| 115 |
+
run: pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
|
| 116 |
- name: Install backend dependencies
|
| 117 |
working-directory: backend
|
| 118 |
run: pip install -r requirements.txt -r requirements-dev.txt
|
|
|
|
| 135 |
# resolve. Audit the resolved environment after install
|
| 136 |
# (pip-audit -r doesn't understand editable lines).
|
| 137 |
run: |
|
| 138 |
+
pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
|
| 139 |
pip install -r backend/requirements.txt
|
| 140 |
pip install bandit[toml]==1.9.4 pip-audit==2.10.0
|
| 141 |
- name: bandit (static analysis)
|
|
|
|
| 217 |
cache: npm
|
| 218 |
cache-dependency-path: frontend/package-lock.json
|
| 219 |
- name: Install saknussemm (editable, sibling package)
|
| 220 |
+
run: pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
|
| 221 |
- name: Install backend dependencies
|
| 222 |
working-directory: backend
|
| 223 |
run: pip install -r requirements.txt -r requirements-dev.txt
|
|
@@ -778,7 +778,9 @@ async def get_job_layout(job: JobManifest = Depends(get_completed_job)) -> dict:
|
|
| 778 |
if job.document_manifest is None:
|
| 779 |
raise HTTPException(status_code=500, detail="Job has no document_manifest.")
|
| 780 |
# Wave-3 review — same offload rationale as /diff.
|
| 781 |
-
data = await asyncio.to_thread(
|
|
|
|
|
|
|
| 782 |
# Plan V2.4 — <img> cannot set headers: append a short-lived,
|
| 783 |
# images-scoped signed credential to each URL (this response is
|
| 784 |
# itself token-gated, so only the job owner receives them).
|
|
|
|
| 778 |
if job.document_manifest is None:
|
| 779 |
raise HTTPException(status_code=500, detail="Job has no document_manifest.")
|
| 780 |
# Wave-3 review — same offload rationale as /diff.
|
| 781 |
+
data = await asyncio.to_thread(
|
| 782 |
+
build_layout, job.job_id, job.document_manifest, job.images, job.report
|
| 783 |
+
)
|
| 784 |
# Plan V2.4 — <img> cannot set headers: append a short-lived,
|
| 785 |
# images-scoped signed credential to each URL (this response is
|
| 786 |
# itself token-gated, so only the job owner receives them).
|
|
@@ -11,7 +11,7 @@ job, guard the HTTP preconditions, and delegate here.
|
|
| 11 |
|
| 12 |
from __future__ import annotations
|
| 13 |
|
| 14 |
-
from app.schemas import DocumentManifest, HyphenRole
|
| 15 |
|
| 16 |
|
| 17 |
def build_diff(job_id: str, document_manifest: DocumentManifest) -> dict:
|
|
@@ -69,13 +69,38 @@ def build_layout(
|
|
| 69 |
job_id: str,
|
| 70 |
document_manifest: DocumentManifest,
|
| 71 |
images: dict[str, str],
|
|
|
|
| 72 |
) -> dict:
|
| 73 |
"""Structural layout: blocks + lines with ALTO coordinates, per page.
|
| 74 |
|
| 75 |
Page dimensions are derived from line coordinates when the source Page
|
| 76 |
element omits WIDTH/HEIGHT. ``images`` maps source_file → image filename;
|
| 77 |
a matching entry becomes the page's ``image_url``.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
"""
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 79 |
pages_out = []
|
| 80 |
for page in document_manifest.pages:
|
| 81 |
line_by_id = {lm.line_id: lm for lm in page.lines}
|
|
@@ -99,6 +124,15 @@ def build_layout(
|
|
| 99 |
"corrected_text": corrected,
|
| 100 |
"modified": corrected != lm.ocr_text,
|
| 101 |
"hyphen_role": lm.hyphen_role.value,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 102 |
}
|
| 103 |
)
|
| 104 |
blocks_out.append(
|
|
|
|
| 11 |
|
| 12 |
from __future__ import annotations
|
| 13 |
|
| 14 |
+
from app.schemas import CorrectionReport, DocumentManifest, HyphenRole
|
| 15 |
|
| 16 |
|
| 17 |
def build_diff(job_id: str, document_manifest: DocumentManifest) -> dict:
|
|
|
|
| 69 |
job_id: str,
|
| 70 |
document_manifest: DocumentManifest,
|
| 71 |
images: dict[str, str],
|
| 72 |
+
report: CorrectionReport | None = None,
|
| 73 |
) -> dict:
|
| 74 |
"""Structural layout: blocks + lines with ALTO coordinates, per page.
|
| 75 |
|
| 76 |
Page dimensions are derived from line coordinates when the source Page
|
| 77 |
element omits WIDTH/HEIGHT. ``images`` maps source_file → image filename;
|
| 78 |
a matching entry becomes the page's ``image_url``.
|
| 79 |
+
|
| 80 |
+
``report`` adds, per line, **what the engine decided and why**: the verdict
|
| 81 |
+
code and the text the producer actually proposed. Without it a reviewer
|
| 82 |
+
sees that a line was left alone but not whether that was because nothing
|
| 83 |
+
was proposed, because a guard refused a hallucination, or because a hyphen
|
| 84 |
+
pair could not reconcile — three situations that look identical in the
|
| 85 |
+
geometry and demand different judgements. Optional so a job whose report
|
| 86 |
+
is not yet written still renders its layout.
|
| 87 |
"""
|
| 88 |
+
verdicts: dict[tuple[str, str], dict] = {}
|
| 89 |
+
if report is not None:
|
| 90 |
+
for trace in report.lines:
|
| 91 |
+
reason = trace.decision.reason
|
| 92 |
+
proposed = trace.proposal.output_text if trace.proposal else None
|
| 93 |
+
verdicts[(trace.page_id, trace.line_id)] = {
|
| 94 |
+
"verdict": reason.code if reason is not None else trace.decision.status,
|
| 95 |
+
"verdict_detail": reason.detail if reason is not None else None,
|
| 96 |
+
"proposed_text": proposed,
|
| 97 |
+
# A proposal the engine declined is the reviewer's most
|
| 98 |
+
# interesting case: something was on offer and was refused.
|
| 99 |
+
"proposal_declined": bool(
|
| 100 |
+
reason is not None and proposed is not None and proposed != trace.source_text
|
| 101 |
+
),
|
| 102 |
+
}
|
| 103 |
+
|
| 104 |
pages_out = []
|
| 105 |
for page in document_manifest.pages:
|
| 106 |
line_by_id = {lm.line_id: lm for lm in page.lines}
|
|
|
|
| 124 |
"corrected_text": corrected,
|
| 125 |
"modified": corrected != lm.ocr_text,
|
| 126 |
"hyphen_role": lm.hyphen_role.value,
|
| 127 |
+
**verdicts.get(
|
| 128 |
+
(page.page_id, lm.line_id),
|
| 129 |
+
{
|
| 130 |
+
"verdict": None,
|
| 131 |
+
"verdict_detail": None,
|
| 132 |
+
"proposed_text": None,
|
| 133 |
+
"proposal_declined": False,
|
| 134 |
+
},
|
| 135 |
+
),
|
| 136 |
}
|
| 137 |
)
|
| 138 |
blocks_out.append(
|
|
@@ -0,0 +1,143 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Human review — where a reader's judgement on a line is recorded.
|
| 2 |
+
|
| 3 |
+
The corpora this demo runs on have **no ground truth**. Gallica's `text.txt`
|
| 4 |
+
is the same OCR layer as its ALTO, so nothing on disk can say whether a
|
| 5 |
+
correction was right or a refusal was justified; the only source of truth is
|
| 6 |
+
a person reading the scan. This module is where that reading is kept.
|
| 7 |
+
|
| 8 |
+
**Why it is worth keeping rather than just displaying.** A reviewer who can
|
| 9 |
+
only look produces an impression. A reviewer who can adjudicate produces the
|
| 10 |
+
dataset that is missing: `(line, what the OCR read, what the model proposed,
|
| 11 |
+
what a human says the scan actually shows)`. That is exactly the input a
|
| 12 |
+
bench needs to answer "was the correction right", which no amount of
|
| 13 |
+
measurement inside the engine can answer on its own.
|
| 14 |
+
|
| 15 |
+
**Three verdicts, and the third is the useful one.** `accepted` and
|
| 16 |
+
`refused` grade what the engine did. `transcribed` carries what the reviewer
|
| 17 |
+
read on the image, which stands on its own whatever the engine decided — and
|
| 18 |
+
accumulates into ground truth line by line.
|
| 19 |
+
|
| 20 |
+
Reviews are keyed by ``(page_id, line_id)`` because a line id repeats across
|
| 21 |
+
files (`ADR-001`); keying on the bare id would silently merge two documents'
|
| 22 |
+
judgements the first time a job carries more than one ALTO.
|
| 23 |
+
"""
|
| 24 |
+
|
| 25 |
+
from __future__ import annotations
|
| 26 |
+
|
| 27 |
+
from datetime import UTC, datetime
|
| 28 |
+
from enum import StrEnum
|
| 29 |
+
|
| 30 |
+
from fastapi import APIRouter, Depends, HTTPException, status
|
| 31 |
+
from pydantic import BaseModel, Field
|
| 32 |
+
|
| 33 |
+
from app.api.deps import get_job_store
|
| 34 |
+
from app.api.jobs import require_job_access
|
| 35 |
+
from app.protocols import JobStore
|
| 36 |
+
from app.schemas import JobManifest
|
| 37 |
+
|
| 38 |
+
router = APIRouter(prefix="/api/jobs", tags=["review"])
|
| 39 |
+
|
| 40 |
+
|
| 41 |
+
class ReviewVerdict(StrEnum):
|
| 42 |
+
"""What the reader concluded about the engine's decision on this line."""
|
| 43 |
+
|
| 44 |
+
#: The engine's outcome is right — whether it corrected or refused.
|
| 45 |
+
ACCEPTED = "accepted"
|
| 46 |
+
#: The engine's outcome is wrong. `note` should say how.
|
| 47 |
+
REFUSED = "refused"
|
| 48 |
+
#: Neither: the reader is recording what the scan actually shows.
|
| 49 |
+
TRANSCRIBED = "transcribed"
|
| 50 |
+
|
| 51 |
+
|
| 52 |
+
class LineReview(BaseModel):
|
| 53 |
+
"""One reader's judgement on one line."""
|
| 54 |
+
|
| 55 |
+
page_id: str = Field(min_length=1, max_length=256)
|
| 56 |
+
line_id: str = Field(min_length=1, max_length=256)
|
| 57 |
+
verdict: ReviewVerdict
|
| 58 |
+
#: What the reader read on the image. Required for ``transcribed``; free
|
| 59 |
+
#: to accompany the other two when the reader wants to be precise about
|
| 60 |
+
#: what the engine got wrong.
|
| 61 |
+
transcription: str | None = Field(default=None, max_length=4000)
|
| 62 |
+
note: str | None = Field(default=None, max_length=2000)
|
| 63 |
+
reviewed_at: str | None = None
|
| 64 |
+
|
| 65 |
+
@property
|
| 66 |
+
def key(self) -> tuple[str, str]:
|
| 67 |
+
return (self.page_id, self.line_id)
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
class ReviewBatch(BaseModel):
|
| 71 |
+
"""Reviews arrive in batches: a reader works through a page, not a line."""
|
| 72 |
+
|
| 73 |
+
reviews: list[LineReview] = Field(max_length=2000)
|
| 74 |
+
|
| 75 |
+
|
| 76 |
+
class ReviewsResponse(BaseModel):
|
| 77 |
+
job_id: str
|
| 78 |
+
reviews: list[LineReview]
|
| 79 |
+
|
| 80 |
+
|
| 81 |
+
def _key(review: LineReview) -> str:
|
| 82 |
+
"""``(page_id, line_id)`` flattened for storage, NUL-separated.
|
| 83 |
+
|
| 84 |
+
A NUL cannot occur in an XML id, so the pair round-trips unambiguously; a
|
| 85 |
+
space could, and would merge two lines the day a producer emits one.
|
| 86 |
+
"""
|
| 87 |
+
return f"{review.page_id}\x00{review.line_id}"
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
def _existing(job: JobManifest) -> dict[str, LineReview]:
|
| 91 |
+
"""Reviews already on the job, back as models.
|
| 92 |
+
|
| 93 |
+
The manifest stores plain dicts so the schema layer never imports the API
|
| 94 |
+
layer's models; re-validating here is what keeps that separation from
|
| 95 |
+
costing type safety at the edge.
|
| 96 |
+
"""
|
| 97 |
+
return {key: LineReview.model_validate(raw) for key, raw in (job.reviews or {}).items()}
|
| 98 |
+
|
| 99 |
+
|
| 100 |
+
@router.put("/{job_id}/reviews", response_model=ReviewsResponse)
|
| 101 |
+
async def put_reviews(
|
| 102 |
+
job_id: str,
|
| 103 |
+
batch: ReviewBatch,
|
| 104 |
+
job: JobManifest = Depends(require_job_access),
|
| 105 |
+
store: JobStore = Depends(get_job_store),
|
| 106 |
+
) -> ReviewsResponse:
|
| 107 |
+
"""Record or replace judgements on lines of this job.
|
| 108 |
+
|
| 109 |
+
Idempotent per line: sending the same line twice replaces its review
|
| 110 |
+
rather than appending, so a reader who changes their mind is not fighting
|
| 111 |
+
an append-only log. The timestamp is stamped here rather than trusted
|
| 112 |
+
from the client — a review's date is a fact about the server.
|
| 113 |
+
"""
|
| 114 |
+
for review in batch.reviews:
|
| 115 |
+
if review.verdict is ReviewVerdict.TRANSCRIBED and not review.transcription:
|
| 116 |
+
raise HTTPException(
|
| 117 |
+
status.HTTP_422_UNPROCESSABLE_CONTENT,
|
| 118 |
+
f"line {review.line_id!r}: a 'transcribed' review must carry the "
|
| 119 |
+
"text the reader read on the image — that text IS the review.",
|
| 120 |
+
)
|
| 121 |
+
|
| 122 |
+
merged = _existing(job)
|
| 123 |
+
stamped = datetime.now(UTC).isoformat(timespec="seconds")
|
| 124 |
+
for review in batch.reviews:
|
| 125 |
+
merged[f"{_key(review)}"] = review.model_copy(update={"reviewed_at": stamped})
|
| 126 |
+
store.update_job(job_id, reviews={k: v.model_dump() for k, v in merged.items()})
|
| 127 |
+
return ReviewsResponse(
|
| 128 |
+
job_id=job_id, reviews=sorted(merged.values(), key=lambda r: (r.page_id, r.line_id))
|
| 129 |
+
)
|
| 130 |
+
|
| 131 |
+
|
| 132 |
+
@router.get("/{job_id}/reviews", response_model=ReviewsResponse)
|
| 133 |
+
async def get_reviews(
|
| 134 |
+
job_id: str,
|
| 135 |
+
job: JobManifest = Depends(require_job_access),
|
| 136 |
+
) -> ReviewsResponse:
|
| 137 |
+
return ReviewsResponse(
|
| 138 |
+
job_id=job_id,
|
| 139 |
+
reviews=sorted(_existing(job).values(), key=lambda r: (r.page_id, r.line_id)),
|
| 140 |
+
)
|
| 141 |
+
|
| 142 |
+
|
| 143 |
+
__all__ = ["LineReview", "ReviewBatch", "ReviewVerdict", "router"]
|
|
@@ -26,7 +26,17 @@ from saknussemm import (
|
|
| 26 |
)
|
| 27 |
from saknussemm.core.events import ReconcileStats
|
| 28 |
from saknussemm.core.protocols import ProviderPermanentError
|
| 29 |
-
from saknussemm.core.schemas import
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 30 |
|
| 31 |
from app.jobs.events import JobEventType
|
| 32 |
from app.jobs.observers import CompositeObserver, JobStoreObserver, LoggingObserver
|
|
@@ -36,6 +46,92 @@ from app.schemas import DocumentManifest, JobStatus
|
|
| 36 |
logger = logging.getLogger(__name__)
|
| 37 |
|
| 38 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
def _default_timeout_from_env() -> int:
|
| 40 |
"""Resolve JOB_TIMEOUT_SECONDS once at import. 0 disables the timeout.
|
| 41 |
|
|
@@ -91,6 +187,7 @@ class JobRunner:
|
|
| 91 |
source_files: dict[str, Path],
|
| 92 |
provider: BaseProvider | None = None,
|
| 93 |
pairing_policy: PairingPolicy | None = None,
|
|
|
|
| 94 |
timeout_seconds: int = 1800,
|
| 95 |
should_abort: Callable[[], bool] | None = None,
|
| 96 |
) -> None:
|
|
@@ -110,6 +207,10 @@ class JobRunner:
|
|
| 110 |
`should_abort`: Plan V2.2 — cooperative cancellation probe,
|
| 111 |
forwarded to the pipeline (polled between pages and chunks).
|
| 112 |
When it trips, the job lands in CANCELLED with no output promoted.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 113 |
"""
|
| 114 |
if provider is None:
|
| 115 |
from app.providers import get_provider
|
|
@@ -132,6 +233,7 @@ class JobRunner:
|
|
| 132 |
output_writer=output_writer,
|
| 133 |
source_files=source_files,
|
| 134 |
pairing_policy=pairing_policy,
|
|
|
|
| 135 |
should_abort=should_abort,
|
| 136 |
),
|
| 137 |
timeout=timeout,
|
|
@@ -318,9 +420,14 @@ class JobRunner:
|
|
| 318 |
output_writer: OutputWriter,
|
| 319 |
source_files: dict[str, Path],
|
| 320 |
pairing_policy: PairingPolicy | None = None,
|
|
|
|
| 321 |
should_abort: Callable[[], bool] | None = None,
|
| 322 |
) -> CorrectionResult:
|
| 323 |
-
"""Drive the pure pipeline and persist its counters back.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 324 |
self.job_store.update_job(job_id, status=JobStatus.STARTED)
|
| 325 |
self.job_store.emit(job_id, JobEventType.STARTED, {"job_id": job_id})
|
| 326 |
|
|
@@ -338,25 +445,78 @@ class JobRunner:
|
|
| 338 |
# §5.1 resorption — credentials go into the producer (via the
|
| 339 |
# for_provider convenience), never into run(): the pipeline surface
|
| 340 |
# carries no api_key anywhere.
|
| 341 |
-
|
| 342 |
-
|
| 343 |
-
|
| 344 |
-
|
| 345 |
-
|
| 346 |
-
|
| 347 |
-
|
| 348 |
-
|
| 349 |
-
|
| 350 |
-
|
| 351 |
-
|
| 352 |
-
#
|
| 353 |
-
|
| 354 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 355 |
# `run_id` is saknussemm's generic identifier; we feed it the
|
| 356 |
# server-side `job_id` so trace.json correlates with the API.
|
| 357 |
result = await pipeline.run(
|
| 358 |
document_manifest=document_manifest,
|
| 359 |
source_files=source_files,
|
|
|
|
| 360 |
run_id=job_id,
|
| 361 |
# Plan V2.2 — the cancel endpoint's event, polled by the
|
| 362 |
# pipeline between pages and chunks.
|
|
|
|
| 26 |
)
|
| 27 |
from saknussemm.core.events import ReconcileStats
|
| 28 |
from saknussemm.core.protocols import ProviderPermanentError
|
| 29 |
+
from saknussemm.core.schemas import (
|
| 30 |
+
GuardConfig,
|
| 31 |
+
ImageAsset,
|
| 32 |
+
ImageTransform,
|
| 33 |
+
ModelCapabilities,
|
| 34 |
+
PageImage,
|
| 35 |
+
PageManifest,
|
| 36 |
+
PairingPolicy,
|
| 37 |
+
)
|
| 38 |
+
from saknussemm.errors import ConfigurationError
|
| 39 |
+
from saknussemm.integrations.vision import build_image_asset
|
| 40 |
|
| 41 |
from app.jobs.events import JobEventType
|
| 42 |
from app.jobs.observers import CompositeObserver, JobStoreObserver, LoggingObserver
|
|
|
|
| 46 |
logger = logging.getLogger(__name__)
|
| 47 |
|
| 48 |
|
| 49 |
+
def page_image_assets(
|
| 50 |
+
document_manifest: DocumentManifest, images_by_source: dict[str, Path]
|
| 51 |
+
) -> dict[str, ImageAsset]:
|
| 52 |
+
"""``{page_id: ImageAsset}`` from the demo's per-SOURCE-FILE image map.
|
| 53 |
+
|
| 54 |
+
The two models disagree and the disagreement matters. This backend stores
|
| 55 |
+
one image per uploaded source file, keyed by its stem; the library wants one
|
| 56 |
+
per *physical page*, and ``require_page_images`` says why in as many words:
|
| 57 |
+
"never one per source file: a multipage XML has as many scans as pages, and
|
| 58 |
+
flattening them to a single per-file ref sent the producer the wrong image
|
| 59 |
+
for every page but the first".
|
| 60 |
+
|
| 61 |
+
So a source file carrying more than one page is **refused**, not flattened.
|
| 62 |
+
A wrong scan produces a confident correction of a line that is not on it,
|
| 63 |
+
and nothing downstream can see that — the projection invariant compares the
|
| 64 |
+
artefact to the decisions the artefact was built from.
|
| 65 |
+
"""
|
| 66 |
+
assets: dict[str, ImageAsset] = {}
|
| 67 |
+
pages_per_source: dict[str, int] = {}
|
| 68 |
+
for page in document_manifest.pages:
|
| 69 |
+
pages_per_source[page.source_file] = pages_per_source.get(page.source_file, 0) + 1
|
| 70 |
+
|
| 71 |
+
for page in document_manifest.pages:
|
| 72 |
+
stem = Path(page.source_file).stem.lower()
|
| 73 |
+
image_path = images_by_source.get(stem)
|
| 74 |
+
if image_path is None:
|
| 75 |
+
continue
|
| 76 |
+
if pages_per_source[page.source_file] > 1:
|
| 77 |
+
raise ValueError(
|
| 78 |
+
f"{page.source_file!r} carries {pages_per_source[page.source_file]} "
|
| 79 |
+
"pages but only one image was uploaded for it. A vision run needs "
|
| 80 |
+
"one scan per physical page; sending the same image for every page "
|
| 81 |
+
"would correct each line against the wrong scan. Upload one file "
|
| 82 |
+
"per page, or run this document without vision."
|
| 83 |
+
)
|
| 84 |
+
asset = build_image_asset(page_id=page.page_id, path=image_path)
|
| 85 |
+
assets[page.page_id] = asset.model_copy(update={"transform": _transform_for(page, asset)})
|
| 86 |
+
return assets
|
| 87 |
+
|
| 88 |
+
|
| 89 |
+
def _transform_for(page: PageManifest, asset: ImageAsset) -> ImageTransform | None:
|
| 90 |
+
"""Map the ALTO/PAGE coordinate space onto the scan's pixels.
|
| 91 |
+
|
| 92 |
+
**Not optional, and the reason is a measured near-miss.** A digital
|
| 93 |
+
library serves a downscaled derivative — Gallica's IIIF ``!1600,1600``
|
| 94 |
+
gives a 1193x1600 JPEG for a page whose ALTO measures 6802x9121, a factor
|
| 95 |
+
of **0.1754**. With no transform the crop is taken at scale 1.0, so a line
|
| 96 |
+
at ``hpos=4798`` is cropped from an image 1193 pixels wide: off the
|
| 97 |
+
canvas, and every crop comes back blank or from the wrong place.
|
| 98 |
+
|
| 99 |
+
Nothing downstream notices. The VLM is handed a picture of nothing, says
|
| 100 |
+
something plausible about the OCR text it was also given, and the guards
|
| 101 |
+
judge a proposal made without the evidence they think it was made with.
|
| 102 |
+
|
| 103 |
+
Derived rather than configured: the page declares its own dimensions and
|
| 104 |
+
the decoded scan declares its pixels, so the ratio is available without
|
| 105 |
+
asking a caller for a DPI it may not know. Returns ``None`` when either
|
| 106 |
+
side is missing — a wrong scale is worse than a declared unknown.
|
| 107 |
+
"""
|
| 108 |
+
if not (asset.pixel_width and asset.pixel_height):
|
| 109 |
+
return None
|
| 110 |
+
if not (page.page_width and page.page_height):
|
| 111 |
+
return None
|
| 112 |
+
return ImageTransform(
|
| 113 |
+
scale_x=asset.pixel_width / page.page_width,
|
| 114 |
+
scale_y=asset.pixel_height / page.page_height,
|
| 115 |
+
)
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
def _media_type_for(path: Path) -> str:
|
| 119 |
+
"""Decoded from the bytes, not guessed from the extension.
|
| 120 |
+
|
| 121 |
+
A ``.jpg`` that is really a PNG makes the vendor reject the data URI, and
|
| 122 |
+
the schema field documents itself as "determined from the bytes rather than
|
| 123 |
+
guessed from the extension".
|
| 124 |
+
"""
|
| 125 |
+
header = path.read_bytes()[:12]
|
| 126 |
+
if header.startswith(b"\x89PNG\r\n\x1a\n"):
|
| 127 |
+
return "image/png"
|
| 128 |
+
if header[:2] == b"\xff\xd8":
|
| 129 |
+
return "image/jpeg"
|
| 130 |
+
if header[:4] in (b"II*\x00", b"MM\x00*"):
|
| 131 |
+
return "image/tiff"
|
| 132 |
+
return "application/octet-stream"
|
| 133 |
+
|
| 134 |
+
|
| 135 |
def _default_timeout_from_env() -> int:
|
| 136 |
"""Resolve JOB_TIMEOUT_SECONDS once at import. 0 disables the timeout.
|
| 137 |
|
|
|
|
| 187 |
source_files: dict[str, Path],
|
| 188 |
provider: BaseProvider | None = None,
|
| 189 |
pairing_policy: PairingPolicy | None = None,
|
| 190 |
+
page_images: dict[str, PageImage] | None = None,
|
| 191 |
timeout_seconds: int = 1800,
|
| 192 |
should_abort: Callable[[], bool] | None = None,
|
| 193 |
) -> None:
|
|
|
|
| 207 |
`should_abort`: Plan V2.2 — cooperative cancellation probe,
|
| 208 |
forwarded to the pipeline (polled between pages and chunks).
|
| 209 |
When it trips, the job lands in CANCELLED with no output promoted.
|
| 210 |
+
`page_images`: page_id → scan, from ``page_image_assets()``. Present
|
| 211 |
+
means a VISION run: the producer becomes ``VisionEditProducer`` and the
|
| 212 |
+
guards become ``GuardConfig.vision()``. Absent (the default) keeps the
|
| 213 |
+
historical text path byte for byte.
|
| 214 |
"""
|
| 215 |
if provider is None:
|
| 216 |
from app.providers import get_provider
|
|
|
|
| 233 |
output_writer=output_writer,
|
| 234 |
source_files=source_files,
|
| 235 |
pairing_policy=pairing_policy,
|
| 236 |
+
page_images=page_images,
|
| 237 |
should_abort=should_abort,
|
| 238 |
),
|
| 239 |
timeout=timeout,
|
|
|
|
| 420 |
output_writer: OutputWriter,
|
| 421 |
source_files: dict[str, Path],
|
| 422 |
pairing_policy: PairingPolicy | None = None,
|
| 423 |
+
page_images: dict[str, PageImage] | None = None,
|
| 424 |
should_abort: Callable[[], bool] | None = None,
|
| 425 |
) -> CorrectionResult:
|
| 426 |
+
"""Drive the pure pipeline and persist its counters back.
|
| 427 |
+
|
| 428 |
+
`page_images`: page_id → scan. Present means a VISION run — the
|
| 429 |
+
producer changes, and so does the guard config.
|
| 430 |
+
"""
|
| 431 |
self.job_store.update_job(job_id, status=JobStatus.STARTED)
|
| 432 |
self.job_store.emit(job_id, JobEventType.STARTED, {"job_id": job_id})
|
| 433 |
|
|
|
|
| 445 |
# §5.1 resorption — credentials go into the producer (via the
|
| 446 |
# for_provider convenience), never into run(): the pipeline surface
|
| 447 |
# carries no api_key anywhere.
|
| 448 |
+
observer = CompositeObserver([JobStoreObserver(self.job_store, job_id), LoggingObserver()])
|
| 449 |
+
if page_images:
|
| 450 |
+
# Vision run. `for_provider` cannot serve it: it always builds the
|
| 451 |
+
# TEXT producer, around a client whose seam carries no images on
|
| 452 |
+
# purpose. So the producer is assembled here, with the three things
|
| 453 |
+
# a VLM run needs and a text run must not have.
|
| 454 |
+
from saknussemm.integrations.vision import (
|
| 455 |
+
MultimodalStructuredClient,
|
| 456 |
+
VisionEditProducer,
|
| 457 |
+
)
|
| 458 |
+
|
| 459 |
+
# Checked, not assumed. A text provider here would fail at the
|
| 460 |
+
# first call, deep inside a run, with a message about a missing
|
| 461 |
+
# method rather than about the job being misconfigured — and the
|
| 462 |
+
# run would already have spent its first chunk getting there.
|
| 463 |
+
if not isinstance(provider, MultimodalStructuredClient):
|
| 464 |
+
raise ConfigurationError(
|
| 465 |
+
f"a vision run needs a multimodal provider, and "
|
| 466 |
+
f"{type(provider).__name__} implements the text seam only. "
|
| 467 |
+
"Pass MistralMultimodalProvider, or run this job without "
|
| 468 |
+
"page images."
|
| 469 |
+
)
|
| 470 |
+
max_images = getattr(provider, "MAX_IMAGES_PER_CALL", None)
|
| 471 |
+
producer = VisionEditProducer(
|
| 472 |
+
provider=provider,
|
| 473 |
+
api_key=api_key,
|
| 474 |
+
model=model,
|
| 475 |
+
# `max_images` is what makes core/batching.py split a chunk.
|
| 476 |
+
# Without it the engine sends every line's crop in one call and
|
| 477 |
+
# the vendor refuses the request outright.
|
| 478 |
+
capabilities=ModelCapabilities(
|
| 479 |
+
text=True,
|
| 480 |
+
vision=True,
|
| 481 |
+
structured_output=True,
|
| 482 |
+
max_images=max_images,
|
| 483 |
+
),
|
| 484 |
+
)
|
| 485 |
+
pipeline = CorrectionPipeline(
|
| 486 |
+
producer=producer,
|
| 487 |
+
observer=observer,
|
| 488 |
+
pairing_policy=pairing_policy,
|
| 489 |
+
# A VLM reads the image, so a CORRECT reading of a badly
|
| 490 |
+
# garbled line diverges from the OCR text further than the
|
| 491 |
+
# text-tuned guard tolerates (0.35 → 0.15). Running vision
|
| 492 |
+
# under the text config refuses the model's best work: measured
|
| 493 |
+
# 18 refusals out of 20 lines, every one
|
| 494 |
+
# `too_different_from_source`.
|
| 495 |
+
guard_config=GuardConfig.vision(),
|
| 496 |
+
)
|
| 497 |
+
else:
|
| 498 |
+
# §5.1 resorption — credentials go into the producer (via the
|
| 499 |
+
# for_provider convenience), never into run(): the pipeline surface
|
| 500 |
+
# carries no api_key anywhere.
|
| 501 |
+
pipeline = CorrectionPipeline.for_provider(
|
| 502 |
+
provider,
|
| 503 |
+
api_key=api_key,
|
| 504 |
+
model=model,
|
| 505 |
+
provider_name=provider_name,
|
| 506 |
+
observer=observer,
|
| 507 |
+
# Provenance parity — the pipeline hashes this into the §11
|
| 508 |
+
# config fingerprint. Passing the job's actual policy (default
|
| 509 |
+
# or geometric_pairing opt-out) keeps the stamped fingerprint
|
| 510 |
+
# honest; omitting it silently reverted to
|
| 511 |
+
# DEFAULT_PAIRING_POLICY.
|
| 512 |
+
pairing_policy=pairing_policy,
|
| 513 |
+
)
|
| 514 |
# `run_id` is saknussemm's generic identifier; we feed it the
|
| 515 |
# server-side `job_id` so trace.json correlates with the API.
|
| 516 |
result = await pipeline.run(
|
| 517 |
document_manifest=document_manifest,
|
| 518 |
source_files=source_files,
|
| 519 |
+
page_images=page_images,
|
| 520 |
run_id=job_id,
|
| 521 |
# Plan V2.2 — the cancel endpoint's event, polled by the
|
| 522 |
# pipeline between pages and chunks.
|
|
@@ -139,6 +139,7 @@ class JobStore:
|
|
| 139 |
fallbacks: int | None = None,
|
| 140 |
duration_seconds: float | None = None,
|
| 141 |
error: str | None = None,
|
|
|
|
| 142 |
images: dict[str, str] | None = None,
|
| 143 |
report: CorrectionReport | None = None,
|
| 144 |
token_hash: str | None = None,
|
|
@@ -172,6 +173,8 @@ class JobStore:
|
|
| 172 |
job.duration_seconds = duration_seconds
|
| 173 |
if error is not None:
|
| 174 |
job.error = error
|
|
|
|
|
|
|
| 175 |
if images is not None:
|
| 176 |
job.images = images
|
| 177 |
if report is not None:
|
|
|
|
| 139 |
fallbacks: int | None = None,
|
| 140 |
duration_seconds: float | None = None,
|
| 141 |
error: str | None = None,
|
| 142 |
+
reviews: dict[str, dict] | None = None,
|
| 143 |
images: dict[str, str] | None = None,
|
| 144 |
report: CorrectionReport | None = None,
|
| 145 |
token_hash: str | None = None,
|
|
|
|
| 173 |
job.duration_seconds = duration_seconds
|
| 174 |
if error is not None:
|
| 175 |
job.error = error
|
| 176 |
+
if reviews is not None:
|
| 177 |
+
job.reviews = reviews
|
| 178 |
if images is not None:
|
| 179 |
job.images = images
|
| 180 |
if report is not None:
|
|
@@ -24,6 +24,7 @@ from app.api.health import router as health_router
|
|
| 24 |
from app.api.jobs import router as jobs_router
|
| 25 |
from app.api.providers import router as providers_router
|
| 26 |
from app.api.rate_limit import limiter
|
|
|
|
| 27 |
from app.api.upload_guard import UploadAdmissionMiddleware, UploadSizeLimitMiddleware
|
| 28 |
from app.frontend_static import INDEX_HTML as _INDEX_HTML
|
| 29 |
from app.frontend_static import STATIC_DIR as _STATIC_DIR
|
|
@@ -278,6 +279,9 @@ def create_app() -> FastAPI:
|
|
| 278 |
# ------------------------------------------------------------------
|
| 279 |
app.include_router(providers_router, prefix="/api/providers", tags=["providers"])
|
| 280 |
app.include_router(jobs_router, prefix="/api/jobs", tags=["jobs"])
|
|
|
|
|
|
|
|
|
|
| 281 |
|
| 282 |
# ------------------------------------------------------------------
|
| 283 |
# Static frontend (HF Spaces single-container mode)
|
|
|
|
| 24 |
from app.api.jobs import router as jobs_router
|
| 25 |
from app.api.providers import router as providers_router
|
| 26 |
from app.api.rate_limit import limiter
|
| 27 |
+
from app.api.review import router as review_router
|
| 28 |
from app.api.upload_guard import UploadAdmissionMiddleware, UploadSizeLimitMiddleware
|
| 29 |
from app.frontend_static import INDEX_HTML as _INDEX_HTML
|
| 30 |
from app.frontend_static import STATIC_DIR as _STATIC_DIR
|
|
|
|
| 279 |
# ------------------------------------------------------------------
|
| 280 |
app.include_router(providers_router, prefix="/api/providers", tags=["providers"])
|
| 281 |
app.include_router(jobs_router, prefix="/api/jobs", tags=["jobs"])
|
| 282 |
+
# Human review sits on the same job paths and carries its own prefix,
|
| 283 |
+
# so it registers without one here.
|
| 284 |
+
app.include_router(review_router)
|
| 285 |
|
| 286 |
# ------------------------------------------------------------------
|
| 287 |
# Static frontend (HF Spaces single-container mode)
|
|
@@ -43,6 +43,7 @@ class JobStore(Protocol):
|
|
| 43 |
error: str | None = None,
|
| 44 |
images: dict[str, str] | None = None,
|
| 45 |
report: Any | None = None,
|
|
|
|
| 46 |
token_hash: str | None = None,
|
| 47 |
) -> None: ...
|
| 48 |
|
|
|
|
| 43 |
error: str | None = None,
|
| 44 |
images: dict[str, str] | None = None,
|
| 45 |
report: Any | None = None,
|
| 46 |
+
reviews: dict[str, Any] | None = None,
|
| 47 |
token_hash: str | None = None,
|
| 48 |
) -> None: ...
|
| 49 |
|
|
@@ -0,0 +1,165 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Mistral multimodal provider — the half of the library the demo could not run.
|
| 2 |
+
|
| 3 |
+
``saknussemm`` ships a complete vision chain: ``VisionEditProducer`` crops each
|
| 4 |
+
line from the page scan, ``VISION_SYSTEM_PROMPT`` tells the model the image is
|
| 5 |
+
authoritative, ``GuardConfig.vision()`` loosens the similarity guard because a
|
| 6 |
+
correct reading of a garbled line necessarily diverges from the OCR text, and
|
| 7 |
+
``core/batching.py`` splits a chunk to respect ``ModelCapabilities.max_images``.
|
| 8 |
+
|
| 9 |
+
None of it could run from here. Every bundled provider implements
|
| 10 |
+
``complete_structured`` only, so the backend accepted page images, stored them,
|
| 11 |
+
served them to the UI — and never sent one to a model.
|
| 12 |
+
|
| 13 |
+
**This is not a regression.** ``complete_structured_multimodal`` has never
|
| 14 |
+
appeared in this repository's history. The capability lived for 23 days as a CLI
|
| 15 |
+
tool at the root of the *library* repo (``scripts/run_vision.py`` +
|
| 16 |
+
``scripts/providers_multimodal.py``, 2026-07-24 to 2026-08-16), a sibling of the
|
| 17 |
+
demo rather than a part of it, and left with the bench archive — swept along by a
|
| 18 |
+
defect proven in its *neighbour*: ``vision_benchmark.py`` built its manifest from
|
| 19 |
+
the reference ALTO and then overwrote the text with the OCR's, a state no real
|
| 20 |
+
run reaches. The client itself carries no such defect. Nothing re-homed the
|
| 21 |
+
capability in a consumer. This does.
|
| 22 |
+
|
| 23 |
+
**Credit where it is due.** Two facts below come from that archived client
|
| 24 |
+
(``cinoc/campaigns/tooling/providers_multimodal.py``), which measured them
|
| 25 |
+
against the live API rather than reading them off a doc page. Both are
|
| 26 |
+
load-bearing and both were rediscovered the hard way before the archive was
|
| 27 |
+
read:
|
| 28 |
+
|
| 29 |
+
* **A ninth image is refused**, HTTP 400 ``"Total number of images exceeds the
|
| 30 |
+
maximum allowed of 8."`` (code 3051). So :data:`MAX_IMAGES_PER_CALL` is 8 and
|
| 31 |
+
is declared through ``ModelCapabilities.max_images``, which is what makes the
|
| 32 |
+
engine split a chunk instead of issuing a request that cannot succeed.
|
| 33 |
+
* **The whole-page scan is the wrong image.** Attaching one page to a chunk and
|
| 34 |
+
asking for OCR corrections makes the model *describe* the photograph: a
|
| 35 |
+
caption came back rewritten from `PHOTOGRAPHIE DE L'ÉBOULEMENT D'UNE FALAISE A
|
| 36 |
+
BOULOGNE-SUR-MER` to `DEUX ASPECTS DE LA FALAISE ÉBOULÉE, MONTRANT LE
|
| 37 |
+
GLISSEMENT DES TERRES`. Per-line crops are not an optimisation; they are what
|
| 38 |
+
makes the task legible to the model.
|
| 39 |
+
|
| 40 |
+
The transport is this backend's ``base.call_llm`` rather than the archived
|
| 41 |
+
client's own httpx block, and that is an upgrade rather than a shortcut: it
|
| 42 |
+
strips a rejected sampling parameter by reading the vendor's own message,
|
| 43 |
+
instead of consulting a hardcoded list of models that refuse ``temperature``,
|
| 44 |
+
and the archived list had already gone stale once.
|
| 45 |
+
"""
|
| 46 |
+
|
| 47 |
+
from __future__ import annotations
|
| 48 |
+
|
| 49 |
+
import base64
|
| 50 |
+
import json
|
| 51 |
+
from typing import Any
|
| 52 |
+
|
| 53 |
+
from app.providers.base import call_llm, extract_chat_text, extract_usage, get_json
|
| 54 |
+
from app.schemas import ModelInfo, Usage
|
| 55 |
+
|
| 56 |
+
_BASE = "https://api.mistral.ai"
|
| 57 |
+
|
| 58 |
+
|
| 59 |
+
class MistralMultimodalProvider:
|
| 60 |
+
"""A ``MultimodalStructuredClient`` for Mistral's chat-completions API.
|
| 61 |
+
|
| 62 |
+
Deliberately a sibling of :class:`~app.providers.mistral_provider.MistralProvider`
|
| 63 |
+
rather than a widening of it: the library keeps the text seam image-free on
|
| 64 |
+
purpose, and mixing the two would make every text provider carry a vision
|
| 65 |
+
signature it cannot honour.
|
| 66 |
+
"""
|
| 67 |
+
|
| 68 |
+
#: Measured, not documented: a ninth image is refused with 400/3051.
|
| 69 |
+
#: Declared to the engine through ``ModelCapabilities.max_images`` so the
|
| 70 |
+
#: chunk is split upstream of the request.
|
| 71 |
+
MAX_IMAGES_PER_CALL = 8
|
| 72 |
+
|
| 73 |
+
def _headers(self, api_key: str) -> dict[str, str]:
|
| 74 |
+
return {
|
| 75 |
+
"Authorization": f"Bearer {api_key}",
|
| 76 |
+
"Content-Type": "application/json",
|
| 77 |
+
}
|
| 78 |
+
|
| 79 |
+
async def list_models(self, api_key: str) -> list[ModelInfo]:
|
| 80 |
+
"""Only the models whose live capabilities include vision.
|
| 81 |
+
|
| 82 |
+
A text-only id here would 400 on the first image, and the catalogue
|
| 83 |
+
moves — so the filter reads the account's own answer instead of a
|
| 84 |
+
hardcoded list.
|
| 85 |
+
"""
|
| 86 |
+
data = await get_json(url=f"{_BASE}/v1/models", headers=self._headers(api_key))
|
| 87 |
+
models = []
|
| 88 |
+
for entry in data.get("data", []):
|
| 89 |
+
caps = entry.get("capabilities", {})
|
| 90 |
+
if not caps.get("completion_chat", False) or not caps.get("vision", False):
|
| 91 |
+
continue
|
| 92 |
+
model_id = entry.get("id", "")
|
| 93 |
+
models.append(ModelInfo(id=model_id, label=entry.get("name") or model_id))
|
| 94 |
+
models.sort(key=lambda m: m.id)
|
| 95 |
+
return models
|
| 96 |
+
|
| 97 |
+
def _content_blocks(
|
| 98 |
+
self, user_payload: dict[str, Any], images: list[Any]
|
| 99 |
+
) -> list[dict[str, Any]]:
|
| 100 |
+
"""Each crop announced with the line it depicts, then the JSON payload.
|
| 101 |
+
|
| 102 |
+
The reply is keyed by ``line_id`` and the model has no other way to
|
| 103 |
+
pair a picture with a line, so the label is part of the contract rather
|
| 104 |
+
than decoration.
|
| 105 |
+
"""
|
| 106 |
+
blocks: list[dict[str, Any]] = []
|
| 107 |
+
for part in images:
|
| 108 |
+
blocks.append({"type": "text", "text": f"Image de la ligne {part.line_id} :"})
|
| 109 |
+
blocks.append(
|
| 110 |
+
{
|
| 111 |
+
"type": "image_url",
|
| 112 |
+
"image_url": (
|
| 113 |
+
f"data:{part.media_type};base64,"
|
| 114 |
+
f"{base64.standard_b64encode(part.data).decode('ascii')}"
|
| 115 |
+
),
|
| 116 |
+
}
|
| 117 |
+
)
|
| 118 |
+
blocks.append({"type": "text", "text": json.dumps(user_payload, ensure_ascii=False)})
|
| 119 |
+
return blocks
|
| 120 |
+
|
| 121 |
+
async def complete_structured_multimodal(
|
| 122 |
+
self,
|
| 123 |
+
*,
|
| 124 |
+
api_key: str,
|
| 125 |
+
model: str,
|
| 126 |
+
system_prompt: str,
|
| 127 |
+
user_payload: dict[str, Any],
|
| 128 |
+
images: list[Any],
|
| 129 |
+
json_schema: dict[str, Any],
|
| 130 |
+
temperature: float = 0.0,
|
| 131 |
+
) -> tuple[dict[str, Any], Usage | None]:
|
| 132 |
+
if len(images) > self.MAX_IMAGES_PER_CALL:
|
| 133 |
+
# Reaching this means the engine was not told the limit: the
|
| 134 |
+
# capability descriptor is what splits the chunk, and a request
|
| 135 |
+
# issued anyway would 400 after the crops were already encoded.
|
| 136 |
+
raise ValueError(
|
| 137 |
+
f"{len(images)} crops in one call but Mistral accepts at most "
|
| 138 |
+
f"{self.MAX_IMAGES_PER_CALL}. Declare "
|
| 139 |
+
f"ModelCapabilities(max_images={self.MAX_IMAGES_PER_CALL}) on the "
|
| 140 |
+
"producer so the engine splits the chunk upstream."
|
| 141 |
+
)
|
| 142 |
+
|
| 143 |
+
body: dict[str, Any] = {
|
| 144 |
+
"model": model,
|
| 145 |
+
"temperature": temperature,
|
| 146 |
+
"messages": [
|
| 147 |
+
{"role": "system", "content": system_prompt},
|
| 148 |
+
{
|
| 149 |
+
"role": "user",
|
| 150 |
+
"content": self._content_blocks(user_payload, images),
|
| 151 |
+
},
|
| 152 |
+
],
|
| 153 |
+
"response_format": {"type": "json_schema", "json_schema": json_schema},
|
| 154 |
+
}
|
| 155 |
+
# Same two-step the text provider relies on: some models reject the
|
| 156 |
+
# strict schema form and accept plain json_object.
|
| 157 |
+
fallback_body = {**body, "response_format": {"type": "json_object"}}
|
| 158 |
+
|
| 159 |
+
data = await call_llm(
|
| 160 |
+
url=f"{_BASE}/v1/chat/completions",
|
| 161 |
+
headers=self._headers(api_key),
|
| 162 |
+
body=body,
|
| 163 |
+
fallback_body=fallback_body,
|
| 164 |
+
)
|
| 165 |
+
return extract_chat_text(data, "Mistral (vision)"), extract_usage(data)
|
|
@@ -89,6 +89,12 @@ class JobManifest(BaseModel):
|
|
| 89 |
# job keeps (the former parallel ``line_traces`` dict was redundant —
|
| 90 |
# nothing read it — and is gone).
|
| 91 |
report: CorrectionReport | None = None
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 92 |
|
| 93 |
|
| 94 |
__all__ = ["TERMINAL_SUCCESS_STATES", "JobManifest", "JobStatus", "Provider"]
|
|
|
|
| 89 |
# job keeps (the former parallel ``line_traces`` dict was redundant —
|
| 90 |
# nothing read it — and is gone).
|
| 91 |
report: CorrectionReport | None = None
|
| 92 |
+
#: Human review, keyed ``"<page_id> <line_id>"`` — see `app.api.review`.
|
| 93 |
+
#: Keyed on the PAIR because a line id repeats across files (`ADR-001`);
|
| 94 |
+
#: the bare id would merge two documents' judgements the first time a job
|
| 95 |
+
#: carries more than one ALTO. Untyped here so the schema layer stays
|
| 96 |
+
#: free of the API layer's models.
|
| 97 |
+
reviews: dict[str, dict] = Field(default_factory=dict)
|
| 98 |
|
| 99 |
|
| 100 |
__all__ = ["TERMINAL_SUCCESS_STATES", "JobManifest", "JobStatus", "Provider"]
|
|
@@ -13,13 +13,21 @@
|
|
| 13 |
# -o requirements-lock.txt requirements.in
|
| 14 |
# 3. pytest
|
| 15 |
#
|
| 16 |
-
# NOTE: saknussemm is
|
| 17 |
-
#
|
| 18 |
-
#
|
|
|
|
|
|
|
| 19 |
#
|
| 20 |
-
# pip install
|
| 21 |
# pip install -r backend/requirements.txt -r backend/requirements-dev.txt
|
| 22 |
#
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
# CI follows this two-step install. The Dockerfiles do too. See
|
| 24 |
# CONTRIBUTING.md for the local-dev recipe.
|
| 25 |
fastapi==0.139.0
|
|
|
|
| 13 |
# -o requirements-lock.txt requirements.in
|
| 14 |
# 3. pytest
|
| 15 |
#
|
| 16 |
+
# NOTE: saknussemm is NOT listed here — it is installed separately, and it
|
| 17 |
+
# must be installed WITH ITS VISION EXTRA. `packages/saknussemm` has not
|
| 18 |
+
# existed since the library left this repository on 2026-08-16 and its tree was
|
| 19 |
+
# flattened; the instruction below said otherwise for long enough that nobody
|
| 20 |
+
# could have followed it.
|
| 21 |
#
|
| 22 |
+
# pip install 'saknussemm[vision] @ git+https://github.com/maribakulj/saknussemm@main'
|
| 23 |
# pip install -r backend/requirements.txt -r backend/requirements-dev.txt
|
| 24 |
#
|
| 25 |
+
# The `[vision]` extra is Pillow, and it is not optional for this backend:
|
| 26 |
+
# `VisionEditProducer` crops each line from the page scan, and without Pillow
|
| 27 |
+
# every vision job fails at the first crop with ModuleNotFoundError — measured,
|
| 28 |
+
# not supposed. A text-only install is a backend that silently cannot run half
|
| 29 |
+
# of what it offers.
|
| 30 |
+
#
|
| 31 |
# CI follows this two-step install. The Dockerfiles do too. See
|
| 32 |
# CONTRIBUTING.md for the local-dev recipe.
|
| 33 |
fastapi==0.139.0
|
|
@@ -60,6 +60,15 @@ _LAYOUT_LINE_KEYS = {
|
|
| 60 |
"corrected_text",
|
| 61 |
"modified",
|
| 62 |
"hyphen_role",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
}
|
| 64 |
|
| 65 |
|
|
|
|
| 60 |
"corrected_text",
|
| 61 |
"modified",
|
| 62 |
"hyphen_role",
|
| 63 |
+
# Added for human review: geometry alone cannot tell a reviewer WHY a line
|
| 64 |
+
# kept its OCR text. "Nothing was proposed", "a guard refused a
|
| 65 |
+
# hallucination" and "a hyphen pair could not reconcile" look identical on
|
| 66 |
+
# the page and demand different judgements. `None` when the job has no
|
| 67 |
+
# report yet — the layout still renders.
|
| 68 |
+
"verdict",
|
| 69 |
+
"verdict_detail",
|
| 70 |
+
"proposed_text",
|
| 71 |
+
"proposal_declined",
|
| 72 |
}
|
| 73 |
|
| 74 |
|
|
@@ -0,0 +1,150 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Human review — the only place ground truth can come from.
|
| 2 |
+
|
| 3 |
+
These corpora have none. Gallica ships a `text.txt` that is the *same OCR
|
| 4 |
+
layer* as its ALTO, so nothing on disk can say whether a correction was right
|
| 5 |
+
or a refusal justified. A reader looking at the scan is the only source, and
|
| 6 |
+
this is where that reading is kept so it accumulates instead of evaporating.
|
| 7 |
+
"""
|
| 8 |
+
|
| 9 |
+
from __future__ import annotations
|
| 10 |
+
|
| 11 |
+
import pytest
|
| 12 |
+
from fastapi.testclient import TestClient
|
| 13 |
+
|
| 14 |
+
from app.api.review import LineReview, ReviewVerdict
|
| 15 |
+
from app.jobs.store import JobStore
|
| 16 |
+
from app.schemas import Provider
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
@pytest.fixture
|
| 20 |
+
def client(tmp_path, monkeypatch):
|
| 21 |
+
"""A TestClient on an isolated storage dir — the demo's usual shape."""
|
| 22 |
+
from app import storage as storage_module
|
| 23 |
+
from app.main import create_app
|
| 24 |
+
|
| 25 |
+
monkeypatch.setattr(storage_module, "_BASE_DIR", tmp_path / "jobs")
|
| 26 |
+
with TestClient(create_app(), raise_server_exceptions=False) as test_client:
|
| 27 |
+
yield test_client
|
| 28 |
+
|
| 29 |
+
|
| 30 |
+
@pytest.fixture
|
| 31 |
+
def job(client: TestClient) -> str:
|
| 32 |
+
"""A store-created job: no token hash, so the capability gate stays open.
|
| 33 |
+
|
| 34 |
+
Jobs created outside the HTTP layer are ungated by design (`P1-7`), which
|
| 35 |
+
is what lets a test exercise the route without minting a token.
|
| 36 |
+
"""
|
| 37 |
+
store: JobStore = client.app.state.job_store
|
| 38 |
+
return store.create_job(provider=Provider.MISTRAL, model="m")
|
| 39 |
+
|
| 40 |
+
|
| 41 |
+
def _put(client: TestClient, job_id: str, reviews: list[dict]):
|
| 42 |
+
return client.put(f"/api/jobs/{job_id}/reviews", json={"reviews": reviews})
|
| 43 |
+
|
| 44 |
+
|
| 45 |
+
def test_a_review_survives_and_comes_back(client: TestClient, job: str) -> None:
|
| 46 |
+
response = _put(
|
| 47 |
+
client,
|
| 48 |
+
job,
|
| 49 |
+
[
|
| 50 |
+
{
|
| 51 |
+
"page_id": "PAG_1",
|
| 52 |
+
"line_id": "TL000323",
|
| 53 |
+
"verdict": "refused",
|
| 54 |
+
"note": "l'OCR lit ETANTS, le scan dit ENFANTS",
|
| 55 |
+
}
|
| 56 |
+
],
|
| 57 |
+
)
|
| 58 |
+
assert response.status_code == 200, response.text
|
| 59 |
+
|
| 60 |
+
fetched = client.get(f"/api/jobs/{job}/reviews").json()["reviews"]
|
| 61 |
+
assert len(fetched) == 1
|
| 62 |
+
assert fetched[0]["verdict"] == "refused"
|
| 63 |
+
assert "ENFANTS" in fetched[0]["note"]
|
| 64 |
+
assert fetched[0]["reviewed_at"], "the server stamps the date, not the client"
|
| 65 |
+
|
| 66 |
+
|
| 67 |
+
def test_reviewing_a_line_twice_replaces_rather_than_appends(client: TestClient, job: str) -> None:
|
| 68 |
+
"""A reader who changes their mind must not be fighting an append log."""
|
| 69 |
+
line = {"page_id": "PAG_1", "line_id": "TL1", "verdict": "accepted"}
|
| 70 |
+
_put(client, job, [line])
|
| 71 |
+
_put(client, job, [{**line, "verdict": "refused", "note": "en y regardant mieux"}])
|
| 72 |
+
|
| 73 |
+
reviews = client.get(f"/api/jobs/{job}/reviews").json()["reviews"]
|
| 74 |
+
assert len(reviews) == 1, reviews
|
| 75 |
+
assert reviews[0]["verdict"] == "refused"
|
| 76 |
+
|
| 77 |
+
|
| 78 |
+
def test_two_files_sharing_a_line_id_stay_separate(client: TestClient, job: str) -> None:
|
| 79 |
+
"""`ADR-001` — a line id repeats across files; only the PAIR is unique.
|
| 80 |
+
|
| 81 |
+
Keyed on the bare id, these two judgements would merge the first time a
|
| 82 |
+
job carries more than one ALTO, and the second reader would silently
|
| 83 |
+
overwrite the first.
|
| 84 |
+
"""
|
| 85 |
+
_put(
|
| 86 |
+
client,
|
| 87 |
+
job,
|
| 88 |
+
[
|
| 89 |
+
{"page_id": "PAG_1", "line_id": "TL1", "verdict": "accepted"},
|
| 90 |
+
{"page_id": "PAG_2", "line_id": "TL1", "verdict": "refused"},
|
| 91 |
+
],
|
| 92 |
+
)
|
| 93 |
+
reviews = client.get(f"/api/jobs/{job}/reviews").json()["reviews"]
|
| 94 |
+
assert len(reviews) == 2, reviews
|
| 95 |
+
assert {r["page_id"] for r in reviews} == {"PAG_1", "PAG_2"}
|
| 96 |
+
|
| 97 |
+
|
| 98 |
+
def test_a_transcription_without_text_is_refused(client: TestClient, job: str) -> None:
|
| 99 |
+
"""The text IS the review, so a `transcribed` verdict cannot be empty.
|
| 100 |
+
|
| 101 |
+
Accepting it would fill the ground-truth set with rows that assert
|
| 102 |
+
nothing, and they would be indistinguishable later from real ones.
|
| 103 |
+
"""
|
| 104 |
+
response = _put(
|
| 105 |
+
client,
|
| 106 |
+
job,
|
| 107 |
+
[{"page_id": "PAG_1", "line_id": "TL1", "verdict": "transcribed"}],
|
| 108 |
+
)
|
| 109 |
+
assert response.status_code == 422
|
| 110 |
+
assert "text the reader read" in response.text
|
| 111 |
+
|
| 112 |
+
|
| 113 |
+
def test_a_transcription_is_what_the_reader_read(client: TestClient, job: str) -> None:
|
| 114 |
+
"""The verdict that produces ground truth rather than grading the engine."""
|
| 115 |
+
_put(
|
| 116 |
+
client,
|
| 117 |
+
job,
|
| 118 |
+
[
|
| 119 |
+
{
|
| 120 |
+
"page_id": "PAG_1",
|
| 121 |
+
"line_id": "TL000323",
|
| 122 |
+
"verdict": "transcribed",
|
| 123 |
+
"transcription": "1,500 ENFANTS : LE NOMBRE DES MORTS S'ÉLÈVE",
|
| 124 |
+
}
|
| 125 |
+
],
|
| 126 |
+
)
|
| 127 |
+
review = client.get(f"/api/jobs/{job}/reviews").json()["reviews"][0]
|
| 128 |
+
assert review["transcription"].startswith("1,500 ENFANTS")
|
| 129 |
+
|
| 130 |
+
|
| 131 |
+
def test_an_unknown_job_is_not_reviewable(client: TestClient) -> None:
|
| 132 |
+
response = _put(client, "nope", [{"page_id": "P", "line_id": "L", "verdict": "accepted"}])
|
| 133 |
+
assert response.status_code == 404
|
| 134 |
+
|
| 135 |
+
|
| 136 |
+
def test_the_key_separator_cannot_occur_in_an_id() -> None:
|
| 137 |
+
"""Why the composite key joins on NUL rather than a space.
|
| 138 |
+
|
| 139 |
+
An XML id may legally contain neither, but a producer emitting an id with
|
| 140 |
+
a space is a real possibility and would merge two lines; a NUL is not
|
| 141 |
+
representable in XML at all, so the pair round-trips unambiguously.
|
| 142 |
+
"""
|
| 143 |
+
from app.api.review import _key
|
| 144 |
+
|
| 145 |
+
left = LineReview(page_id="A B", line_id="C", verdict=ReviewVerdict.ACCEPTED)
|
| 146 |
+
right = LineReview(page_id="A", line_id="B C", verdict=ReviewVerdict.ACCEPTED)
|
| 147 |
+
assert _key(left) != _key(right), (
|
| 148 |
+
"two different (page, line) pairs collapsed to one key — a space "
|
| 149 |
+
"separator would do exactly this"
|
| 150 |
+
)
|
|
@@ -0,0 +1,277 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""The vision path — the half of the library this backend could not run.
|
| 2 |
+
|
| 3 |
+
Until now every bundled provider implemented ``complete_structured`` only. The
|
| 4 |
+
backend accepted page images, stored them, served them to the UI, and never
|
| 5 |
+
sent one to a model, so ``saknussemm``'s ``VisionEditProducer`` — crops per
|
| 6 |
+
line, image-is-authoritative prompt, loosened guards, chunk splitting on
|
| 7 |
+
``max_images`` — had no consumer at all.
|
| 8 |
+
|
| 9 |
+
These tests pin the four things that make the difference between a vision run
|
| 10 |
+
and a request that cannot succeed. Each of the four was learned by getting it
|
| 11 |
+
wrong first, which is why each has its own case rather than being folded into
|
| 12 |
+
one happy path.
|
| 13 |
+
"""
|
| 14 |
+
|
| 15 |
+
from __future__ import annotations
|
| 16 |
+
|
| 17 |
+
from pathlib import Path
|
| 18 |
+
from typing import Any
|
| 19 |
+
|
| 20 |
+
import pytest
|
| 21 |
+
from saknussemm.core.schemas import GuardConfig
|
| 22 |
+
from saknussemm.integrations.vision import ImagePart, MultimodalStructuredClient
|
| 23 |
+
|
| 24 |
+
from app.jobs.runner import _media_type_for, page_image_assets
|
| 25 |
+
from app.providers.mistral_multimodal import MistralMultimodalProvider
|
| 26 |
+
from app.schemas import DocumentManifest
|
| 27 |
+
|
| 28 |
+
|
| 29 |
+
def _encoded(fmt: str, size: tuple[int, int] = (40, 12)) -> bytes:
|
| 30 |
+
"""A real, decodable image.
|
| 31 |
+
|
| 32 |
+
Hand-rolled header bytes were enough while the code only sniffed the magic
|
| 33 |
+
number; they stopped being enough when ``page_image_assets`` began
|
| 34 |
+
DECODING the scan to derive the XML-to-pixel scale. A fixture that cannot
|
| 35 |
+
be opened then fails for a reason that has nothing to do with what is
|
| 36 |
+
under test.
|
| 37 |
+
"""
|
| 38 |
+
import io
|
| 39 |
+
|
| 40 |
+
from PIL import Image
|
| 41 |
+
|
| 42 |
+
buffer = io.BytesIO()
|
| 43 |
+
Image.new("RGB", size, (255, 255, 255)).save(buffer, format=fmt)
|
| 44 |
+
return buffer.getvalue()
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
_JPEG = _encoded("JPEG")
|
| 48 |
+
_PNG = _encoded("PNG")
|
| 49 |
+
_TIFF = _encoded("TIFF")
|
| 50 |
+
|
| 51 |
+
|
| 52 |
+
def test_the_provider_satisfies_the_librarys_multimodal_protocol() -> None:
|
| 53 |
+
"""The seam is a Protocol, so conformance is checkable rather than hoped for.
|
| 54 |
+
|
| 55 |
+
``VisionEditProducer`` takes a ``MultimodalStructuredClient``; a provider
|
| 56 |
+
that merely looks similar fails at the first call, deep inside a run.
|
| 57 |
+
"""
|
| 58 |
+
assert isinstance(MistralMultimodalProvider(), MultimodalStructuredClient)
|
| 59 |
+
|
| 60 |
+
|
| 61 |
+
def test_the_image_cap_is_declared_and_refused_loudly() -> None:
|
| 62 |
+
"""Eight is measured, not documented, and the ninth image 400s.
|
| 63 |
+
|
| 64 |
+
The cap exists to be handed to the engine through
|
| 65 |
+
``ModelCapabilities.max_images`` — ``core/batching.py`` then splits the
|
| 66 |
+
chunk. This assertion is the backstop for when it is *not* handed over:
|
| 67 |
+
refusing before encoding is better than a 400 after, and the message must
|
| 68 |
+
name the fix rather than the symptom.
|
| 69 |
+
"""
|
| 70 |
+
provider = MistralMultimodalProvider()
|
| 71 |
+
assert provider.MAX_IMAGES_PER_CALL == 8
|
| 72 |
+
|
| 73 |
+
crops = [
|
| 74 |
+
ImagePart(line_id=f"L{i}", media_type="image/jpeg", data=_JPEG, sha256="x")
|
| 75 |
+
for i in range(9)
|
| 76 |
+
]
|
| 77 |
+
with pytest.raises(ValueError, match="max_images"):
|
| 78 |
+
import asyncio
|
| 79 |
+
|
| 80 |
+
asyncio.run(
|
| 81 |
+
provider.complete_structured_multimodal(
|
| 82 |
+
api_key="k",
|
| 83 |
+
model="mistral-small-latest",
|
| 84 |
+
system_prompt="s",
|
| 85 |
+
user_payload={"lines": []},
|
| 86 |
+
images=crops,
|
| 87 |
+
json_schema={},
|
| 88 |
+
)
|
| 89 |
+
)
|
| 90 |
+
|
| 91 |
+
|
| 92 |
+
def test_each_crop_is_labelled_with_the_line_it_depicts() -> None:
|
| 93 |
+
"""The reply is keyed by ``line_id`` and the model has no other pairing.
|
| 94 |
+
|
| 95 |
+
An unlabelled pile of crops is how a VLM ends up captioning: it cannot tell
|
| 96 |
+
which picture belongs to which line, so it answers about the picture.
|
| 97 |
+
"""
|
| 98 |
+
provider = MistralMultimodalProvider()
|
| 99 |
+
crops = [
|
| 100 |
+
ImagePart(line_id="L1", media_type="image/jpeg", data=_JPEG, sha256="a"),
|
| 101 |
+
ImagePart(line_id="L2", media_type="image/png", data=_PNG, sha256="b"),
|
| 102 |
+
]
|
| 103 |
+
blocks = provider._content_blocks({"lines": [{"line_id": "L1"}]}, crops)
|
| 104 |
+
|
| 105 |
+
labels = [b["text"] for b in blocks if b["type"] == "text"]
|
| 106 |
+
assert any("L1" in label for label in labels), labels
|
| 107 |
+
assert any("L2" in label for label in labels), labels
|
| 108 |
+
images = [b for b in blocks if b["type"] == "image_url"]
|
| 109 |
+
assert len(images) == 2
|
| 110 |
+
assert images[0]["image_url"].startswith("data:image/jpeg;base64,")
|
| 111 |
+
assert images[1]["image_url"].startswith("data:image/png;base64,")
|
| 112 |
+
# The payload rides last, after the crops it describes.
|
| 113 |
+
assert blocks[-1]["type"] == "text"
|
| 114 |
+
assert "line_id" in blocks[-1]["text"]
|
| 115 |
+
|
| 116 |
+
|
| 117 |
+
@pytest.mark.parametrize(
|
| 118 |
+
("raw", "expected"),
|
| 119 |
+
[(_JPEG, "image/jpeg"), (_PNG, "image/png"), (_TIFF, "image/tiff")],
|
| 120 |
+
)
|
| 121 |
+
def test_the_media_type_is_read_from_the_bytes(tmp_path: Path, raw: bytes, expected: str) -> None:
|
| 122 |
+
"""A ``.jpg`` that is really a PNG makes the vendor reject the data URI.
|
| 123 |
+
|
| 124 |
+
``ImageAsset.media_type`` documents itself as "determined from the bytes
|
| 125 |
+
rather than guessed from the extension", so every file here is named
|
| 126 |
+
``.jpg`` regardless of what it contains.
|
| 127 |
+
"""
|
| 128 |
+
path = tmp_path / "scan.jpg"
|
| 129 |
+
path.write_bytes(raw)
|
| 130 |
+
assert _media_type_for(path) == expected
|
| 131 |
+
|
| 132 |
+
|
| 133 |
+
def _manifest_with(pages_per_file: dict[str, int]) -> DocumentManifest:
|
| 134 |
+
"""A manifest with the given number of pages per source file."""
|
| 135 |
+
from saknussemm.core.schemas import PageManifest
|
| 136 |
+
|
| 137 |
+
pages = []
|
| 138 |
+
for source, count in pages_per_file.items():
|
| 139 |
+
for index in range(count):
|
| 140 |
+
pages.append(
|
| 141 |
+
PageManifest(
|
| 142 |
+
page_id=f"{Path(source).stem}_P{index}",
|
| 143 |
+
source_file=source,
|
| 144 |
+
page_index=index,
|
| 145 |
+
page_width=100,
|
| 146 |
+
page_height=100,
|
| 147 |
+
blocks=[],
|
| 148 |
+
lines=[],
|
| 149 |
+
)
|
| 150 |
+
)
|
| 151 |
+
return DocumentManifest(pages=pages, source_files=sorted(pages_per_file))
|
| 152 |
+
|
| 153 |
+
|
| 154 |
+
def test_one_page_per_file_maps_cleanly(tmp_path: Path) -> None:
|
| 155 |
+
scan = tmp_path / "a.jpg"
|
| 156 |
+
scan.write_bytes(_JPEG)
|
| 157 |
+
manifest = _manifest_with({"a.xml": 1})
|
| 158 |
+
|
| 159 |
+
assets = page_image_assets(manifest, {"a": scan})
|
| 160 |
+
|
| 161 |
+
assert list(assets) == ["a_P0"]
|
| 162 |
+
asset = assets["a_P0"]
|
| 163 |
+
assert asset.page_id == "a_P0"
|
| 164 |
+
assert asset.media_type == "image/jpeg"
|
| 165 |
+
assert asset.sha256 and len(asset.sha256) == 64
|
| 166 |
+
|
| 167 |
+
|
| 168 |
+
def test_a_multipage_file_is_refused_rather_than_flattened(tmp_path: Path) -> None:
|
| 169 |
+
"""The one failure mode nothing downstream could ever see.
|
| 170 |
+
|
| 171 |
+
This backend keys images by SOURCE FILE; the library wants one per physical
|
| 172 |
+
page and ``require_page_images`` spells out why — "flattening them to a
|
| 173 |
+
single per-file ref sent the producer the wrong image for every page but
|
| 174 |
+
the first". A line corrected against the wrong scan comes back confident
|
| 175 |
+
and wrong, and the projection invariant cannot notice: it compares the
|
| 176 |
+
artefact to the decisions the artefact was built from.
|
| 177 |
+
"""
|
| 178 |
+
scan = tmp_path / "vol.jpg"
|
| 179 |
+
scan.write_bytes(_JPEG)
|
| 180 |
+
manifest = _manifest_with({"vol.xml": 3})
|
| 181 |
+
|
| 182 |
+
with pytest.raises(ValueError, match="one scan per physical page"):
|
| 183 |
+
page_image_assets(manifest, {"vol": scan})
|
| 184 |
+
|
| 185 |
+
|
| 186 |
+
def test_a_file_without_a_scan_is_simply_absent(tmp_path: Path) -> None:
|
| 187 |
+
"""Not every uploaded XML has a matching image, and that is not an error.
|
| 188 |
+
|
| 189 |
+
``require_page_images`` refuses a vision run whose coverage is incomplete,
|
| 190 |
+
so the decision belongs there — this helper reports what exists.
|
| 191 |
+
"""
|
| 192 |
+
manifest = _manifest_with({"a.xml": 1, "b.xml": 1})
|
| 193 |
+
scan = tmp_path / "a.jpg"
|
| 194 |
+
scan.write_bytes(_JPEG)
|
| 195 |
+
|
| 196 |
+
assets = page_image_assets(manifest, {"a": scan})
|
| 197 |
+
assert list(assets) == ["a_P0"]
|
| 198 |
+
|
| 199 |
+
|
| 200 |
+
def test_the_vision_guard_config_is_looser_than_the_text_one() -> None:
|
| 201 |
+
"""Why the branch changes the guards and not only the producer.
|
| 202 |
+
|
| 203 |
+
A VLM reads the image, so a CORRECT reading of a badly garbled line
|
| 204 |
+
diverges from the OCR text further than the text-tuned threshold tolerates.
|
| 205 |
+
Measured once with the text config by mistake: 18 refusals out of 20 lines,
|
| 206 |
+
every one ``too_different_from_source``.
|
| 207 |
+
"""
|
| 208 |
+
from saknussemm.core.guards import DEFAULT_GUARD_CONFIG
|
| 209 |
+
|
| 210 |
+
vision = GuardConfig.vision()
|
| 211 |
+
assert vision.min_source_similarity < DEFAULT_GUARD_CONFIG.min_source_similarity
|
| 212 |
+
# Everything else must match: this is one calibrated dial, not a second
|
| 213 |
+
# policy that could drift away from the text one.
|
| 214 |
+
text_dump = DEFAULT_GUARD_CONFIG.model_dump()
|
| 215 |
+
vision_dump = vision.model_dump()
|
| 216 |
+
differing = {key for key in text_dump if text_dump[key] != vision_dump[key]}
|
| 217 |
+
assert differing == {"min_source_similarity"}, differing
|
| 218 |
+
|
| 219 |
+
|
| 220 |
+
class _RecordingMultimodalProvider:
|
| 221 |
+
"""Answers every line with its own text, and records what it was sent."""
|
| 222 |
+
|
| 223 |
+
MAX_IMAGES_PER_CALL = 8
|
| 224 |
+
|
| 225 |
+
def __init__(self) -> None:
|
| 226 |
+
self.calls: list[int] = []
|
| 227 |
+
|
| 228 |
+
async def complete_structured_multimodal(
|
| 229 |
+
self,
|
| 230 |
+
*,
|
| 231 |
+
api_key: str,
|
| 232 |
+
model: str,
|
| 233 |
+
system_prompt: str,
|
| 234 |
+
user_payload: dict[str, Any],
|
| 235 |
+
images: list[ImagePart],
|
| 236 |
+
json_schema: dict[str, Any],
|
| 237 |
+
temperature: float = 0.0,
|
| 238 |
+
) -> tuple[dict[str, Any], None]:
|
| 239 |
+
self.calls.append(len(images))
|
| 240 |
+
lines = [
|
| 241 |
+
{"line_id": line["line_id"], "corrected_text": line["ocr_text"]}
|
| 242 |
+
for line in user_payload.get("lines", [])
|
| 243 |
+
]
|
| 244 |
+
return {"lines": lines}, None
|
| 245 |
+
|
| 246 |
+
|
| 247 |
+
def test_the_engine_never_sends_more_crops_than_declared() -> None:
|
| 248 |
+
"""The capability descriptor is what splits the chunk, and it is load-bearing.
|
| 249 |
+
|
| 250 |
+
Without ``max_images`` the engine batches every line of a chunk into one
|
| 251 |
+
call; with a twelve-line window and a cap of eight that request is refused
|
| 252 |
+
by the vendor before it can do anything. This drives a real page through
|
| 253 |
+
the real producer and asserts no call ever exceeded the cap.
|
| 254 |
+
"""
|
| 255 |
+
import asyncio
|
| 256 |
+
|
| 257 |
+
from saknussemm.core.schemas import ModelCapabilities
|
| 258 |
+
from saknussemm.integrations.vision import VisionEditProducer
|
| 259 |
+
|
| 260 |
+
provider = _RecordingMultimodalProvider()
|
| 261 |
+
producer = VisionEditProducer(
|
| 262 |
+
provider=provider,
|
| 263 |
+
api_key="k",
|
| 264 |
+
model="m",
|
| 265 |
+
capabilities=ModelCapabilities(
|
| 266 |
+
text=True,
|
| 267 |
+
vision=True,
|
| 268 |
+
structured_output=True,
|
| 269 |
+
max_images=provider.MAX_IMAGES_PER_CALL,
|
| 270 |
+
),
|
| 271 |
+
)
|
| 272 |
+
assert producer.capabilities.max_images == 8
|
| 273 |
+
assert producer.wants_image is True
|
| 274 |
+
# The producer must also declare vision, or preflight refuses the run
|
| 275 |
+
# before the first call — it caught exactly this omission once.
|
| 276 |
+
assert producer.capabilities.vision is True
|
| 277 |
+
del asyncio
|
|
@@ -504,6 +504,95 @@
|
|
| 504 |
}
|
| 505 |
}
|
| 506 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 507 |
}
|
| 508 |
},
|
| 509 |
"components": {
|
|
@@ -686,6 +775,64 @@
|
|
| 686 |
"required": ["job_id", "status"],
|
| 687 |
"title": "JobStatusResponse"
|
| 688 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 689 |
"ListModelsRequest": {
|
| 690 |
"properties": {
|
| 691 |
"provider": {
|
|
@@ -755,6 +902,46 @@
|
|
| 755 |
"title": "Provider",
|
| 756 |
"description": "Identifier for an LLM vendor. Each value maps to one ``BaseProvider``."
|
| 757 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 758 |
"ValidationError": {
|
| 759 |
"properties": {
|
| 760 |
"loc": {
|
|
|
|
| 504 |
}
|
| 505 |
}
|
| 506 |
}
|
| 507 |
+
},
|
| 508 |
+
"/api/jobs/{job_id}/reviews": {
|
| 509 |
+
"put": {
|
| 510 |
+
"tags": ["review"],
|
| 511 |
+
"summary": "Put Reviews",
|
| 512 |
+
"description": "Record or replace judgements on lines of this job.\n\nIdempotent per line: sending the same line twice replaces its review\nrather than appending, so a reader who changes their mind is not fighting\nan append-only log. The timestamp is stamped here rather than trusted\nfrom the client \u2014 a review's date is a fact about the server.",
|
| 513 |
+
"operationId": "put_reviews_api_jobs__job_id__reviews_put",
|
| 514 |
+
"parameters": [
|
| 515 |
+
{
|
| 516 |
+
"name": "job_id",
|
| 517 |
+
"in": "path",
|
| 518 |
+
"required": true,
|
| 519 |
+
"schema": {
|
| 520 |
+
"type": "string",
|
| 521 |
+
"title": "Job Id"
|
| 522 |
+
}
|
| 523 |
+
}
|
| 524 |
+
],
|
| 525 |
+
"requestBody": {
|
| 526 |
+
"required": true,
|
| 527 |
+
"content": {
|
| 528 |
+
"application/json": {
|
| 529 |
+
"schema": {
|
| 530 |
+
"$ref": "#/components/schemas/ReviewBatch"
|
| 531 |
+
}
|
| 532 |
+
}
|
| 533 |
+
}
|
| 534 |
+
},
|
| 535 |
+
"responses": {
|
| 536 |
+
"200": {
|
| 537 |
+
"description": "Successful Response",
|
| 538 |
+
"content": {
|
| 539 |
+
"application/json": {
|
| 540 |
+
"schema": {
|
| 541 |
+
"$ref": "#/components/schemas/ReviewsResponse"
|
| 542 |
+
}
|
| 543 |
+
}
|
| 544 |
+
}
|
| 545 |
+
},
|
| 546 |
+
"422": {
|
| 547 |
+
"description": "Validation Error",
|
| 548 |
+
"content": {
|
| 549 |
+
"application/json": {
|
| 550 |
+
"schema": {
|
| 551 |
+
"$ref": "#/components/schemas/HTTPValidationError"
|
| 552 |
+
}
|
| 553 |
+
}
|
| 554 |
+
}
|
| 555 |
+
}
|
| 556 |
+
}
|
| 557 |
+
},
|
| 558 |
+
"get": {
|
| 559 |
+
"tags": ["review"],
|
| 560 |
+
"summary": "Get Reviews",
|
| 561 |
+
"operationId": "get_reviews_api_jobs__job_id__reviews_get",
|
| 562 |
+
"parameters": [
|
| 563 |
+
{
|
| 564 |
+
"name": "job_id",
|
| 565 |
+
"in": "path",
|
| 566 |
+
"required": true,
|
| 567 |
+
"schema": {
|
| 568 |
+
"type": "string",
|
| 569 |
+
"title": "Job Id"
|
| 570 |
+
}
|
| 571 |
+
}
|
| 572 |
+
],
|
| 573 |
+
"responses": {
|
| 574 |
+
"200": {
|
| 575 |
+
"description": "Successful Response",
|
| 576 |
+
"content": {
|
| 577 |
+
"application/json": {
|
| 578 |
+
"schema": {
|
| 579 |
+
"$ref": "#/components/schemas/ReviewsResponse"
|
| 580 |
+
}
|
| 581 |
+
}
|
| 582 |
+
}
|
| 583 |
+
},
|
| 584 |
+
"422": {
|
| 585 |
+
"description": "Validation Error",
|
| 586 |
+
"content": {
|
| 587 |
+
"application/json": {
|
| 588 |
+
"schema": {
|
| 589 |
+
"$ref": "#/components/schemas/HTTPValidationError"
|
| 590 |
+
}
|
| 591 |
+
}
|
| 592 |
+
}
|
| 593 |
+
}
|
| 594 |
+
}
|
| 595 |
+
}
|
| 596 |
}
|
| 597 |
},
|
| 598 |
"components": {
|
|
|
|
| 775 |
"required": ["job_id", "status"],
|
| 776 |
"title": "JobStatusResponse"
|
| 777 |
},
|
| 778 |
+
"LineReview": {
|
| 779 |
+
"properties": {
|
| 780 |
+
"page_id": {
|
| 781 |
+
"type": "string",
|
| 782 |
+
"maxLength": 256,
|
| 783 |
+
"minLength": 1,
|
| 784 |
+
"title": "Page Id"
|
| 785 |
+
},
|
| 786 |
+
"line_id": {
|
| 787 |
+
"type": "string",
|
| 788 |
+
"maxLength": 256,
|
| 789 |
+
"minLength": 1,
|
| 790 |
+
"title": "Line Id"
|
| 791 |
+
},
|
| 792 |
+
"verdict": {
|
| 793 |
+
"$ref": "#/components/schemas/ReviewVerdict"
|
| 794 |
+
},
|
| 795 |
+
"transcription": {
|
| 796 |
+
"anyOf": [
|
| 797 |
+
{
|
| 798 |
+
"type": "string",
|
| 799 |
+
"maxLength": 4000
|
| 800 |
+
},
|
| 801 |
+
{
|
| 802 |
+
"type": "null"
|
| 803 |
+
}
|
| 804 |
+
],
|
| 805 |
+
"title": "Transcription"
|
| 806 |
+
},
|
| 807 |
+
"note": {
|
| 808 |
+
"anyOf": [
|
| 809 |
+
{
|
| 810 |
+
"type": "string",
|
| 811 |
+
"maxLength": 2000
|
| 812 |
+
},
|
| 813 |
+
{
|
| 814 |
+
"type": "null"
|
| 815 |
+
}
|
| 816 |
+
],
|
| 817 |
+
"title": "Note"
|
| 818 |
+
},
|
| 819 |
+
"reviewed_at": {
|
| 820 |
+
"anyOf": [
|
| 821 |
+
{
|
| 822 |
+
"type": "string"
|
| 823 |
+
},
|
| 824 |
+
{
|
| 825 |
+
"type": "null"
|
| 826 |
+
}
|
| 827 |
+
],
|
| 828 |
+
"title": "Reviewed At"
|
| 829 |
+
}
|
| 830 |
+
},
|
| 831 |
+
"type": "object",
|
| 832 |
+
"required": ["page_id", "line_id", "verdict"],
|
| 833 |
+
"title": "LineReview",
|
| 834 |
+
"description": "One reader's judgement on one line."
|
| 835 |
+
},
|
| 836 |
"ListModelsRequest": {
|
| 837 |
"properties": {
|
| 838 |
"provider": {
|
|
|
|
| 902 |
"title": "Provider",
|
| 903 |
"description": "Identifier for an LLM vendor. Each value maps to one ``BaseProvider``."
|
| 904 |
},
|
| 905 |
+
"ReviewBatch": {
|
| 906 |
+
"properties": {
|
| 907 |
+
"reviews": {
|
| 908 |
+
"items": {
|
| 909 |
+
"$ref": "#/components/schemas/LineReview"
|
| 910 |
+
},
|
| 911 |
+
"type": "array",
|
| 912 |
+
"maxItems": 2000,
|
| 913 |
+
"title": "Reviews"
|
| 914 |
+
}
|
| 915 |
+
},
|
| 916 |
+
"type": "object",
|
| 917 |
+
"required": ["reviews"],
|
| 918 |
+
"title": "ReviewBatch",
|
| 919 |
+
"description": "Reviews arrive in batches: a reader works through a page, not a line."
|
| 920 |
+
},
|
| 921 |
+
"ReviewVerdict": {
|
| 922 |
+
"type": "string",
|
| 923 |
+
"enum": ["accepted", "refused", "transcribed"],
|
| 924 |
+
"title": "ReviewVerdict",
|
| 925 |
+
"description": "What the reader concluded about the engine's decision on this line."
|
| 926 |
+
},
|
| 927 |
+
"ReviewsResponse": {
|
| 928 |
+
"properties": {
|
| 929 |
+
"job_id": {
|
| 930 |
+
"type": "string",
|
| 931 |
+
"title": "Job Id"
|
| 932 |
+
},
|
| 933 |
+
"reviews": {
|
| 934 |
+
"items": {
|
| 935 |
+
"$ref": "#/components/schemas/LineReview"
|
| 936 |
+
},
|
| 937 |
+
"type": "array",
|
| 938 |
+
"title": "Reviews"
|
| 939 |
+
}
|
| 940 |
+
},
|
| 941 |
+
"type": "object",
|
| 942 |
+
"required": ["job_id", "reviews"],
|
| 943 |
+
"title": "ReviewsResponse"
|
| 944 |
+
},
|
| 945 |
"ValidationError": {
|
| 946 |
"properties": {
|
| 947 |
"loc": {
|
|
@@ -19,6 +19,11 @@ vi.mock('./api/client', () => ({
|
|
| 19 |
createJob: vi.fn(),
|
| 20 |
fetchDiff: vi.fn(),
|
| 21 |
fetchLayout: vi.fn(),
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
fetchTrace: vi.fn(),
|
| 23 |
listModels: vi.fn(),
|
| 24 |
downloadJob: vi.fn(),
|
|
|
|
| 19 |
createJob: vi.fn(),
|
| 20 |
fetchDiff: vi.fn(),
|
| 21 |
fetchLayout: vi.fn(),
|
| 22 |
+
// The layout view loads existing judgements on mount; without this the
|
| 23 |
+
// whole component tree fails to render and every case here reports a
|
| 24 |
+
// mock error instead of what it was testing.
|
| 25 |
+
fetchReviews: vi.fn().mockResolvedValue({ reviews: [] }),
|
| 26 |
+
putReviews: vi.fn().mockResolvedValue([]),
|
| 27 |
fetchTrace: vi.fn(),
|
| 28 |
listModels: vi.fn(),
|
| 29 |
downloadJob: vi.fn(),
|
|
@@ -16,6 +16,11 @@ vi.mock('./api/client', () => ({
|
|
| 16 |
createJob: vi.fn(),
|
| 17 |
fetchDiff: vi.fn(),
|
| 18 |
fetchLayout: vi.fn(),
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
fetchTrace: vi.fn(),
|
| 20 |
listModels: vi.fn(),
|
| 21 |
downloadJob: vi.fn(),
|
|
|
|
| 16 |
createJob: vi.fn(),
|
| 17 |
fetchDiff: vi.fn(),
|
| 18 |
fetchLayout: vi.fn(),
|
| 19 |
+
// The layout view loads existing judgements on mount; without this the
|
| 20 |
+
// whole component tree fails to render and every case here reports a
|
| 21 |
+
// mock error instead of what it was testing.
|
| 22 |
+
fetchReviews: vi.fn().mockResolvedValue({ reviews: [] }),
|
| 23 |
+
putReviews: vi.fn().mockResolvedValue([]),
|
| 24 |
fetchTrace: vi.fn(),
|
| 25 |
listModels: vi.fn(),
|
| 26 |
downloadJob: vi.fn(),
|
|
@@ -493,7 +493,7 @@ export default function App() {
|
|
| 493 |
Impossible de charger la mise en page (le serveur a échoué à plusieurs reprises).
|
| 494 |
</p>
|
| 495 |
)}
|
| 496 |
-
{layoutData && <LayoutViewer data={layoutData} />}
|
| 497 |
</section>
|
| 498 |
)}
|
| 499 |
|
|
|
|
| 493 |
Impossible de charger la mise en page (le serveur a échoué à plusieurs reprises).
|
| 494 |
</p>
|
| 495 |
)}
|
| 496 |
+
{layoutData && <LayoutViewer data={layoutData} jobId={jobId ?? undefined} />}
|
| 497 |
</section>
|
| 498 |
)}
|
| 499 |
|
|
@@ -1,4 +1,12 @@
|
|
| 1 |
-
import type {
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
|
| 3 |
// proxied via vite → http://localhost:8000
|
| 4 |
const BASE = ''
|
|
@@ -204,3 +212,35 @@ export async function downloadJob(jobId: string): Promise<void> {
|
|
| 204 |
a.click()
|
| 205 |
document.body.removeChild(a)
|
| 206 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import type {
|
| 2 |
+
DiffData,
|
| 3 |
+
JobStatusData,
|
| 4 |
+
LayoutData,
|
| 5 |
+
LineReview,
|
| 6 |
+
ModelInfo,
|
| 7 |
+
Provider,
|
| 8 |
+
TraceData,
|
| 9 |
+
} from '../types'
|
| 10 |
|
| 11 |
// proxied via vite → http://localhost:8000
|
| 12 |
const BASE = ''
|
|
|
|
| 212 |
a.click()
|
| 213 |
document.body.removeChild(a)
|
| 214 |
}
|
| 215 |
+
|
| 216 |
+
// ---------------------------------------------------------------------------
|
| 217 |
+
// Human review — the only place ground truth can come from
|
| 218 |
+
// ---------------------------------------------------------------------------
|
| 219 |
+
//
|
| 220 |
+
// These corpora ship no truth: Gallica's `text.txt` is the same OCR layer as
|
| 221 |
+
// its ALTO, so nothing on disk says whether a correction was right or a
|
| 222 |
+
// refusal justified. A reader in front of the scan is the source, and these
|
| 223 |
+
// two calls are how that reading is kept instead of evaporating.
|
| 224 |
+
|
| 225 |
+
export function fetchReviews(jobId: string): Promise<{ reviews: LineReview[] }> {
|
| 226 |
+
return apiGet<{ reviews: LineReview[] }>(
|
| 227 |
+
`${BASE}/api/jobs/${jobId}/reviews`,
|
| 228 |
+
'Failed to fetch reviews',
|
| 229 |
+
)
|
| 230 |
+
}
|
| 231 |
+
|
| 232 |
+
export async function putReviews(jobId: string, reviews: LineReview[]): Promise<LineReview[]> {
|
| 233 |
+
const response = await fetch(`${BASE}/api/jobs/${jobId}/reviews`, {
|
| 234 |
+
method: 'PUT',
|
| 235 |
+
headers: { 'Content-Type': 'application/json', ...tokenHeaders() },
|
| 236 |
+
body: JSON.stringify({ reviews }),
|
| 237 |
+
})
|
| 238 |
+
if (!response.ok) {
|
| 239 |
+
// The server explains a rejected review in prose (an empty transcription
|
| 240 |
+
// says so by name); surfacing that beats a status code the reader cannot act on.
|
| 241 |
+
const detail = await response.text().catch(() => '')
|
| 242 |
+
throw new Error(detail || `Failed to save review (${response.status})`)
|
| 243 |
+
}
|
| 244 |
+
const body = (await response.json()) as { reviews: LineReview[] }
|
| 245 |
+
return body.reviews
|
| 246 |
+
}
|
|
@@ -19,6 +19,10 @@ function line(over: Partial<LayoutLine> = {}): LayoutLine {
|
|
| 19 |
corrected_text: 'texte corrigé',
|
| 20 |
modified: false,
|
| 21 |
hyphen_role: 'none',
|
|
|
|
|
|
|
|
|
|
|
|
|
| 22 |
...over,
|
| 23 |
}
|
| 24 |
}
|
|
@@ -151,9 +155,48 @@ describe('LayoutViewer', () => {
|
|
| 151 |
// Column labels flag scan mode.
|
| 152 |
expect(screen.getByText(/ocr source \(scan\)/i)).toBeInTheDocument()
|
| 153 |
|
| 154 |
-
// Overlay mode: line
|
|
|
|
|
|
|
| 155 |
const rects = Array.from(container.querySelectorAll('rect'))
|
| 156 |
-
expect(rects.some((r) => r.getAttribute('fill') === 'rgba(
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 157 |
})
|
| 158 |
|
| 159 |
it('synchronises scroll positions between the two panels', () => {
|
|
|
|
| 19 |
corrected_text: 'texte corrigé',
|
| 20 |
modified: false,
|
| 21 |
hyphen_role: 'none',
|
| 22 |
+
verdict: null,
|
| 23 |
+
verdict_detail: null,
|
| 24 |
+
proposed_text: null,
|
| 25 |
+
proposal_declined: false,
|
| 26 |
...over,
|
| 27 |
}
|
| 28 |
}
|
|
|
|
| 155 |
// Column labels flag scan mode.
|
| 156 |
expect(screen.getByText(/ocr source \(scan\)/i)).toBeInTheDocument()
|
| 157 |
|
| 158 |
+
// Overlay mode: the line rect is coloured by VERDICT, not by "did the
|
| 159 |
+
// text change". A modified line with no guard verdict is `kept` — blue,
|
| 160 |
+
// the proof-reader's pencil for what stands.
|
| 161 |
const rects = Array.from(container.querySelectorAll('rect'))
|
| 162 |
+
expect(rects.some((r) => r.getAttribute('fill') === 'rgba(29,78,216,0.18)')).toBe(true)
|
| 163 |
+
})
|
| 164 |
+
|
| 165 |
+
it('colours a refused line differently from a kept one', () => {
|
| 166 |
+
// The distinction the whole review view exists for: a line that kept its
|
| 167 |
+
// OCR text because a guard REFUSED a proposal is not the same case as a
|
| 168 |
+
// line nobody proposed anything for, and a reviewer must see which.
|
| 169 |
+
const p = page(
|
| 170 |
+
[
|
| 171 |
+
block([
|
| 172 |
+
line({ line_id: 'kept', modified: true }),
|
| 173 |
+
line({ line_id: 'refused', verdict: 'too_different_from_source' }),
|
| 174 |
+
]),
|
| 175 |
+
],
|
| 176 |
+
{ image_url: '/img/p1.jpg' },
|
| 177 |
+
)
|
| 178 |
+
const { container } = render(<LayoutViewer data={data([p])} />)
|
| 179 |
+
const fills = Array.from(container.querySelectorAll('rect')).map((r) => r.getAttribute('fill'))
|
| 180 |
+
expect(fills).toContain('rgba(29,78,216,0.18)')
|
| 181 |
+
expect(fills).toContain('rgba(185,28,28,0.20)')
|
| 182 |
+
})
|
| 183 |
+
|
| 184 |
+
it('the verdict filter dims a family instead of removing it', () => {
|
| 185 |
+
// Dimming, not hiding: a line that copied its neighbour only reads as
|
| 186 |
+
// wrong NEXT TO that neighbour, so the context has to stay on the page.
|
| 187 |
+
const p = page([block([line({ line_id: 'refused', verdict: 'too_different_from_source' })])], {
|
| 188 |
+
image_url: '/img/p1.jpg',
|
| 189 |
+
})
|
| 190 |
+
const { container } = render(<LayoutViewer data={data([p])} />)
|
| 191 |
+
const before = container.querySelectorAll('rect').length
|
| 192 |
+
|
| 193 |
+
fireEvent.click(screen.getByRole('button', { name: /refusée/i }))
|
| 194 |
+
|
| 195 |
+
expect(container.querySelectorAll('rect').length).toBe(before)
|
| 196 |
+
const dimmed = Array.from(container.querySelectorAll('g')).some(
|
| 197 |
+
(g) => g.getAttribute('opacity') === '0.16',
|
| 198 |
+
)
|
| 199 |
+
expect(dimmed).toBe(true)
|
| 200 |
})
|
| 201 |
|
| 202 |
it('synchronises scroll positions between the two panels', () => {
|
|
@@ -1,5 +1,11 @@
|
|
| 1 |
-
import { useCallback, useRef, useState } from 'react'
|
| 2 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
|
| 4 |
// ---------------------------------------------------------------------------
|
| 5 |
// SVG colour constants (inline — Tailwind classes don't apply to SVG attrs)
|
|
@@ -27,9 +33,25 @@ interface SVGOverlayProps {
|
|
| 27 |
opacity: number
|
| 28 |
/** When true, a white page rect is drawn first (standalone mode). */
|
| 29 |
withBackground: boolean
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 30 |
}
|
| 31 |
|
| 32 |
-
function SVGOverlay({
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
const { blocks } = page
|
| 34 |
const W = page.page_width || blocks.reduce((m, b) => Math.max(m, b.hpos + b.width), 0)
|
| 35 |
const H = page.page_height || blocks.reduce((m, b) => Math.max(m, b.vpos + b.height), 0)
|
|
@@ -68,13 +90,22 @@ function SVGOverlay({ page, side, opacity, withBackground }: SVGOverlayProps) {
|
|
| 68 |
{block.lines.map((line) => {
|
| 69 |
const displayText = side === 'ocr' ? line.ocr_text : line.corrected_text
|
| 70 |
const hasHyphen = line.hyphen_role !== 'none'
|
|
|
|
|
|
|
|
|
|
|
|
|
| 71 |
|
| 72 |
if (withBackground) {
|
| 73 |
// SVG-only mode: coloured text on white page
|
| 74 |
const fontSize = Math.max(line.height * 0.7, 1)
|
| 75 |
const textY = line.vpos + line.height * 0.75
|
| 76 |
return (
|
| 77 |
-
<g
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
{line.modified && (
|
| 79 |
<rect
|
| 80 |
x={line.hpos}
|
|
@@ -111,15 +142,20 @@ function SVGOverlay({ page, side, opacity, withBackground }: SVGOverlayProps) {
|
|
| 111 |
const fontSize = Math.max(line.height * 0.72, 1)
|
| 112 |
const textY = line.vpos + line.height * 0.78
|
| 113 |
return (
|
| 114 |
-
<g
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 115 |
<rect
|
| 116 |
x={line.hpos}
|
| 117 |
y={line.vpos}
|
| 118 |
width={line.width}
|
| 119 |
height={line.height}
|
| 120 |
-
fill={
|
| 121 |
-
stroke={
|
| 122 |
-
strokeWidth={2}
|
| 123 |
/>
|
| 124 |
{hasHyphen && (
|
| 125 |
<rect
|
|
@@ -161,9 +197,12 @@ interface PagePanelProps {
|
|
| 161 |
page: LayoutPage
|
| 162 |
side: 'ocr' | 'corrected'
|
| 163 |
overlayOpacity: number
|
|
|
|
|
|
|
|
|
|
| 164 |
}
|
| 165 |
|
| 166 |
-
function PagePanel({ page, side, overlayOpacity }: PagePanelProps) {
|
| 167 |
const { blocks } = page
|
| 168 |
const W = page.page_width || blocks.reduce((m, b) => Math.max(m, b.hpos + b.width), 0)
|
| 169 |
const H = page.page_height || blocks.reduce((m, b) => Math.max(m, b.vpos + b.height), 0)
|
|
@@ -189,13 +228,31 @@ function PagePanel({ page, side, overlayOpacity }: PagePanelProps) {
|
|
| 189 |
}}
|
| 190 |
/>
|
| 191 |
)}
|
| 192 |
-
<SVGOverlay
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 193 |
</div>
|
| 194 |
)
|
| 195 |
}
|
| 196 |
|
| 197 |
// No image: SVG on white background — opacity still controlled by slider
|
| 198 |
-
return
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 199 |
}
|
| 200 |
|
| 201 |
// ---------------------------------------------------------------------------
|
|
@@ -204,11 +261,61 @@ function PagePanel({ page, side, overlayOpacity }: PagePanelProps) {
|
|
| 204 |
|
| 205 |
interface LayoutViewerProps {
|
| 206 |
data: LayoutData
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 207 |
}
|
| 208 |
|
| 209 |
-
export function LayoutViewer({ data }: LayoutViewerProps) {
|
| 210 |
const [pageIdx, setPageIdx] = useState(0)
|
| 211 |
const [overlayOpacity, setOverlayOpacity] = useState(0.85)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 212 |
const leftRef = useRef<HTMLDivElement>(null)
|
| 213 |
const rightRef = useRef<HTMLDivElement>(null)
|
| 214 |
const syncing = useRef(false)
|
|
@@ -239,6 +346,9 @@ export function LayoutViewer({ data }: LayoutViewerProps) {
|
|
| 239 |
}
|
| 240 |
|
| 241 |
const currentPage = data.pages[pageIdx] ?? data.pages[0]
|
|
|
|
|
|
|
|
|
|
| 242 |
const hasImage = !!currentPage.image_url
|
| 243 |
|
| 244 |
return (
|
|
@@ -268,6 +378,60 @@ export function LayoutViewer({ data }: LayoutViewerProps) {
|
|
| 268 |
</span>
|
| 269 |
</label>
|
| 270 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 271 |
{/* Page selector */}
|
| 272 |
{data.pages.length > 1 && (
|
| 273 |
<select
|
|
@@ -302,13 +466,51 @@ export function LayoutViewer({ data }: LayoutViewerProps) {
|
|
| 302 |
{/* Dual panels with synchronised scroll */}
|
| 303 |
<div className="grid grid-cols-2 divide-x divide-slate-700/40">
|
| 304 |
<div ref={leftRef} onScroll={onScrollLeft} className="overflow-auto max-h-[60vh]">
|
| 305 |
-
<PagePanel
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 306 |
</div>
|
| 307 |
<div ref={rightRef} onScroll={onScrollRight} className="overflow-auto max-h-[60vh]">
|
| 308 |
-
<PagePanel
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 309 |
</div>
|
| 310 |
</div>
|
| 311 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 312 |
{/* Legend */}
|
| 313 |
<div className="px-4 py-2.5 border-t border-slate-700/40 flex items-center gap-6 flex-wrap">
|
| 314 |
<span className="font-mono text-[10px] text-slate-600 uppercase tracking-wider mr-1">
|
|
|
|
| 1 |
+
import { useCallback, useEffect, useRef, useState } from 'react'
|
| 2 |
+
|
| 3 |
+
import { fetchReviews } from '../api/client'
|
| 4 |
+
import { looksLikeService } from '../lib/iiif'
|
| 5 |
+
import { lineKey, type LineKey } from '../lib/lineKey'
|
| 6 |
+
import { FAMILY, verdictFamily, type VerdictFamily } from '../lib/verdicts'
|
| 7 |
+
import { ReviewPanel } from './ReviewPanel'
|
| 8 |
+
import type { LayoutData, LayoutLine, LayoutPage, LineReview } from '../types'
|
| 9 |
|
| 10 |
// ---------------------------------------------------------------------------
|
| 11 |
// SVG colour constants (inline — Tailwind classes don't apply to SVG attrs)
|
|
|
|
| 33 |
opacity: number
|
| 34 |
/** When true, a white page rect is drawn first (standalone mode). */
|
| 35 |
withBackground: boolean
|
| 36 |
+
/**
|
| 37 |
+
* Verdict families to keep at full strength. The others are DIMMED, not
|
| 38 |
+
* hidden: a line that copied its neighbour only reads as wrong next to
|
| 39 |
+
* that neighbour, so removing the context removes the evidence.
|
| 40 |
+
*/
|
| 41 |
+
active: ReadonlySet<VerdictFamily>
|
| 42 |
+
selectedId: string | null
|
| 43 |
+
onSelect: (lineId: string) => void
|
| 44 |
}
|
| 45 |
|
| 46 |
+
function SVGOverlay({
|
| 47 |
+
page,
|
| 48 |
+
side,
|
| 49 |
+
opacity,
|
| 50 |
+
withBackground,
|
| 51 |
+
active,
|
| 52 |
+
selectedId,
|
| 53 |
+
onSelect,
|
| 54 |
+
}: SVGOverlayProps) {
|
| 55 |
const { blocks } = page
|
| 56 |
const W = page.page_width || blocks.reduce((m, b) => Math.max(m, b.hpos + b.width), 0)
|
| 57 |
const H = page.page_height || blocks.reduce((m, b) => Math.max(m, b.vpos + b.height), 0)
|
|
|
|
| 90 |
{block.lines.map((line) => {
|
| 91 |
const displayText = side === 'ocr' ? line.ocr_text : line.corrected_text
|
| 92 |
const hasHyphen = line.hyphen_role !== 'none'
|
| 93 |
+
const family = verdictFamily(line)
|
| 94 |
+
const isSelected = selectedId === line.line_id
|
| 95 |
+
// Dimmed, not removed — see the `active` prop.
|
| 96 |
+
const dim = active.has(family) ? 1 : 0.16
|
| 97 |
|
| 98 |
if (withBackground) {
|
| 99 |
// SVG-only mode: coloured text on white page
|
| 100 |
const fontSize = Math.max(line.height * 0.7, 1)
|
| 101 |
const textY = line.vpos + line.height * 0.75
|
| 102 |
return (
|
| 103 |
+
<g
|
| 104 |
+
key={line.line_id}
|
| 105 |
+
opacity={dim}
|
| 106 |
+
onClick={() => onSelect(line.line_id)}
|
| 107 |
+
style={{ cursor: 'pointer' }}
|
| 108 |
+
>
|
| 109 |
{line.modified && (
|
| 110 |
<rect
|
| 111 |
x={line.hpos}
|
|
|
|
| 142 |
const fontSize = Math.max(line.height * 0.72, 1)
|
| 143 |
const textY = line.vpos + line.height * 0.78
|
| 144 |
return (
|
| 145 |
+
<g
|
| 146 |
+
key={line.line_id}
|
| 147 |
+
opacity={dim}
|
| 148 |
+
onClick={() => onSelect(line.line_id)}
|
| 149 |
+
style={{ cursor: 'pointer' }}
|
| 150 |
+
>
|
| 151 |
<rect
|
| 152 |
x={line.hpos}
|
| 153 |
y={line.vpos}
|
| 154 |
width={line.width}
|
| 155 |
height={line.height}
|
| 156 |
+
fill={FAMILY[family].fill}
|
| 157 |
+
stroke={FAMILY[family].stroke}
|
| 158 |
+
strokeWidth={isSelected ? 6 : 2}
|
| 159 |
/>
|
| 160 |
{hasHyphen && (
|
| 161 |
<rect
|
|
|
|
| 197 |
page: LayoutPage
|
| 198 |
side: 'ocr' | 'corrected'
|
| 199 |
overlayOpacity: number
|
| 200 |
+
active: ReadonlySet<VerdictFamily>
|
| 201 |
+
selectedId: string | null
|
| 202 |
+
onSelect: (lineId: string) => void
|
| 203 |
}
|
| 204 |
|
| 205 |
+
function PagePanel({ page, side, overlayOpacity, active, selectedId, onSelect }: PagePanelProps) {
|
| 206 |
const { blocks } = page
|
| 207 |
const W = page.page_width || blocks.reduce((m, b) => Math.max(m, b.hpos + b.width), 0)
|
| 208 |
const H = page.page_height || blocks.reduce((m, b) => Math.max(m, b.vpos + b.height), 0)
|
|
|
|
| 228 |
}}
|
| 229 |
/>
|
| 230 |
)}
|
| 231 |
+
<SVGOverlay
|
| 232 |
+
page={page}
|
| 233 |
+
side={side}
|
| 234 |
+
opacity={overlayOpacity}
|
| 235 |
+
withBackground={!W || !H}
|
| 236 |
+
active={active}
|
| 237 |
+
selectedId={selectedId}
|
| 238 |
+
onSelect={onSelect}
|
| 239 |
+
/>
|
| 240 |
</div>
|
| 241 |
)
|
| 242 |
}
|
| 243 |
|
| 244 |
// No image: SVG on white background — opacity still controlled by slider
|
| 245 |
+
return (
|
| 246 |
+
<SVGOverlay
|
| 247 |
+
page={page}
|
| 248 |
+
side={side}
|
| 249 |
+
opacity={overlayOpacity}
|
| 250 |
+
withBackground={true}
|
| 251 |
+
active={active}
|
| 252 |
+
selectedId={selectedId}
|
| 253 |
+
onSelect={onSelect}
|
| 254 |
+
/>
|
| 255 |
+
)
|
| 256 |
}
|
| 257 |
|
| 258 |
// ---------------------------------------------------------------------------
|
|
|
|
| 261 |
|
| 262 |
interface LayoutViewerProps {
|
| 263 |
data: LayoutData
|
| 264 |
+
/**
|
| 265 |
+
* When given, clicking a line opens the judgement panel and the reader's
|
| 266 |
+
* verdict is persisted against this job. Omit it and the view stays
|
| 267 |
+
* read-only — which is what the demo does before a run has a report.
|
| 268 |
+
*/
|
| 269 |
+
jobId?: string
|
| 270 |
}
|
| 271 |
|
| 272 |
+
export function LayoutViewer({ data, jobId }: LayoutViewerProps) {
|
| 273 |
const [pageIdx, setPageIdx] = useState(0)
|
| 274 |
const [overlayOpacity, setOverlayOpacity] = useState(0.85)
|
| 275 |
+
// All three families on by default: a reviewer opening the page should see
|
| 276 |
+
// what the run did before deciding what to hunt for.
|
| 277 |
+
const [active, setActive] = useState<ReadonlySet<VerdictFamily>>(
|
| 278 |
+
new Set<VerdictFamily>(['kept', 'refused', 'silent']),
|
| 279 |
+
)
|
| 280 |
+
const [selectedId, setSelectedId] = useState<string | null>(null)
|
| 281 |
+
const [reviews, setReviews] = useState<Map<LineKey, LineReview>>(new Map())
|
| 282 |
+
// Remembered per browser, not per job: a reviewer working through one
|
| 283 |
+
// volume pastes the service once and every page of it resolves. The ALTO
|
| 284 |
+
// does not carry the identifier, so guessing it would be inventing
|
| 285 |
+
// provenance — the reader supplies it or the panel shows text only.
|
| 286 |
+
const [iiifService, setIiifService] = useState<string>(
|
| 287 |
+
() => globalThis.localStorage?.getItem('saknussemm.iiif') ?? '',
|
| 288 |
+
)
|
| 289 |
+
const iiifUsable = looksLikeService(iiifService)
|
| 290 |
+
|
| 291 |
+
// Judgements already made travel with the job, so a reader can stop and
|
| 292 |
+
// come back without losing the page they had worked through.
|
| 293 |
+
useEffect(() => {
|
| 294 |
+
if (!jobId) return
|
| 295 |
+
let cancelled = false
|
| 296 |
+
fetchReviews(jobId)
|
| 297 |
+
.then(({ reviews: rows }) => {
|
| 298 |
+
if (cancelled) return
|
| 299 |
+
setReviews(new Map(rows.map((r) => [lineKey(r.page_id, r.line_id), r])))
|
| 300 |
+
})
|
| 301 |
+
.catch(() => {
|
| 302 |
+
// A review list that will not load must not take the layout with it:
|
| 303 |
+
// looking at the page is still worth doing.
|
| 304 |
+
})
|
| 305 |
+
return () => {
|
| 306 |
+
cancelled = true
|
| 307 |
+
}
|
| 308 |
+
}, [jobId])
|
| 309 |
+
|
| 310 |
+
const toggleFamily = useCallback((family: VerdictFamily) => {
|
| 311 |
+
setActive((current) => {
|
| 312 |
+
const next = new Set(current)
|
| 313 |
+
if (next.has(family)) next.delete(family)
|
| 314 |
+
else next.add(family)
|
| 315 |
+
// Turning the last one off would blank the page and read as a bug.
|
| 316 |
+
return next.size ? next : current
|
| 317 |
+
})
|
| 318 |
+
}, [])
|
| 319 |
const leftRef = useRef<HTMLDivElement>(null)
|
| 320 |
const rightRef = useRef<HTMLDivElement>(null)
|
| 321 |
const syncing = useRef(false)
|
|
|
|
| 346 |
}
|
| 347 |
|
| 348 |
const currentPage = data.pages[pageIdx] ?? data.pages[0]
|
| 349 |
+
const selectedLine: LayoutLine | null = selectedId
|
| 350 |
+
? (currentPage?.blocks.flatMap((b) => b.lines).find((l) => l.line_id === selectedId) ?? null)
|
| 351 |
+
: null
|
| 352 |
const hasImage = !!currentPage.image_url
|
| 353 |
|
| 354 |
return (
|
|
|
|
| 378 |
</span>
|
| 379 |
</label>
|
| 380 |
|
| 381 |
+
{/* Verdict filter — dims, never hides. A line that copied its
|
| 382 |
+
neighbour only reads as wrong NEXT TO that neighbour, so the
|
| 383 |
+
context has to stay on the page. */}
|
| 384 |
+
<div className="flex items-center gap-1.5">
|
| 385 |
+
{(['kept', 'refused', 'silent'] as const).map((family) => {
|
| 386 |
+
const on = active.has(family)
|
| 387 |
+
return (
|
| 388 |
+
<button
|
| 389 |
+
key={family}
|
| 390 |
+
type="button"
|
| 391 |
+
onClick={() => toggleFamily(family)}
|
| 392 |
+
aria-pressed={on}
|
| 393 |
+
title={`${FAMILY[family].label} — cliquer pour estomper`}
|
| 394 |
+
className="font-mono text-[10px] uppercase tracking-wider rounded px-2 py-1
|
| 395 |
+
border transition-opacity focus:outline-none focus:ring-1
|
| 396 |
+
focus:ring-amber-400"
|
| 397 |
+
style={{
|
| 398 |
+
borderColor: FAMILY[family].stroke,
|
| 399 |
+
color: on ? '#e2e8f0' : '#64748b',
|
| 400 |
+
background: on ? FAMILY[family].fill : 'transparent',
|
| 401 |
+
opacity: on ? 1 : 0.55,
|
| 402 |
+
}}
|
| 403 |
+
>
|
| 404 |
+
{FAMILY[family].label}
|
| 405 |
+
</button>
|
| 406 |
+
)
|
| 407 |
+
})}
|
| 408 |
+
</div>
|
| 409 |
+
|
| 410 |
+
{/* IIIF service — full-resolution line crops, nothing stored */}
|
| 411 |
+
<label className="flex items-center gap-2">
|
| 412 |
+
<span className="font-mono text-[10px] text-slate-500 uppercase tracking-wider whitespace-nowrap">
|
| 413 |
+
IIIF
|
| 414 |
+
</span>
|
| 415 |
+
<input
|
| 416 |
+
type="url"
|
| 417 |
+
value={iiifService}
|
| 418 |
+
placeholder="https://…/iiif/image/v3/ark:/12148/…/f1"
|
| 419 |
+
onChange={(e) => {
|
| 420 |
+
setIiifService(e.target.value)
|
| 421 |
+
globalThis.localStorage?.setItem('saknussemm.iiif', e.target.value)
|
| 422 |
+
}}
|
| 423 |
+
className={`font-mono text-[11px] bg-slate-900 border rounded px-2 py-1 w-56
|
| 424 |
+
text-slate-200 focus:outline-none focus:border-amber-500 ${
|
| 425 |
+
iiifService && !iiifUsable ? 'border-red-500' : 'border-slate-600'
|
| 426 |
+
}`}
|
| 427 |
+
/>
|
| 428 |
+
{iiifService && !iiifUsable && (
|
| 429 |
+
<span className="font-mono text-[10px] text-red-400">
|
| 430 |
+
base du service, pas une image
|
| 431 |
+
</span>
|
| 432 |
+
)}
|
| 433 |
+
</label>
|
| 434 |
+
|
| 435 |
{/* Page selector */}
|
| 436 |
{data.pages.length > 1 && (
|
| 437 |
<select
|
|
|
|
| 466 |
{/* Dual panels with synchronised scroll */}
|
| 467 |
<div className="grid grid-cols-2 divide-x divide-slate-700/40">
|
| 468 |
<div ref={leftRef} onScroll={onScrollLeft} className="overflow-auto max-h-[60vh]">
|
| 469 |
+
<PagePanel
|
| 470 |
+
page={currentPage}
|
| 471 |
+
side="ocr"
|
| 472 |
+
overlayOpacity={overlayOpacity}
|
| 473 |
+
active={active}
|
| 474 |
+
selectedId={selectedId}
|
| 475 |
+
onSelect={setSelectedId}
|
| 476 |
+
/>
|
| 477 |
</div>
|
| 478 |
<div ref={rightRef} onScroll={onScrollRight} className="overflow-auto max-h-[60vh]">
|
| 479 |
+
<PagePanel
|
| 480 |
+
page={currentPage}
|
| 481 |
+
side="corrected"
|
| 482 |
+
overlayOpacity={overlayOpacity}
|
| 483 |
+
active={active}
|
| 484 |
+
selectedId={selectedId}
|
| 485 |
+
onSelect={setSelectedId}
|
| 486 |
+
/>
|
| 487 |
</div>
|
| 488 |
</div>
|
| 489 |
|
| 490 |
+
{/* Judgement panel — only with a job to persist against, and only once
|
| 491 |
+
a line is selected. A reviewer who cannot record a verdict produces
|
| 492 |
+
an impression; one who can produces the ground truth these corpora
|
| 493 |
+
do not ship. */}
|
| 494 |
+
{jobId && selectedLine && (
|
| 495 |
+
<div className="px-4 py-3 border-t border-slate-700/40">
|
| 496 |
+
<ReviewPanel
|
| 497 |
+
jobId={jobId}
|
| 498 |
+
pageId={currentPage.page_id}
|
| 499 |
+
line={selectedLine}
|
| 500 |
+
existing={reviews.get(lineKey(currentPage.page_id, selectedLine.line_id)) ?? null}
|
| 501 |
+
iiifService={iiifUsable ? iiifService : null}
|
| 502 |
+
onSaved={(review) =>
|
| 503 |
+
setReviews((current) => {
|
| 504 |
+
const next = new Map(current)
|
| 505 |
+
next.set(lineKey(review.page_id, review.line_id), review)
|
| 506 |
+
return next
|
| 507 |
+
})
|
| 508 |
+
}
|
| 509 |
+
onClose={() => setSelectedId(null)}
|
| 510 |
+
/>
|
| 511 |
+
</div>
|
| 512 |
+
)}
|
| 513 |
+
|
| 514 |
{/* Legend */}
|
| 515 |
<div className="px-4 py-2.5 border-t border-slate-700/40 flex items-center gap-6 flex-wrap">
|
| 516 |
<span className="font-mono text-[10px] text-slate-600 uppercase tracking-wider mr-1">
|
|
@@ -0,0 +1,186 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import { fireEvent, render, screen, waitFor } from '@testing-library/react'
|
| 2 |
+
import { beforeEach, describe, expect, it, vi } from 'vitest'
|
| 3 |
+
|
| 4 |
+
import type { LayoutLine, LineReview } from '../types'
|
| 5 |
+
import { ReviewPanel } from './ReviewPanel'
|
| 6 |
+
|
| 7 |
+
vi.mock('../api/client', () => ({
|
| 8 |
+
putReviews: vi.fn(),
|
| 9 |
+
}))
|
| 10 |
+
|
| 11 |
+
import { putReviews } from '../api/client'
|
| 12 |
+
|
| 13 |
+
const line = (over: Partial<LayoutLine> = {}): LayoutLine => ({
|
| 14 |
+
line_id: 'TL000323',
|
| 15 |
+
hpos: 0,
|
| 16 |
+
vpos: 0,
|
| 17 |
+
width: 100,
|
| 18 |
+
height: 20,
|
| 19 |
+
ocr_text: '1,500 ETANTS : LE NOMBRE DES MORTS',
|
| 20 |
+
corrected_text: '1,500 ETANTS : LE NOMBRE DES MORTS',
|
| 21 |
+
modified: false,
|
| 22 |
+
hyphen_role: 'none',
|
| 23 |
+
verdict: 'absorbs_previous_line',
|
| 24 |
+
verdict_detail: null,
|
| 25 |
+
proposed_text: "AU COURS D'UNE MATINÉE 1,500 ÉTANTS : LE NOMBRE DES MORTS",
|
| 26 |
+
proposal_declined: true,
|
| 27 |
+
...over,
|
| 28 |
+
})
|
| 29 |
+
|
| 30 |
+
function mount(over: Partial<LayoutLine> = {}, existing: LineReview | null = null) {
|
| 31 |
+
const onSaved = vi.fn()
|
| 32 |
+
render(
|
| 33 |
+
<ReviewPanel
|
| 34 |
+
jobId="job-1"
|
| 35 |
+
pageId="PAG_1"
|
| 36 |
+
line={line(over)}
|
| 37 |
+
existing={existing}
|
| 38 |
+
onSaved={onSaved}
|
| 39 |
+
onClose={vi.fn()}
|
| 40 |
+
/>,
|
| 41 |
+
)
|
| 42 |
+
return { onSaved }
|
| 43 |
+
}
|
| 44 |
+
|
| 45 |
+
describe('ReviewPanel', () => {
|
| 46 |
+
beforeEach(() => {
|
| 47 |
+
vi.mocked(putReviews).mockReset()
|
| 48 |
+
})
|
| 49 |
+
|
| 50 |
+
it('shows what was proposed and that it was declined', () => {
|
| 51 |
+
// The reviewer's most valuable case: something WAS on the table and a
|
| 52 |
+
// guard refused it. Hiding the proposal would hide whether the refusal
|
| 53 |
+
// saved a hallucination or threw away a good correction.
|
| 54 |
+
mount()
|
| 55 |
+
expect(screen.getByText(/proposé — et refusé/i)).toBeInTheDocument()
|
| 56 |
+
expect(screen.getByText(/AU COURS D'UNE MATINÉE/)).toBeInTheDocument()
|
| 57 |
+
})
|
| 58 |
+
|
| 59 |
+
it('refuses an empty transcription before it reaches the server', () => {
|
| 60 |
+
// The text IS the review. An empty one would land in the ground-truth set
|
| 61 |
+
// asserting nothing, indistinguishable later from a real reading.
|
| 62 |
+
mount()
|
| 63 |
+
fireEvent.click(screen.getByRole('button', { name: /ce que je lis/i }))
|
| 64 |
+
|
| 65 |
+
expect(screen.getByRole('alert')).toHaveTextContent(/n'affirme rien/i)
|
| 66 |
+
expect(putReviews).not.toHaveBeenCalled()
|
| 67 |
+
})
|
| 68 |
+
|
| 69 |
+
it('sends the transcription the reader typed', async () => {
|
| 70 |
+
vi.mocked(putReviews).mockResolvedValue([
|
| 71 |
+
{
|
| 72 |
+
page_id: 'PAG_1',
|
| 73 |
+
line_id: 'TL000323',
|
| 74 |
+
verdict: 'transcribed',
|
| 75 |
+
transcription: '1,500 ENFANTS : LE NOMBRE DES MORTS',
|
| 76 |
+
note: null,
|
| 77 |
+
reviewed_at: '2026-08-18T10:00:00+00:00',
|
| 78 |
+
},
|
| 79 |
+
])
|
| 80 |
+
const { onSaved } = mount()
|
| 81 |
+
|
| 82 |
+
fireEvent.change(screen.getByRole('textbox', { name: /ce que je lis sur le scan/i }), {
|
| 83 |
+
target: { value: '1,500 ENFANTS : LE NOMBRE DES MORTS' },
|
| 84 |
+
})
|
| 85 |
+
fireEvent.click(screen.getByRole('button', { name: /ce que je lis/i }))
|
| 86 |
+
|
| 87 |
+
await waitFor(() => expect(putReviews).toHaveBeenCalledTimes(1))
|
| 88 |
+
const [jobId, reviews] = vi.mocked(putReviews).mock.calls[0]
|
| 89 |
+
expect(jobId).toBe('job-1')
|
| 90 |
+
expect(reviews[0]).toMatchObject({
|
| 91 |
+
page_id: 'PAG_1',
|
| 92 |
+
line_id: 'TL000323',
|
| 93 |
+
verdict: 'transcribed',
|
| 94 |
+
transcription: '1,500 ENFANTS : LE NOMBRE DES MORTS',
|
| 95 |
+
})
|
| 96 |
+
await waitFor(() => expect(onSaved).toHaveBeenCalled())
|
| 97 |
+
})
|
| 98 |
+
|
| 99 |
+
it('grades the engine without needing a transcription', async () => {
|
| 100 |
+
vi.mocked(putReviews).mockResolvedValue([
|
| 101 |
+
{
|
| 102 |
+
page_id: 'PAG_1',
|
| 103 |
+
line_id: 'TL000323',
|
| 104 |
+
verdict: 'accepted',
|
| 105 |
+
transcription: null,
|
| 106 |
+
note: null,
|
| 107 |
+
reviewed_at: '2026-08-18T10:00:00+00:00',
|
| 108 |
+
},
|
| 109 |
+
])
|
| 110 |
+
mount()
|
| 111 |
+
fireEvent.click(screen.getByRole('button', { name: /le moteur a eu raison/i }))
|
| 112 |
+
await waitFor(() => expect(putReviews).toHaveBeenCalledTimes(1))
|
| 113 |
+
expect(vi.mocked(putReviews).mock.calls[0][1][0].verdict).toBe('accepted')
|
| 114 |
+
})
|
| 115 |
+
|
| 116 |
+
it("surfaces the server's own words when it rejects a review", async () => {
|
| 117 |
+
// A status code the reader cannot act on is worse than no feedback: the
|
| 118 |
+
// server explains an empty transcription by name, so show that.
|
| 119 |
+
vi.mocked(putReviews).mockRejectedValue(new Error('a transcribed review must carry the text'))
|
| 120 |
+
mount()
|
| 121 |
+
fireEvent.click(screen.getByRole('button', { name: /le moteur a eu tort/i }))
|
| 122 |
+
await waitFor(() => expect(screen.getByRole('alert')).toHaveTextContent(/must carry the text/i))
|
| 123 |
+
})
|
| 124 |
+
|
| 125 |
+
it('shows an existing judgement as the pressed one', () => {
|
| 126 |
+
mount(
|
| 127 |
+
{},
|
| 128 |
+
{
|
| 129 |
+
page_id: 'PAG_1',
|
| 130 |
+
line_id: 'TL000323',
|
| 131 |
+
verdict: 'refused',
|
| 132 |
+
transcription: null,
|
| 133 |
+
note: 'le scan dit ENFANTS',
|
| 134 |
+
reviewed_at: '2026-08-18T09:00:00+00:00',
|
| 135 |
+
},
|
| 136 |
+
)
|
| 137 |
+
expect(screen.getByRole('button', { name: /le moteur a eu tort/i })).toHaveAttribute(
|
| 138 |
+
'aria-pressed',
|
| 139 |
+
'true',
|
| 140 |
+
)
|
| 141 |
+
expect(screen.getByDisplayValue('le scan dit ENFANTS')).toBeInTheDocument()
|
| 142 |
+
})
|
| 143 |
+
})
|
| 144 |
+
|
| 145 |
+
describe('ReviewPanel — le scan en pleine résolution', () => {
|
| 146 |
+
it('shows the line from IIIF when a service is known', () => {
|
| 147 |
+
// Judging a word from a downscaled preview means judging ~13 pixels of
|
| 148 |
+
// newspaper line. The region request costs nothing to store and returns
|
| 149 |
+
// the scan's native pixels.
|
| 150 |
+
render(
|
| 151 |
+
<ReviewPanel
|
| 152 |
+
jobId="job-1"
|
| 153 |
+
pageId="PAG_1"
|
| 154 |
+
line={line({ hpos: 1000, vpos: 2000, width: 500, height: 100 })}
|
| 155 |
+
existing={null}
|
| 156 |
+
iiifService="https://example.org/iiif/f1"
|
| 157 |
+
onSaved={vi.fn()}
|
| 158 |
+
onClose={vi.fn()}
|
| 159 |
+
/>,
|
| 160 |
+
)
|
| 161 |
+
const img = screen.getByRole('img', { name: /ligne TL000323 sur le scan/i })
|
| 162 |
+
expect(img).toHaveAttribute(
|
| 163 |
+
'src',
|
| 164 |
+
'https://example.org/iiif/f1/995,1980,510,140/max/0/default.jpg',
|
| 165 |
+
)
|
| 166 |
+
})
|
| 167 |
+
|
| 168 |
+
it('falls back to text alone when no service is known', () => {
|
| 169 |
+
// The ALTO does not carry a IIIF identifier, so its absence is normal and
|
| 170 |
+
// must not cost the reviewer the rest of the panel.
|
| 171 |
+
render(
|
| 172 |
+
<ReviewPanel
|
| 173 |
+
jobId="job-1"
|
| 174 |
+
pageId="PAG_1"
|
| 175 |
+
line={line()}
|
| 176 |
+
existing={null}
|
| 177 |
+
onSaved={vi.fn()}
|
| 178 |
+
onClose={vi.fn()}
|
| 179 |
+
/>,
|
| 180 |
+
)
|
| 181 |
+
expect(screen.queryByRole('img')).not.toBeInTheDocument()
|
| 182 |
+
// The source and the retained text are identical on a refused line, so
|
| 183 |
+
// both nodes carry it — the point is that the panel still renders.
|
| 184 |
+
expect(screen.getAllByText(/1,500 ETANTS/).length).toBeGreaterThan(0)
|
| 185 |
+
})
|
| 186 |
+
})
|
|
@@ -0,0 +1,215 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import { useEffect, useState } from 'react'
|
| 2 |
+
|
| 3 |
+
import { putReviews } from '../api/client'
|
| 4 |
+
import { lineRegionUrl } from '../lib/iiif'
|
| 5 |
+
import type { LayoutLine, LineReview, ReviewVerdict } from '../types'
|
| 6 |
+
|
| 7 |
+
/**
|
| 8 |
+
* Where a reader's judgement on one line is made.
|
| 9 |
+
*
|
| 10 |
+
* The corpora this demo runs on carry no ground truth — Gallica's `text.txt`
|
| 11 |
+
* is the same OCR layer as its ALTO — so nothing on disk can say whether a
|
| 12 |
+
* correction was right or a refusal justified. Only a person reading the scan
|
| 13 |
+
* can, and this panel is where that reading is captured.
|
| 14 |
+
*
|
| 15 |
+
* Three verdicts, and the third is the one that compounds. `accepted` and
|
| 16 |
+
* `refused` grade what the engine did, which is useful for tuning it.
|
| 17 |
+
* `transcribed` records what the reader actually read, which stands on its own
|
| 18 |
+
* whatever the engine decided — and accumulates, line by line, into the ground
|
| 19 |
+
* truth the bench needs to answer "was the correction right".
|
| 20 |
+
*/
|
| 21 |
+
|
| 22 |
+
interface ReviewPanelProps {
|
| 23 |
+
jobId: string
|
| 24 |
+
pageId: string
|
| 25 |
+
line: LayoutLine
|
| 26 |
+
/** The reader's existing judgement on this line, if any. */
|
| 27 |
+
existing: LineReview | null
|
| 28 |
+
/**
|
| 29 |
+
* IIIF Image API service base for this page, when the reader has one.
|
| 30 |
+
* With it, the line is shown at the scan's NATIVE resolution and nothing is
|
| 31 |
+
* stored — the alternative is judging a word from a downscaled preview,
|
| 32 |
+
* which on a newspaper line is roughly 13 pixels tall.
|
| 33 |
+
*/
|
| 34 |
+
iiifService?: string | null
|
| 35 |
+
onSaved: (review: LineReview) => void
|
| 36 |
+
onClose: () => void
|
| 37 |
+
}
|
| 38 |
+
|
| 39 |
+
const VERDICT_LABEL: Record<ReviewVerdict, string> = {
|
| 40 |
+
accepted: 'Le moteur a eu raison',
|
| 41 |
+
refused: 'Le moteur a eu tort',
|
| 42 |
+
transcribed: 'Voici ce que je lis',
|
| 43 |
+
}
|
| 44 |
+
|
| 45 |
+
export function ReviewPanel({
|
| 46 |
+
jobId,
|
| 47 |
+
pageId,
|
| 48 |
+
line,
|
| 49 |
+
existing,
|
| 50 |
+
iiifService,
|
| 51 |
+
onSaved,
|
| 52 |
+
onClose,
|
| 53 |
+
}: ReviewPanelProps) {
|
| 54 |
+
const [transcription, setTranscription] = useState(existing?.transcription ?? '')
|
| 55 |
+
const [note, setNote] = useState(existing?.note ?? '')
|
| 56 |
+
const [saving, setSaving] = useState(false)
|
| 57 |
+
const [error, setError] = useState<string | null>(null)
|
| 58 |
+
|
| 59 |
+
// A reader moves line to line; the fields must follow the selection rather
|
| 60 |
+
// than carry the previous line's text into the next one's judgement.
|
| 61 |
+
useEffect(() => {
|
| 62 |
+
setTranscription(existing?.transcription ?? '')
|
| 63 |
+
setNote(existing?.note ?? '')
|
| 64 |
+
setError(null)
|
| 65 |
+
}, [line.line_id, pageId, existing])
|
| 66 |
+
|
| 67 |
+
async function save(verdict: ReviewVerdict) {
|
| 68 |
+
if (verdict === 'transcribed' && !transcription.trim()) {
|
| 69 |
+
setError("Une transcription vide n'affirme rien : écrivez ce que vous lisez sur le scan.")
|
| 70 |
+
return
|
| 71 |
+
}
|
| 72 |
+
setSaving(true)
|
| 73 |
+
setError(null)
|
| 74 |
+
try {
|
| 75 |
+
const review: LineReview = {
|
| 76 |
+
page_id: pageId,
|
| 77 |
+
line_id: line.line_id,
|
| 78 |
+
verdict,
|
| 79 |
+
transcription: transcription.trim() || null,
|
| 80 |
+
note: note.trim() || null,
|
| 81 |
+
}
|
| 82 |
+
const saved = await putReviews(jobId, [review])
|
| 83 |
+
const mine = saved.find((r) => r.page_id === pageId && r.line_id === line.line_id)
|
| 84 |
+
if (mine) onSaved(mine)
|
| 85 |
+
} catch (caught) {
|
| 86 |
+
setError(caught instanceof Error ? caught.message : String(caught))
|
| 87 |
+
} finally {
|
| 88 |
+
setSaving(false)
|
| 89 |
+
}
|
| 90 |
+
}
|
| 91 |
+
|
| 92 |
+
const declined = line.proposal_declined
|
| 93 |
+
|
| 94 |
+
return (
|
| 95 |
+
<aside
|
| 96 |
+
className="border border-slate-700 bg-slate-800 rounded p-4 flex flex-col gap-3"
|
| 97 |
+
aria-label={`Jugement sur la ligne ${line.line_id}`}
|
| 98 |
+
>
|
| 99 |
+
<header className="flex items-baseline justify-between gap-3">
|
| 100 |
+
<span className="font-mono text-[10px] uppercase tracking-wider text-slate-400">
|
| 101 |
+
{line.line_id}
|
| 102 |
+
</span>
|
| 103 |
+
<button
|
| 104 |
+
type="button"
|
| 105 |
+
onClick={onClose}
|
| 106 |
+
className="font-mono text-[10px] text-slate-400 hover:text-slate-200"
|
| 107 |
+
>
|
| 108 |
+
fermer
|
| 109 |
+
</button>
|
| 110 |
+
</header>
|
| 111 |
+
|
| 112 |
+
{iiifService && (
|
| 113 |
+
<figure className="m-0">
|
| 114 |
+
<img
|
| 115 |
+
src={lineRegionUrl(iiifService, line)}
|
| 116 |
+
alt={`Ligne ${line.line_id} sur le scan`}
|
| 117 |
+
className="w-full rounded border border-slate-600 bg-white"
|
| 118 |
+
/>
|
| 119 |
+
<figcaption className="font-mono text-[10px] text-slate-500 mt-1">
|
| 120 |
+
résolution native, servie par IIIF — rien n'est stocké
|
| 121 |
+
</figcaption>
|
| 122 |
+
</figure>
|
| 123 |
+
)}
|
| 124 |
+
|
| 125 |
+
<dl className="flex flex-col gap-2 text-xs">
|
| 126 |
+
<div>
|
| 127 |
+
<dt className="font-mono text-[10px] uppercase tracking-wider text-slate-500">
|
| 128 |
+
OCR source
|
| 129 |
+
</dt>
|
| 130 |
+
<dd className="font-mono text-slate-200 break-words">{line.ocr_text}</dd>
|
| 131 |
+
</div>
|
| 132 |
+
{line.proposed_text && line.proposed_text !== line.ocr_text && (
|
| 133 |
+
<div>
|
| 134 |
+
<dt className="font-mono text-[10px] uppercase tracking-wider text-slate-500">
|
| 135 |
+
Proposé{declined ? ' — et refusé' : ''}
|
| 136 |
+
</dt>
|
| 137 |
+
<dd className={`font-mono break-words ${declined ? 'text-red-300' : 'text-blue-300'}`}>
|
| 138 |
+
{line.proposed_text}
|
| 139 |
+
</dd>
|
| 140 |
+
</div>
|
| 141 |
+
)}
|
| 142 |
+
<div>
|
| 143 |
+
<dt className="font-mono text-[10px] uppercase tracking-wider text-slate-500">Retenu</dt>
|
| 144 |
+
<dd className="font-mono text-slate-100 break-words">{line.corrected_text}</dd>
|
| 145 |
+
</div>
|
| 146 |
+
{line.verdict && (
|
| 147 |
+
<div>
|
| 148 |
+
<dt className="font-mono text-[10px] uppercase tracking-wider text-slate-500">
|
| 149 |
+
Verdict du moteur
|
| 150 |
+
</dt>
|
| 151 |
+
<dd className="font-mono text-amber-300 break-words">
|
| 152 |
+
{line.verdict}
|
| 153 |
+
{line.verdict_detail ? ` — ${line.verdict_detail}` : ''}
|
| 154 |
+
</dd>
|
| 155 |
+
</div>
|
| 156 |
+
)}
|
| 157 |
+
</dl>
|
| 158 |
+
|
| 159 |
+
<label className="flex flex-col gap-1">
|
| 160 |
+
<span className="font-mono text-[10px] uppercase tracking-wider text-slate-500">
|
| 161 |
+
Ce que je lis sur le scan
|
| 162 |
+
</span>
|
| 163 |
+
<textarea
|
| 164 |
+
value={transcription}
|
| 165 |
+
onChange={(e) => setTranscription(e.target.value)}
|
| 166 |
+
rows={2}
|
| 167 |
+
spellCheck={false}
|
| 168 |
+
className="font-mono text-xs bg-slate-900 border border-slate-600 text-slate-100
|
| 169 |
+
rounded px-2 py-1 focus:outline-none focus:border-amber-500"
|
| 170 |
+
/>
|
| 171 |
+
</label>
|
| 172 |
+
|
| 173 |
+
<label className="flex flex-col gap-1">
|
| 174 |
+
<span className="font-mono text-[10px] uppercase tracking-wider text-slate-500">Note</span>
|
| 175 |
+
<input
|
| 176 |
+
value={note}
|
| 177 |
+
onChange={(e) => setNote(e.target.value)}
|
| 178 |
+
className="font-mono text-xs bg-slate-900 border border-slate-600 text-slate-100
|
| 179 |
+
rounded px-2 py-1 focus:outline-none focus:border-amber-500"
|
| 180 |
+
/>
|
| 181 |
+
</label>
|
| 182 |
+
|
| 183 |
+
{error && (
|
| 184 |
+
<p role="alert" className="text-xs text-red-300">
|
| 185 |
+
{error}
|
| 186 |
+
</p>
|
| 187 |
+
)}
|
| 188 |
+
|
| 189 |
+
<div className="flex flex-wrap gap-2">
|
| 190 |
+
{(['accepted', 'refused', 'transcribed'] as const).map((verdict) => (
|
| 191 |
+
<button
|
| 192 |
+
key={verdict}
|
| 193 |
+
type="button"
|
| 194 |
+
disabled={saving}
|
| 195 |
+
onClick={() => save(verdict)}
|
| 196 |
+
aria-pressed={existing?.verdict === verdict}
|
| 197 |
+
className={`font-mono text-[11px] rounded px-3 py-1.5 border transition-colors
|
| 198 |
+
disabled:opacity-50 focus:outline-none focus:ring-1
|
| 199 |
+
focus:ring-amber-400 ${
|
| 200 |
+
existing?.verdict === verdict
|
| 201 |
+
? 'bg-amber-500 border-amber-400 text-slate-900'
|
| 202 |
+
: 'bg-slate-700 border-slate-600 text-slate-200 hover:border-amber-500'
|
| 203 |
+
}`}
|
| 204 |
+
>
|
| 205 |
+
{VERDICT_LABEL[verdict]}
|
| 206 |
+
</button>
|
| 207 |
+
))}
|
| 208 |
+
</div>
|
| 209 |
+
|
| 210 |
+
{existing?.reviewed_at && (
|
| 211 |
+
<p className="font-mono text-[10px] text-slate-500">Jugé le {existing.reviewed_at}</p>
|
| 212 |
+
)}
|
| 213 |
+
</aside>
|
| 214 |
+
)
|
| 215 |
+
}
|
|
@@ -0,0 +1,66 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import { describe, expect, it } from 'vitest'
|
| 2 |
+
|
| 3 |
+
import { lineRegionUrl, looksLikeService } from './iiif'
|
| 4 |
+
|
| 5 |
+
describe('lineRegionUrl', () => {
|
| 6 |
+
it('passes ALTO coordinates through as the IIIF region', () => {
|
| 7 |
+
// The two spaces are the same one — a service serving the full scan
|
| 8 |
+
// declares the page's own dimensions — so no scaling belongs here.
|
| 9 |
+
const url = lineRegionUrl(
|
| 10 |
+
'https://example.org/iiif/f1',
|
| 11 |
+
{
|
| 12 |
+
hpos: 1000,
|
| 13 |
+
vpos: 2000,
|
| 14 |
+
width: 500,
|
| 15 |
+
height: 100,
|
| 16 |
+
},
|
| 17 |
+
0,
|
| 18 |
+
)
|
| 19 |
+
expect(url).toBe('https://example.org/iiif/f1/995,2000,510,100/max/0/default.jpg')
|
| 20 |
+
})
|
| 21 |
+
|
| 22 |
+
it('asks for max rather than a fixed width', () => {
|
| 23 |
+
// Measured against Gallica: a fixed width 400s the moment it would
|
| 24 |
+
// upscale, and lines are ~70px tall so it almost always would.
|
| 25 |
+
const url = lineRegionUrl('https://example.org/iiif/f1', {
|
| 26 |
+
hpos: 0,
|
| 27 |
+
vpos: 0,
|
| 28 |
+
width: 100,
|
| 29 |
+
height: 20,
|
| 30 |
+
})
|
| 31 |
+
expect(url).toContain('/max/0/default.jpg')
|
| 32 |
+
expect(url).not.toMatch(/\/\d+,\/0\//)
|
| 33 |
+
})
|
| 34 |
+
|
| 35 |
+
it('tolerates a trailing slash on the service', () => {
|
| 36 |
+
const url = lineRegionUrl(
|
| 37 |
+
'https://example.org/iiif/f1/',
|
| 38 |
+
{
|
| 39 |
+
hpos: 0,
|
| 40 |
+
vpos: 0,
|
| 41 |
+
width: 10,
|
| 42 |
+
height: 10,
|
| 43 |
+
},
|
| 44 |
+
0,
|
| 45 |
+
)
|
| 46 |
+
expect(url).not.toContain('f1//')
|
| 47 |
+
})
|
| 48 |
+
})
|
| 49 |
+
|
| 50 |
+
describe('looksLikeService', () => {
|
| 51 |
+
it('accepts a service base', () => {
|
| 52 |
+
expect(looksLikeService('https://openapi.bnf.fr/iiif/image/v3/ark:/12148/x/f1')).toBe(true)
|
| 53 |
+
})
|
| 54 |
+
|
| 55 |
+
it('rejects a full image URL, which would nest two region paths', () => {
|
| 56 |
+
expect(
|
| 57 |
+
looksLikeService(
|
| 58 |
+
'https://openapi.bnf.fr/iiif/image/v3/ark:/12148/x/f1/full/max/0/default.jpg',
|
| 59 |
+
),
|
| 60 |
+
).toBe(false)
|
| 61 |
+
})
|
| 62 |
+
|
| 63 |
+
it('rejects anything that is not a URL', () => {
|
| 64 |
+
expect(looksLikeService('bpt6k4607951t')).toBe(false)
|
| 65 |
+
})
|
| 66 |
+
})
|
|
@@ -0,0 +1,50 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
/**
|
| 2 |
+
* IIIF Image API region URLs — a line's own pixels, without storing a thing.
|
| 3 |
+
*
|
| 4 |
+
* ALTO coordinates and the IIIF region parameter live in the SAME space: a
|
| 5 |
+
* service serving the full scan declares exactly the page's own dimensions
|
| 6 |
+
* (measured on Gallica: `info.json` says 6802×9121 for an ALTO page of
|
| 7 |
+
* 6802×9121). So the coordinates go through verbatim — no scale, no
|
| 8 |
+
* transform, and no local derivative to crop.
|
| 9 |
+
*
|
| 10 |
+
* That matters beyond convenience. Cropping a DOWNSCALED derivative locally
|
| 11 |
+
* is how a line ends up cut from the wrong place: Gallica's `!1600,1600`
|
| 12 |
+
* preview is 17.5% of the ALTO space, so a line at hpos=4798 falls off an
|
| 13 |
+
* image 1193 pixels wide, and the crop comes back blank with nothing
|
| 14 |
+
* downstream any the wiser.
|
| 15 |
+
*/
|
| 16 |
+
|
| 17 |
+
export interface Region {
|
| 18 |
+
hpos: number
|
| 19 |
+
vpos: number
|
| 20 |
+
width: number
|
| 21 |
+
height: number
|
| 22 |
+
}
|
| 23 |
+
|
| 24 |
+
/**
|
| 25 |
+
* `size` is `max`, deliberately.
|
| 26 |
+
*
|
| 27 |
+
* Asking for a fixed width is refused with HTTP 400 the moment it would
|
| 28 |
+
* UPSCALE the region — measured against Gallica, whose lines are ~70px tall:
|
| 29 |
+
* `/1200,/` returns 400, `/max/` returns 200. IIIF 3.0 requires an explicit
|
| 30 |
+
* `^` prefix to allow upscaling, and a line does not need it: its native
|
| 31 |
+
* pixels are exactly what a reader wants to see.
|
| 32 |
+
*/
|
| 33 |
+
export function lineRegionUrl(service: string, region: Region, marginRatio = 0.2): string {
|
| 34 |
+
const marginX = Math.round(region.width * 0.01)
|
| 35 |
+
const marginY = Math.round(region.height * marginRatio)
|
| 36 |
+
const x = Math.max(region.hpos - marginX, 0)
|
| 37 |
+
const y = Math.max(region.vpos - marginY, 0)
|
| 38 |
+
const w = region.width + 2 * marginX
|
| 39 |
+
const h = region.height + 2 * marginY
|
| 40 |
+
return `${service.replace(/\/+$/, '')}/${x},${y},${w},${h}/max/0/default.jpg`
|
| 41 |
+
}
|
| 42 |
+
|
| 43 |
+
/** Whether a pasted string looks like an Image API service base, not a full image URL. */
|
| 44 |
+
export function looksLikeService(value: string): boolean {
|
| 45 |
+
const trimmed = value.trim()
|
| 46 |
+
if (!/^https?:\/\//i.test(trimmed)) return false
|
| 47 |
+
// A full image URL already carries region/size/rotation/quality; using it as
|
| 48 |
+
// a base would produce nonsense like `.../full/max/0/default.jpg/12,3,4,5/max/...`.
|
| 49 |
+
return !/\/(full|square|\d+,\d+,\d+,\d+)\/[^/]+\/\d+\/[^/]+\.\w+$/i.test(trimmed)
|
| 50 |
+
}
|
|
@@ -0,0 +1,45 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
// ---------------------------------------------------------------------------
|
| 2 |
+
// Verdict palette — what the engine decided, and why
|
| 3 |
+
// ---------------------------------------------------------------------------
|
| 4 |
+
//
|
| 5 |
+
// Three families, because a reviewer scans for three different things and
|
| 6 |
+
// should not have to read a legend to tell them apart:
|
| 7 |
+
//
|
| 8 |
+
// kept the correction survived every guard
|
| 9 |
+
// refused something was proposed and a guard declined it — the cases
|
| 10 |
+
// worth a human eye, since each is either a caught hallucination
|
| 11 |
+
// or a good correction thrown away
|
| 12 |
+
// silent nothing was proposed, or nothing changed; no judgement to make
|
| 13 |
+
//
|
| 14 |
+
// Colours are the proof-reader's two pencils: blue marks what stands, red
|
| 15 |
+
// marks what was struck out. Amber sits between them for a line the engine
|
| 16 |
+
// never got an answer for.
|
| 17 |
+
|
| 18 |
+
export type VerdictFamily = 'kept' | 'refused' | 'silent'
|
| 19 |
+
|
| 20 |
+
const REFUSAL_CODES = new Set([
|
| 21 |
+
'too_different_from_source',
|
| 22 |
+
'closer_to_previous_line',
|
| 23 |
+
'closer_to_next_line',
|
| 24 |
+
'absorbs_previous_line',
|
| 25 |
+
'absorbs_next_line',
|
| 26 |
+
'hyphen_pair_fallback',
|
| 27 |
+
'boundary_migration_forward',
|
| 28 |
+
'boundary_migration_backward',
|
| 29 |
+
'adjacent_duplicate_detected',
|
| 30 |
+
'adjacent_duplicate_pair_atomicity',
|
| 31 |
+
'orphan_hyphen_completed',
|
| 32 |
+
'hyphen_unit_fallback',
|
| 33 |
+
])
|
| 34 |
+
|
| 35 |
+
export function verdictFamily(line: { verdict: string | null; modified: boolean }): VerdictFamily {
|
| 36 |
+
if (line.verdict && REFUSAL_CODES.has(line.verdict)) return 'refused'
|
| 37 |
+
if (line.verdict === 'all_attempts_exhausted') return 'silent'
|
| 38 |
+
return line.modified ? 'kept' : 'silent'
|
| 39 |
+
}
|
| 40 |
+
|
| 41 |
+
export const FAMILY = {
|
| 42 |
+
kept: { stroke: '#1d4ed8', fill: 'rgba(29,78,216,0.18)', label: 'Retenue' },
|
| 43 |
+
refused: { stroke: '#b91c1c', fill: 'rgba(185,28,28,0.20)', label: 'Refusée' },
|
| 44 |
+
silent: { stroke: '#a1a1aa', fill: 'rgba(161,161,170,0.10)', label: 'Sans objet' },
|
| 45 |
+
} as const
|
|
@@ -289,6 +289,32 @@ export interface paths {
|
|
| 289 |
patch?: never
|
| 290 |
trace?: never
|
| 291 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 292 |
}
|
| 293 |
export type webhooks = Record<string, never>
|
| 294 |
export interface components {
|
|
@@ -397,6 +423,23 @@ export interface components {
|
|
| 397 |
/** Error */
|
| 398 |
error?: string | null
|
| 399 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 400 |
/** ListModelsRequest */
|
| 401 |
ListModelsRequest: {
|
| 402 |
provider: components['schemas']['Provider']
|
|
@@ -432,6 +475,27 @@ export interface components {
|
|
| 432 |
* @enum {string}
|
| 433 |
*/
|
| 434 |
Provider: 'openai' | 'anthropic' | 'mistral' | 'google'
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 435 |
/** ValidationError */
|
| 436 |
ValidationError: {
|
| 437 |
/** Location */
|
|
@@ -837,4 +901,70 @@ export interface operations {
|
|
| 837 |
}
|
| 838 |
}
|
| 839 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 840 |
}
|
|
|
|
| 289 |
patch?: never
|
| 290 |
trace?: never
|
| 291 |
}
|
| 292 |
+
'/api/jobs/{job_id}/reviews': {
|
| 293 |
+
parameters: {
|
| 294 |
+
query?: never
|
| 295 |
+
header?: never
|
| 296 |
+
path?: never
|
| 297 |
+
cookie?: never
|
| 298 |
+
}
|
| 299 |
+
/** Get Reviews */
|
| 300 |
+
get: operations['get_reviews_api_jobs__job_id__reviews_get']
|
| 301 |
+
/**
|
| 302 |
+
* Put Reviews
|
| 303 |
+
* @description Record or replace judgements on lines of this job.
|
| 304 |
+
*
|
| 305 |
+
* Idempotent per line: sending the same line twice replaces its review
|
| 306 |
+
* rather than appending, so a reader who changes their mind is not fighting
|
| 307 |
+
* an append-only log. The timestamp is stamped here rather than trusted
|
| 308 |
+
* from the client — a review's date is a fact about the server.
|
| 309 |
+
*/
|
| 310 |
+
put: operations['put_reviews_api_jobs__job_id__reviews_put']
|
| 311 |
+
post?: never
|
| 312 |
+
delete?: never
|
| 313 |
+
options?: never
|
| 314 |
+
head?: never
|
| 315 |
+
patch?: never
|
| 316 |
+
trace?: never
|
| 317 |
+
}
|
| 318 |
}
|
| 319 |
export type webhooks = Record<string, never>
|
| 320 |
export interface components {
|
|
|
|
| 423 |
/** Error */
|
| 424 |
error?: string | null
|
| 425 |
}
|
| 426 |
+
/**
|
| 427 |
+
* LineReview
|
| 428 |
+
* @description One reader's judgement on one line.
|
| 429 |
+
*/
|
| 430 |
+
LineReview: {
|
| 431 |
+
/** Page Id */
|
| 432 |
+
page_id: string
|
| 433 |
+
/** Line Id */
|
| 434 |
+
line_id: string
|
| 435 |
+
verdict: components['schemas']['ReviewVerdict']
|
| 436 |
+
/** Transcription */
|
| 437 |
+
transcription?: string | null
|
| 438 |
+
/** Note */
|
| 439 |
+
note?: string | null
|
| 440 |
+
/** Reviewed At */
|
| 441 |
+
reviewed_at?: string | null
|
| 442 |
+
}
|
| 443 |
/** ListModelsRequest */
|
| 444 |
ListModelsRequest: {
|
| 445 |
provider: components['schemas']['Provider']
|
|
|
|
| 475 |
* @enum {string}
|
| 476 |
*/
|
| 477 |
Provider: 'openai' | 'anthropic' | 'mistral' | 'google'
|
| 478 |
+
/**
|
| 479 |
+
* ReviewBatch
|
| 480 |
+
* @description Reviews arrive in batches: a reader works through a page, not a line.
|
| 481 |
+
*/
|
| 482 |
+
ReviewBatch: {
|
| 483 |
+
/** Reviews */
|
| 484 |
+
reviews: components['schemas']['LineReview'][]
|
| 485 |
+
}
|
| 486 |
+
/**
|
| 487 |
+
* ReviewVerdict
|
| 488 |
+
* @description What the reader concluded about the engine's decision on this line.
|
| 489 |
+
* @enum {string}
|
| 490 |
+
*/
|
| 491 |
+
ReviewVerdict: 'accepted' | 'refused' | 'transcribed'
|
| 492 |
+
/** ReviewsResponse */
|
| 493 |
+
ReviewsResponse: {
|
| 494 |
+
/** Job Id */
|
| 495 |
+
job_id: string
|
| 496 |
+
/** Reviews */
|
| 497 |
+
reviews: components['schemas']['LineReview'][]
|
| 498 |
+
}
|
| 499 |
/** ValidationError */
|
| 500 |
ValidationError: {
|
| 501 |
/** Location */
|
|
|
|
| 901 |
}
|
| 902 |
}
|
| 903 |
}
|
| 904 |
+
get_reviews_api_jobs__job_id__reviews_get: {
|
| 905 |
+
parameters: {
|
| 906 |
+
query?: never
|
| 907 |
+
header?: never
|
| 908 |
+
path: {
|
| 909 |
+
job_id: string
|
| 910 |
+
}
|
| 911 |
+
cookie?: never
|
| 912 |
+
}
|
| 913 |
+
requestBody?: never
|
| 914 |
+
responses: {
|
| 915 |
+
/** @description Successful Response */
|
| 916 |
+
200: {
|
| 917 |
+
headers: {
|
| 918 |
+
[name: string]: unknown
|
| 919 |
+
}
|
| 920 |
+
content: {
|
| 921 |
+
'application/json': components['schemas']['ReviewsResponse']
|
| 922 |
+
}
|
| 923 |
+
}
|
| 924 |
+
/** @description Validation Error */
|
| 925 |
+
422: {
|
| 926 |
+
headers: {
|
| 927 |
+
[name: string]: unknown
|
| 928 |
+
}
|
| 929 |
+
content: {
|
| 930 |
+
'application/json': components['schemas']['HTTPValidationError']
|
| 931 |
+
}
|
| 932 |
+
}
|
| 933 |
+
}
|
| 934 |
+
}
|
| 935 |
+
put_reviews_api_jobs__job_id__reviews_put: {
|
| 936 |
+
parameters: {
|
| 937 |
+
query?: never
|
| 938 |
+
header?: never
|
| 939 |
+
path: {
|
| 940 |
+
job_id: string
|
| 941 |
+
}
|
| 942 |
+
cookie?: never
|
| 943 |
+
}
|
| 944 |
+
requestBody: {
|
| 945 |
+
content: {
|
| 946 |
+
'application/json': components['schemas']['ReviewBatch']
|
| 947 |
+
}
|
| 948 |
+
}
|
| 949 |
+
responses: {
|
| 950 |
+
/** @description Successful Response */
|
| 951 |
+
200: {
|
| 952 |
+
headers: {
|
| 953 |
+
[name: string]: unknown
|
| 954 |
+
}
|
| 955 |
+
content: {
|
| 956 |
+
'application/json': components['schemas']['ReviewsResponse']
|
| 957 |
+
}
|
| 958 |
+
}
|
| 959 |
+
/** @description Validation Error */
|
| 960 |
+
422: {
|
| 961 |
+
headers: {
|
| 962 |
+
[name: string]: unknown
|
| 963 |
+
}
|
| 964 |
+
content: {
|
| 965 |
+
'application/json': components['schemas']['HTTPValidationError']
|
| 966 |
+
}
|
| 967 |
+
}
|
| 968 |
+
}
|
| 969 |
+
}
|
| 970 |
}
|
|
@@ -211,6 +211,31 @@ export interface LayoutLine {
|
|
| 211 |
corrected_text: string
|
| 212 |
modified: boolean
|
| 213 |
hyphen_role: 'none' | 'HypPart1' | 'HypPart2'
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 214 |
}
|
| 215 |
|
| 216 |
export interface LayoutBlock {
|
|
|
|
| 211 |
corrected_text: string
|
| 212 |
modified: boolean
|
| 213 |
hyphen_role: 'none' | 'HypPart1' | 'HypPart2'
|
| 214 |
+
/**
|
| 215 |
+
* Why the line ended up as it did — the guard code, or the decision status
|
| 216 |
+
* when no guard spoke. Geometry alone cannot distinguish "nothing was
|
| 217 |
+
* proposed" from "a hallucination was refused", and a reviewer needs to.
|
| 218 |
+
* `null` when the job carries no report yet.
|
| 219 |
+
*/
|
| 220 |
+
verdict: string | null
|
| 221 |
+
verdict_detail: string | null
|
| 222 |
+
/** What the producer actually returned, before the engine judged it. */
|
| 223 |
+
proposed_text: string | null
|
| 224 |
+
/** A proposal was on the table and the engine declined it — the case worth reading. */
|
| 225 |
+
proposal_declined: boolean
|
| 226 |
+
}
|
| 227 |
+
|
| 228 |
+
/** How a reader graded the engine on one line. */
|
| 229 |
+
export type ReviewVerdict = 'accepted' | 'refused' | 'transcribed'
|
| 230 |
+
|
| 231 |
+
export interface LineReview {
|
| 232 |
+
page_id: string
|
| 233 |
+
line_id: string
|
| 234 |
+
verdict: ReviewVerdict
|
| 235 |
+
/** What the reader read on the scan. Required for `transcribed`. */
|
| 236 |
+
transcription?: string | null
|
| 237 |
+
note?: string | null
|
| 238 |
+
reviewed_at?: string | null
|
| 239 |
}
|
| 240 |
|
| 241 |
export interface LayoutBlock {
|
|
@@ -0,0 +1,261 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Run a corpus through the demo's own JobRunner, one arm at a time.
|
| 2 |
+
|
| 3 |
+
python tools/run_corpus_arms.py <corpus-root> <arm> [concurrency]
|
| 4 |
+
|
| 5 |
+
Arms:
|
| 6 |
+
|
| 7 |
+
small-text mistral-small-latest, no image
|
| 8 |
+
small-vision mistral-small-latest, per-line crops
|
| 9 |
+
ministral-text ministral-8b-latest, no image
|
| 10 |
+
ministral-vision ministral-8b-latest, per-line crops
|
| 11 |
+
|
| 12 |
+
`<corpus-root>` is scanned for ``alto.xml`` files, each with an optional
|
| 13 |
+
``image.jpg`` beside it — the layout a digital library's page-by-page export
|
| 14 |
+
produces, and the shape ``gallica_multicolumn`` writes.
|
| 15 |
+
|
| 16 |
+
This drives ``JobRunner`` rather than the HTTP API: same producer assembly,
|
| 17 |
+
same guards, same output writer, no server to stand up. The key is read from
|
| 18 |
+
the macOS keychain and never appears in a flag, an environment dump, or an
|
| 19 |
+
artefact — a flag would land in shell history and in the process list.
|
| 20 |
+
|
| 21 |
+
One JSON per page under ``<corpus-root>/../arm-results/<arm>/``, so an
|
| 22 |
+
interrupted run resumes and nothing measured is lost.
|
| 23 |
+
"""
|
| 24 |
+
|
| 25 |
+
from __future__ import annotations
|
| 26 |
+
|
| 27 |
+
import asyncio
|
| 28 |
+
import json
|
| 29 |
+
import subprocess
|
| 30 |
+
import sys
|
| 31 |
+
import time
|
| 32 |
+
from collections import Counter
|
| 33 |
+
from pathlib import Path
|
| 34 |
+
|
| 35 |
+
sys.path.insert(0, str(Path(__file__).resolve().parent.parent / "backend"))
|
| 36 |
+
|
| 37 |
+
from saknussemm.formats.loader import build_document_manifest
|
| 38 |
+
|
| 39 |
+
from app.jobs.runner import JobRunner, page_image_assets
|
| 40 |
+
from app.jobs.store import JobStore
|
| 41 |
+
from app.providers.mistral_multimodal import MistralMultimodalProvider
|
| 42 |
+
from app.providers.mistral_provider import MistralProvider
|
| 43 |
+
from app.storage.output_writer import FilesystemOutputWriter
|
| 44 |
+
|
| 45 |
+
ARMS: dict[str, tuple[str, bool]] = {
|
| 46 |
+
"small-text": ("mistral-small-latest", False),
|
| 47 |
+
"small-vision": ("mistral-small-latest", True),
|
| 48 |
+
"ministral-text": ("ministral-8b-latest", False),
|
| 49 |
+
"ministral-vision": ("ministral-8b-latest", True),
|
| 50 |
+
}
|
| 51 |
+
|
| 52 |
+
|
| 53 |
+
def keychain_key() -> str:
|
| 54 |
+
return subprocess.run(
|
| 55 |
+
["security", "find-generic-password", "-a", "marcel", "-s", "mistral-api-key", "-w"],
|
| 56 |
+
capture_output=True,
|
| 57 |
+
text=True,
|
| 58 |
+
check=True,
|
| 59 |
+
).stdout.strip()
|
| 60 |
+
|
| 61 |
+
|
| 62 |
+
def find_pages(root: Path) -> list[tuple[str, Path, Path | None]]:
|
| 63 |
+
"""``(label, alto path, image path or None)`` for every page under *root*."""
|
| 64 |
+
pages = []
|
| 65 |
+
for alto in sorted(root.rglob("alto.xml")):
|
| 66 |
+
label = f"{alto.parent.parent.name}_{alto.parent.name.split('_')[-1]}"
|
| 67 |
+
image = alto.parent / "image.jpg"
|
| 68 |
+
pages.append((label, alto, image if image.exists() else None))
|
| 69 |
+
return pages
|
| 70 |
+
|
| 71 |
+
|
| 72 |
+
async def run_one(
|
| 73 |
+
arm: str,
|
| 74 |
+
label: str,
|
| 75 |
+
alto: Path,
|
| 76 |
+
image: Path | None,
|
| 77 |
+
api_key: str,
|
| 78 |
+
results: Path,
|
| 79 |
+
sem: asyncio.Semaphore,
|
| 80 |
+
) -> dict | None:
|
| 81 |
+
target = results / f"{label}.json"
|
| 82 |
+
if target.exists():
|
| 83 |
+
return json.loads(target.read_text())
|
| 84 |
+
|
| 85 |
+
model, wants_image = ARMS[arm]
|
| 86 |
+
if wants_image and image is None:
|
| 87 |
+
return None
|
| 88 |
+
|
| 89 |
+
async with sem:
|
| 90 |
+
from app.schemas import Provider
|
| 91 |
+
|
| 92 |
+
store = JobStore()
|
| 93 |
+
job_id = store.create_job(provider=Provider.MISTRAL, model=model)
|
| 94 |
+
manifest = build_document_manifest([(alto, alto.name)])
|
| 95 |
+
|
| 96 |
+
images = None
|
| 97 |
+
if wants_image:
|
| 98 |
+
images = page_image_assets(manifest, {alto.parent.name.lower(): image})
|
| 99 |
+
if not images:
|
| 100 |
+
# The demo keys images by source-file stem; this corpus names
|
| 101 |
+
# every file `alto.xml`, so map it explicitly by page instead of
|
| 102 |
+
# letting the helper's stem lookup miss silently.
|
| 103 |
+
import hashlib
|
| 104 |
+
|
| 105 |
+
from saknussemm.core.schemas import ImageAsset
|
| 106 |
+
|
| 107 |
+
digest = hashlib.sha256(image.read_bytes()).hexdigest()
|
| 108 |
+
images = {
|
| 109 |
+
page.page_id: ImageAsset(
|
| 110 |
+
page_id=page.page_id,
|
| 111 |
+
uri=str(image),
|
| 112 |
+
sha256=digest,
|
| 113 |
+
media_type="image/jpeg",
|
| 114 |
+
)
|
| 115 |
+
for page in manifest.pages
|
| 116 |
+
}
|
| 117 |
+
|
| 118 |
+
out_dir = results / "outputs" / label
|
| 119 |
+
out_dir.mkdir(parents=True, exist_ok=True)
|
| 120 |
+
writer = FilesystemOutputWriter(out_dir)
|
| 121 |
+
runner = JobRunner(store)
|
| 122 |
+
|
| 123 |
+
started = time.perf_counter()
|
| 124 |
+
await runner.run(
|
| 125 |
+
job_id=job_id,
|
| 126 |
+
document_manifest=manifest,
|
| 127 |
+
provider_name="mistral",
|
| 128 |
+
api_key=api_key,
|
| 129 |
+
model=model,
|
| 130 |
+
output_writer=writer,
|
| 131 |
+
source_files={alto.name: alto},
|
| 132 |
+
provider=(MistralMultimodalProvider() if wants_image else MistralProvider()),
|
| 133 |
+
page_images=images,
|
| 134 |
+
timeout_seconds=0,
|
| 135 |
+
)
|
| 136 |
+
elapsed = time.perf_counter() - started
|
| 137 |
+
|
| 138 |
+
job = store.get_job(job_id)
|
| 139 |
+
report = job.report if job else None
|
| 140 |
+
record: dict = {
|
| 141 |
+
"arm": arm,
|
| 142 |
+
"page": label,
|
| 143 |
+
"model": model,
|
| 144 |
+
"wants_image": wants_image,
|
| 145 |
+
"seconds": round(elapsed, 1),
|
| 146 |
+
"status": job.status.value if job else "unknown",
|
| 147 |
+
"error": job.error if job else None,
|
| 148 |
+
"lines": job.total_lines if job else 0,
|
| 149 |
+
"fallbacks": job.fallbacks if job else 0,
|
| 150 |
+
"retries": job.retries if job else 0,
|
| 151 |
+
}
|
| 152 |
+
if report is not None:
|
| 153 |
+
reasons = Counter(
|
| 154 |
+
line.decision.reason.code
|
| 155 |
+
for line in report.lines
|
| 156 |
+
if line.decision.reason is not None
|
| 157 |
+
)
|
| 158 |
+
usage = report.usage
|
| 159 |
+
record.update(
|
| 160 |
+
input_tokens=getattr(usage, "input_tokens", 0) if usage else 0,
|
| 161 |
+
output_tokens=getattr(usage, "output_tokens", 0) if usage else 0,
|
| 162 |
+
reasons=dict(reasons),
|
| 163 |
+
statuses=dict(Counter(ln.decision.status for ln in report.lines)),
|
| 164 |
+
changed=sum(1 for ln in report.lines if ln.decision.final_text != ln.source_text),
|
| 165 |
+
blocked=sum(
|
| 166 |
+
1
|
| 167 |
+
for ln in report.lines
|
| 168 |
+
if ln.decision.reason is not None
|
| 169 |
+
and ln.proposal is not None
|
| 170 |
+
and ln.proposal.output_text != ln.source_text
|
| 171 |
+
),
|
| 172 |
+
format_losses=dict(report.format_losses or {}),
|
| 173 |
+
samples=[
|
| 174 |
+
{"src": ln.source_text, "out": ln.decision.final_text}
|
| 175 |
+
for ln in report.lines
|
| 176 |
+
if ln.decision.final_text != ln.source_text
|
| 177 |
+
][:5],
|
| 178 |
+
refused=[
|
| 179 |
+
{
|
| 180 |
+
"reason": ln.decision.reason.code,
|
| 181 |
+
"src": ln.source_text,
|
| 182 |
+
"llm": ln.proposal.output_text,
|
| 183 |
+
}
|
| 184 |
+
for ln in report.lines
|
| 185 |
+
if ln.decision.reason is not None
|
| 186 |
+
and ln.proposal is not None
|
| 187 |
+
and ln.proposal.output_text != ln.source_text
|
| 188 |
+
][:5],
|
| 189 |
+
)
|
| 190 |
+
|
| 191 |
+
target.parent.mkdir(parents=True, exist_ok=True)
|
| 192 |
+
target.write_text(json.dumps(record, ensure_ascii=False, indent=1))
|
| 193 |
+
print(
|
| 194 |
+
f"[{arm}] {label:26s} {record['status']:10s} {record.get('lines', 0):5d}L "
|
| 195 |
+
f"{elapsed:6.1f}s {record.get('input_tokens', 0):8d}in/"
|
| 196 |
+
f"{record.get('output_tokens', 0):7d}out",
|
| 197 |
+
flush=True,
|
| 198 |
+
)
|
| 199 |
+
return record
|
| 200 |
+
|
| 201 |
+
|
| 202 |
+
async def main() -> None:
|
| 203 |
+
if len(sys.argv) < 3:
|
| 204 |
+
print(__doc__)
|
| 205 |
+
raise SystemExit(2)
|
| 206 |
+
root = Path(sys.argv[1]).resolve()
|
| 207 |
+
arm = sys.argv[2]
|
| 208 |
+
concurrency = int(sys.argv[3]) if len(sys.argv) > 3 else 3
|
| 209 |
+
if arm not in ARMS:
|
| 210 |
+
raise SystemExit(f"unknown arm {arm!r}; pick from {list(ARMS)}")
|
| 211 |
+
|
| 212 |
+
results = root.parent / "arm-results" / arm
|
| 213 |
+
results.mkdir(parents=True, exist_ok=True)
|
| 214 |
+
pages = find_pages(root)
|
| 215 |
+
print(f"arm {arm} — {len(pages)} pages, concurrency {concurrency}", flush=True)
|
| 216 |
+
|
| 217 |
+
api_key = keychain_key()
|
| 218 |
+
sem = asyncio.Semaphore(concurrency)
|
| 219 |
+
records = await asyncio.gather(
|
| 220 |
+
*(run_one(arm, label, alto, image, api_key, results, sem) for label, alto, image in pages),
|
| 221 |
+
return_exceptions=True,
|
| 222 |
+
)
|
| 223 |
+
|
| 224 |
+
good = [r for r in records if isinstance(r, dict)]
|
| 225 |
+
failed = [r for r in records if isinstance(r, BaseException)]
|
| 226 |
+
totals: Counter[str] = Counter()
|
| 227 |
+
reasons: Counter[str] = Counter()
|
| 228 |
+
for record in good:
|
| 229 |
+
for field in (
|
| 230 |
+
"lines",
|
| 231 |
+
"input_tokens",
|
| 232 |
+
"output_tokens",
|
| 233 |
+
"changed",
|
| 234 |
+
"blocked",
|
| 235 |
+
"fallbacks",
|
| 236 |
+
"retries",
|
| 237 |
+
):
|
| 238 |
+
totals[field] += record.get(field, 0)
|
| 239 |
+
reasons.update(record.get("reasons", {}))
|
| 240 |
+
print(f"\n=== {arm} ===")
|
| 241 |
+
print(f"pages ok {len(good)}, raised {len(failed)}")
|
| 242 |
+
print(f"totals {dict(totals)}")
|
| 243 |
+
print(f"reasons {dict(reasons)}")
|
| 244 |
+
for exc in failed[:3]:
|
| 245 |
+
print(f" raised: {type(exc).__name__}: {exc}")
|
| 246 |
+
(results.parent / f"{arm}-summary.json").write_text(
|
| 247 |
+
json.dumps(
|
| 248 |
+
{
|
| 249 |
+
"arm": arm,
|
| 250 |
+
"pages": len(good),
|
| 251 |
+
"raised": len(failed),
|
| 252 |
+
"totals": dict(totals),
|
| 253 |
+
"reasons": dict(reasons),
|
| 254 |
+
},
|
| 255 |
+
ensure_ascii=False,
|
| 256 |
+
indent=1,
|
| 257 |
+
)
|
| 258 |
+
)
|
| 259 |
+
|
| 260 |
+
|
| 261 |
+
asyncio.run(main())
|