Nutrient · Document Intelligence

Page Stream Segmentation — Unified Leaderboard

One boundary metric, one harness, every contender re-measured: our flagship and open-weight models against cloud VLMs and the field's self-declared numbers. Cells show boundary F1; the small figure is chance-corrected κ.

updated 2026-08-11 · metric: boundary page-F1 (page 0 forced) + Cohen's κ
0.891flagship OpenPSS-long F1 — vs best cloud 0.244
1 modelbeats OpenPSS's two specialists across both slices
~0.0007USD / 1k pages (A40) — cloud VLMs ≈ $0.014 / stream
4.5×lighter open model (doc-split-v1) at near-flagship quality

Boundary F1 · κ — across six evaluation cuts

Model our-200easy, saturated OpenPSS-shortsparse · hardest OpenPSS-longlong streams TABME++test Tobacco800test val-fullour real-doc
Ours
doc-split-v2
flagship · one model for short + long streams · ~1.0B · on-prem
Commercial
0.944.79 0.652.60 0.891.86 0.943.91 0.969.93 0.917.86
doc-split-v1
open-weight · ~4.5× faster
Open-weight
0.936.78 0.585.53 0.859.82 0.704.56 0.820.60 0.918.86
Cloud VLM · image-only · single-prompt
gemini-flash
best cloud on OpenPSS · ~$0.014/stream
Cloud
0.917 0.598.53 0.244.16
gemini-pro
Cloud
0.936 0.530.45 0.196.11
gpt-sol
best cloud on our-200
Cloud
0.942 0.193.17 0.025.02
claude-opus
Cloud
0.318.28 0.047.03
Research — self-declared (their metric / in-domain)
OpenPSS SHORT-specialist
BERT-EffNet ensemble · one of two models
Published
0.76 0.50
OpenPSS LONG-specialist
separate model · cross-slice drops
Published
0.62 0.83
bert-pss (agiagoulas)
only released PSS specialist · text-only
Released
0.915*κ .00 run
≥ 0.85 0.70–0.85 0.50–0.70 < 0.50

One balanced model vs two specialists. No single OpenPSS model wins both slices — their short-specialist craters on long (0.50), their long-specialist on short (0.62). the v4 flagship does short + long with one model, and its long (0.891) tops even their long-specialist (0.83).

Cloud VLMs can't ingest long streams. Per-request image caps (Anthropic ~100 / OpenAI ~200 / Gemini ~500) force predict-none on streams over the cap, which dominates OpenPSS-long → best cloud 0.244 vs flagship 0.891, at ~20× the cost per page.

* bert-pss self-declares ~0.915 accuracy / 0.825 κ on Tobacco800, but collapses to a single class when actually run (κ ≈ 0) — the only released PSS specialist is non-functional off its exact serving harness.

Metric. Ours & cloud: micro boundary-F1 + κ over internal pages, identical harness. OpenPSS rows: their published per-stream page-F1 (ballpark-comparable). Tobacco incumbents report accuracy — compared via κ. AI-Lab-Splitter omitted: data gated, absolute F1 paywalled.