One boundary metric, one harness, every contender re-measured: our flagship and open-weight models against cloud VLMs and the field's self-declared numbers. Cells show boundary F1; the small figure is chance-corrected κ.
| Model | our-200easy, saturated | OpenPSS-shortsparse · hardest | OpenPSS-longlong streams | TABME++test | Tobacco800test | val-fullour real-doc |
|---|---|---|---|---|---|---|
| Ours | ||||||
doc-split-v2 flagship · one model for short + long streams · ~1.0B · on-prem Commercial |
0.944.79 | 0.652.60 | 0.891.86 | 0.943.91 | 0.969.93 | 0.917.86 |
doc-split-v1 open-weight · ~4.5× faster Open-weight |
0.936.78 | 0.585.53 | 0.859.82 | 0.704.56 | 0.820.60 | 0.918.86 |
| Cloud VLM · image-only · single-prompt | ||||||
gemini-flash best cloud on OpenPSS · ~$0.014/stream Cloud |
0.917 | 0.598.53 | 0.244.16 | — | — | — |
gemini-pro Cloud |
0.936 | 0.530.45 | 0.196.11 | — | — | — |
gpt-sol best cloud on our-200 Cloud |
0.942 | 0.193.17 | 0.025.02 | — | — | — |
claude-opus Cloud |
— | 0.318.28 | 0.047.03 | — | — | — |
| Research — self-declared (their metric / in-domain) | ||||||
OpenPSS SHORT-specialist BERT-EffNet ensemble · one of two models Published |
— | 0.76 | 0.50 | — | — | — |
OpenPSS LONG-specialist separate model · cross-slice drops Published |
— | 0.62 | 0.83 | — | — | — |
bert-pss (agiagoulas) only released PSS specialist · text-only Released |
— | — | — | — | 0.915*κ .00 run | — |
One balanced model vs two specialists. No single OpenPSS model wins both slices — their short-specialist craters on long (0.50), their long-specialist on short (0.62). the v4 flagship does short + long with one model, and its long (0.891) tops even their long-specialist (0.83).
Cloud VLMs can't ingest long streams. Per-request image caps (Anthropic ~100 / OpenAI ~200 / Gemini ~500) force predict-none on streams over the cap, which dominates OpenPSS-long → best cloud 0.244 vs flagship 0.891, at ~20× the cost per page.
* bert-pss self-declares ~0.915 accuracy / 0.825 κ on Tobacco800, but collapses to a single class when actually run (κ ≈ 0) — the only released PSS specialist is non-functional off its exact serving harness.
Metric. Ours & cloud: micro boundary-F1 + κ over internal pages, identical harness. OpenPSS rows: their published per-stream page-F1 (ballpark-comparable). Tobacco incumbents report accuracy — compared via κ. AI-Lab-Splitter omitted: data gated, absolute F1 paywalled.