#ocr#ml#production

Why we re-trained the OCR backbone instead of fine-tuning

OCR output degrading on a difficult bank document — the moment a fine-tune isn't enough.

There is a particular kind of fatigue that comes from looking at a model report and feeling it lie to you. The dashboard said the OCR system was improving. Average accuracy looked healthy. Internal demos were clean enough to make everyone briefly optimistic. Then a real customer would upload a low-contrast scan from a branch office, or a phone photo taken under fluorescent light, and the illusion would fall apart in under a second.

At GMO-Z.com RUNSYSTEM, we spent weeks trying to be disciplined about the obvious thing first. Fine-tune the model we already had. Add more examples. Adjust the augmentation. Clean the labels. Tighten the decoder. Do the respectable engineering before asking for the more expensive answer. And still, the same errors kept surfacing. Vietnamese diacritics collapsing into each other. Thin printed Japanese characters disappearing when the document was compressed one too many times. Table headers that looked stable in curated datasets turning brittle in the wild. The kind of failures that were not dramatic enough to make the system crash, but sharp enough to break trust.

What the long tail actually looks like

When we finally stopped staring at the headline accuracy and sliced the failures by what kind of page produced them, the picture became uncomfortable in a useful way. Roughly a third of all customer-visible errors were Vietnamese diacritics collapsing — `ấ` reading as `â`, `ặ` decaying into `a`. Another quarter were thin-stroke Japanese glyphs eaten by JPEG compression. Table headers, low-contrast scans, and a long bucket of "other glyph drift" filled in the rest. Nothing in that distribution looked like a problem you could fix by adding two thousand more clean training pages.

Bar chart of OCR error types in production traffic
chart ·Where the model lost trust — error categories from a week of production logs (illustrative, not benchmark numbers).

This is the part of the long tail that benchmarks do not flatter. The Donut and Nougat papers (Kim et al., 2022; Blecher et al., 2023) both quietly admit something similar — once a document deviates from the printed-page norm, OCR-free or OCR-heavy systems tend to fail in ways that look small per page but compound per workflow. A 2% field-level error rate on identity documents is not a 2% problem; it is the customer who has to retype their address three times before giving up.

Why fine-tuning hits a ceiling

That was when the conversation changed. We stopped asking how to squeeze more out of the current backbone and started asking whether the backbone itself had learned the wrong instincts. It is an uncomfortable moment, because retraining sounds like admitting the previous months were not enough. It stretches timelines. It makes people nervous. It forces product conversations that are much easier to delay than to have.

But once you have seen enough production documents, you realise some problems are not parameter problems. They are representation problems. The model had become too fluent in the world we prepared for it and not fluent enough in the world customers actually inhabited. Bank forms folded at the edges. Photocopied IDs with bruised shadows. Receipts that looked like they had survived rain. Medical paperwork with tiny text and strange spacing. Real documents are rude. They refuse the dignity of a benchmark.

Fine-tuning, when you do it honestly, is a parameter solution. You are nudging weights in a basin that was carved by the original pre-training distribution. If your customers live outside that basin, no amount of nudging gets you out of it — you just smooth the floor while the cliff stays where it always was. The TrOCR paper (Li et al., 2021) is candid about this: a Transformer encoder pre-trained on relatively clean printed-text corpora behaves differently under domain shift than a CNN+CTC stack, and the failure shape is not always linear in the amount of fine-tune data you throw at it. We saw the same shape. After the first few thousand additional Vietnamese examples, the curve began to flatten. After ten thousand, it had a ceiling.

Line chart comparing fine-tune ceiling vs retrain trajectory
chart ·Fine-tuning curves flatten; retraining keeps slope on the hardest quartile of documents.

The retrained backbone, trained from a checkpoint with our domain mixed into pre-training rather than bolted on, did something different. On the easy quartile of documents — the curated scans that had always looked good — it gave up a fraction of accuracy. On the hardest quartile — the folded, faxed, faded edge of the long tail — it kept its slope. We were trading a luxury we did not need for a guarantee customers actually felt.

How we restructured eval as a living loop

So we rebuilt the pipeline with that rudeness in mind. We widened the data sources. We were more aggressive about hard-example mining, in the spirit of OHEM (Shrivastava, Gupta, Girshick, 2016) — letting the model's own mistakes pick the next batch instead of pretending uniform sampling was fair. We treated evaluation as a living argument instead of a ceremony at the end. We made the enrichment loop part of the product rhythm, not a rescue mission after a failure.

Diagram of the new training pipeline
fig ·The pipeline as it actually lived — eval is a loop, not a ceremony.

The shape that emerged looks obvious once you draw it, and that is the trap: it is obvious only after you have lived through the version where it was not. Production failures re-enter mining. Mined examples flow into augmentation. Augmented batches retrain the backbone. The backbone is scored not by a single number but by sliced metrics across script, layout, and lighting — closer in spirit to a model card (Mitchell et al., 2019) than to a leaderboard row. Every release ships with a diff: where it improved, where it regressed, and which slices we explicitly accepted as cost.

The failure mode this guards against is exactly what Sculley et al. (2015) called CACE — Changing Anything Changes Everything. In an OCR system tied to downstream KYC, accounting, and routing, a 1% improvement in the average can quietly hide a 6% regression in the slice that matters most to your biggest customer. We learned to publish both numbers, in that order, every time.

What I would tell my younger self about backbones

The retraining itself was only half the work. The harder half was teaching the team — and myself — to look at error patterns without ego. To stop reading a regression chart as an indictment and start reading it as a map. The LayoutLMv3 work (Huang et al., 2022) makes the point structurally: when text and image are pre-trained together with unified masking, the representation behaves differently from a stack where the layout encoder was retrofitted later. "Bolt-on" representations carry a kind of debt you only see when the documents start arguing back.

What surprised me most was not the lift in accuracy, though that mattered. It was the change in tone when customer feedback came in. Fewer apologetic workarounds. Fewer quiet guesses about why the system had misread a field. More confidence that when it failed, it failed in ways we understood. That distinction matters more than people admit. A system you understand can be improved. A system that merely performs well in meetings becomes superstition very quickly.

I used to think model work was mostly about cleverness. Better losses. Better architectures. Better tricks hidden in notebooks after midnight. There is some truth there. But production AI has made me less romantic about cleverness and more respectful of honesty. The hardest thing is not building a model that shines when the document behaves. It is building one that keeps its footing when the document arrives tired, crooked, and inconveniently human.

Sometimes the most practical decision is also the most humbling one. Some systems do not need more polish. They need a new spine.

References

  1. [1]Li et al. (2021). TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models · arXiv:2109.10282Why a Transformer backbone behaves differently than CNN+CTC stacks under domain shift, and why fine-tune curves on it tend to flatten earlier than people expect.
  2. [2]Kim et al. (2022). Donut: OCR-free Document Understanding Transformer · ECCV 2022Useful counterpoint: even an OCR-free encoder-decoder still inherits long-tail failure shapes from its pre-training corpus.
  3. [3]Blecher et al. (2023). Nougat: Neural Optical Understanding for Academic Documents · Meta AIA modern reminder that document parsing degrades non-uniformly when the input distribution drifts, especially on math, tables, and dense layout.
  4. [4]Huang et al. (2022). LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking · ACM MM 2022The structural argument for representation work over bolt-on adapters: pre-training text and image together changes what the model can be honest about.
  5. [5]Sculley et al. (2015). Hidden Technical Debt in Machine Learning Systems · NeurIPS 2015The CACE principle (Changing Anything Changes Everything) is the cleanest one-line explanation of why a slice-aware eval loop is not optional in production OCR.
  6. [6]Shrivastava, Gupta, Girshick (2016). Training Region-based Object Detectors with Online Hard Example Mining · CVPR 2016OHEM is the canonical reference for letting the model's own mistakes pick the next batch — the same instinct we ported into the OCR mining stage.
  7. [7]Mitchell et al. (2019). Model Cards for Model Reporting · FAT* 2019Model Cards reframe a release note from a number into a slice-aware contract — closer to how production OCR should be reported.
Drag to move · tap to chat · double-click for terminal