Applied AI / Case study
Training AI models to read documents accurately in seconds on local hardware
We trained and optimized AI models for accurate document and product-label recognition in seconds on NVIDIA Jetson AGX Thor. We adapted training and validation methods from a larger mixture-of-experts model to a compact 1B-class model, with runtime tuning for each. The goal was to keep local deployment costs low while retaining control of the adapted models and their serving runtimes.
- Project
- Local AI for document and product-label recognition
- Key outcomes
- Approximately 4.5–4.9 seconds per document with the larger AI model on AGX Thor.
- A compact 1B-class AI model averaged 1–2 seconds per product-label image on the same hardware platform.
- Accelerated and non-accelerated generation achieved the same aggregate exact-identifier match rate on the compact model’s validation workload.
Background
Business documents and product labels contain details that must survive optical character recognition (OCR). One changed character can turn a part number into the wrong identifier. A correctly read price is useful only when later processing can associate it with the correct item.
We also needed control over how the adapted models were served. A cloud service may not support custom low-rank adaptation (LoRA) adapters. Even a compatible provider can become unavailable, and another provider hosting the same base model may not accept those adapters. Falling back to the base model could lose the recognition behavior developed through task-specific training and would require accuracy to be checked again.
A backup provider can also fail. Providers may depend on the same cloud, network, or access services, so one infrastructure outage can affect several alternatives at once. Retaining the adapted model and a validated local runtime gave us a recognition path that did not require an external AI service to be available. This reduced exposure to both provider-specific limitations and shared infrastructure failures.
Deployment cost shaped the hardware choice. We selected NVIDIA Jetson AGX Thor as the local deployment target, then trained and optimized the models for useful recognition accuracy and response times within that hardware budget. The work extended from a larger mixture-of-experts (MoE) model to a compact model with approximately one billion parameters. Accuracy, speed, and the cost of the hardware needed to run the models had to be considered together.
Challenge
The hardest cases combined similar characters with difficult formats. Shapes such as 0/O and 1/l/i could produce plausible but incorrect identifiers. Long strings of zeros, small print, punctuation, and currency marks required exact recognition.
Layout added another constraint. A model could read individual words but omit a narrow column, combine address blocks, or detach a number from its table row. A lower training loss or average error score could hide these failures. The final model also had to retain useful recognition after export and acceleration on the target hardware.
Solution
We defined the information that had to remain exact, then compared source images, training labels, and output. This separated model errors from label errors and harmless formatting differences. Observed failures guided the next training examples and the choice of training checkpoint.
Complete-page inputs preserved small text and its surrounding context. In one diagnostic comparison, an address header was readable in isolation but omitted from the full page. That directed the work toward page coverage as well as character recognition.
We compared adaptation of the visual encoder, the language components, and the connection between them. This helped determine where training could address the observed failures. During adapter training, we used full-precision (FP32) parameters or master weights to protect small learning updates while retaining lower-precision model computation. Exported models were checked again because a change in deployment precision could change a critical character.
We adapted these methods from the larger model to the compact 1B-class model. Each architecture required its own training and runtime work. On AGX Thor, compiled inference and speculative decoding reduced response time. Speculative decoding proposes multiple tokens for the model to verify. We checked speed together with exact-identifier recognition when selecting the execution path.
Result
The larger model reached approximately 4.5–4.9 seconds per document on AGX Thor in the final configuration. The compact model averaged 1–2 seconds per product-label image; the recorded service test averaged about 1.4 seconds.
On the compact model’s validation workload, accelerated and non-accelerated generation achieved the same aggregate exact-identifier match rate. The two model sizes were evaluated on separate document and product-label workloads. The work showed how training and validation methods could be adapted across the two model sizes. Local deployment gave us direct control over the adapted models, their serving runtimes, and the validation needed after a change.