
›Aug 2024 – Nov 2025
AI Engineer
GMO-Z.com RUNSYSTEM
“From 94% to 98% accuracy and ~40% lower latency — every millisecond counts when banks, hospitals, and government agencies depend on your OCR.”
## Metrics
The numbers behind the work — measured in production, not in benchmarks.
94 → 98%
OCR accuracy
multilingual VN/JP, production
−40%
inference latency
via dynamic resize + decoder loop
12
enterprise customers
across VN + JP, banks + gov
## Missions
- 01
Built and deployed a multilingual OCR platform (Vietnamese/Japanese) with an iterative data pipeline (continuous enrichment loop) for dataset expansion, model training, and evaluation; production deployment on cloud and mobile via Triton Inference Server and TensorRT/ONNX, boosting overall accuracy from 94% → 98% and significantly increasing serving throughput.
- 02
Delivered end-to-end computer vision pipelines (YOLO, RT-DETR, Segment Anything) for CAD/technical-drawing understanding, object measurement, license plate recognition, and structured document extraction — at production-grade accuracy across diverse real-world image types.
- 03
Orchestrated a production document-parsing pipeline (FastAPI, PostgreSQL, Docker, S3) producing Markdown/HTML/JSON full-content structures and schema-based extracts using cloud LLM APIs and in-house VLMs, with normalized outputs ready for downstream RAG.
- 04
Implemented an agent-driven Text-to-SQL workflow for stock analysis: semantic table/row search, metadata enrichment, expert few-shot templates, and LLM self-correction/cross-reflection to validate and normalize tabular outputs.
- 05
Optimized models for cloud and edge through distillation, quantization, and efficient pre/postprocessing; introduced dynamic resizing and autoregressive decoder loop management that reduced OCR inference latency by ~40%.
## Trusted by
Production systems delivered to enterprise + government customers across Vietnam and Japan.