#career#vietnam#ai-engineering

From SmartPay to GMO: 18 months of AI engineering in Vietnam

Saigon at dusk, Tokyo at night — the path of an AI engineering career between the two.

There is a very specific smell to a late office in Saigon after the air conditioning has been running too long. Cold coffee. Warm laptops. Someone's instant noodles from an hour ago. The strange calm that settles in after a team has stopped talking in complete sentences. A lot of my early AI engineering life in Vietnam lives in that smell. At SmartPay, everything felt immediate. Fraud detection does that to you. The feedback loop is sharp, because the system is making judgments in a world that keeps moving whether you are ready or not. Features drift. Behavior changes. The business wants confidence, but the data keeps reminding you that confidence is rented, not owned. I learned a lot there about operational humility. A model can look excellent in a notebook and still become awkward the moment it meets a real workflow. I learned to care about retraining pipelines, feature freshness, and the quiet business logic sitting around the model like scaffolding nobody applauds. Then came GMO-Z.com RUNSYSTEM, and the texture of the work changed. Suddenly the problems were wider. OCR that had to survive bad scans, low-end devices, multilingual documents, and customer expectations shaped by enterprise contracts rather than startup speed. Document AI pipelines that touched cloud APIs, in-house models, storage layers, review flows, downstream systems. More surface area. More responsibility. More places where a small oversight could become a long week. I used to think career growth would feel like becoming more certain. It has felt more like becoming harder to impress. Not cynical. Just less easily seduced by nice demos and fast applause. Vietnam's AI scene has its own energy. It is ambitious, scrappy, and often asked to do serious work under practical constraints. Budgets are real. Deadlines are real. Customers are rarely paying for novelty alone. They want throughput, stability, measurable improvement, and someone who will still answer the phone when the edge cases arrive three days before Tết. That environment has been good for me. It has taught me that engineering maturity is not about sounding architectural in meetings. It is about making trade-offs you can still defend when the system is tired and the customer is impatient. It is about knowing which shortcuts are harmless and which ones quietly create debt that another version of you will resent. Some of the most meaningful moments from these eighteen months were not launches. They were clarifications. Realising that latency is product work. Realising that dataset quality is culture work. Realising that the difference between a prototype and a system is often just whether somebody is willing to live with it on a difficult Tuesday. I have also become more protective of curiosity. There were stretches when the work moved so quickly that curiosity started to feel inefficient, almost indulgent. Just ship it. Just patch it. Just make the metric recover. But the engineers I admire most are still paying attention. They still ask why the error clusters that way. They still care about the shape of the failure, not just the count. That kind of attention feels like a moral choice as much as a technical one. Eighteen months is not a grand timeline. I know that. But it has been long enough to show me what kind of builder I want to become. Not the fastest in every room. Not the loudest either. Just someone whose systems feel more honest each year.

Personal skill arc across 18 months
chart ·What I thought I was learning vs what I was actually learning. The honest curve is the lumpy one.

What "production maturity" actually meant in month 9

For the first half of those eighteen months, I had a story I told myself about progress. The story said: each month, I understand a little more, I am a little less surprised, I push back a little harder on bad ideas in design reviews. A nice line going up and to the right. The kind of arc you would put on a self-evaluation form.

The real arc was lumpier. It dipped the week a fraud rule I had quietly defended turned out to be triggering on the wrong slice of merchants. It dipped again when I argued — confidently, in a meeting, in front of people I respected — that a simple class-weight tweak would close a precision gap, and it did not. It dipped a third time around month nine, the week we stopped pretending the OCR backbone was "just under-tuned" and admitted it had learned the wrong instincts on a dataset that no longer matched what customers were sending us.

What I now think of as production maturity is not the absence of those dips. It is the speed at which I let myself fall into them without flinching. Sculley and colleagues have a line in *Hidden Technical Debt in Machine Learning Systems* (2015) that I came back to often: the ML code is the small black box at the center of a vast surrounding system. Most of the cost lives outside that box. Every dip in my own arc was the moment I was forced to look at that surrounding system honestly — the joins, the windowing logic, the silent retries, the human review queue, the report no one had opened in weeks.

The dips taught more than the slopes between them. The slopes were just me getting fluent in a vocabulary the dips had already paid for.

The MLOps boxes nobody applauds

Diagram of the MLOps stack at SmartPay then GMO
fig ·The stack as experienced from inside — most of the work lives in the boxes nobody applauds.

If you draw the SmartPay-then-GMO stack as I actually lived it, the diagram does not look exotic. Data ingestion. Feature freshness. A model registry. A retraining loop. An eval harness. Drift monitoring. A human review flow. An on-call rotation that mattered most around national holidays.

What the diagram hides is which boxes get applause. The notebook gets the screenshot. The training job gets the slack message. The fine-tuned model is the one that ends up in the deck. The boxes with the most leverage — feature freshness, dataset hygiene, the retraining cadence, the eval slices — get exactly zero parties. Sambasivan and colleagues, in *Everyone wants to do the model work, not the data work* (CHI 2021), describe this almost too cleanly. Their fieldwork on "data cascades" matches what I lived: small upstream compromises in labeling, sampling, or schema discipline that quietly compound until a downstream metric drops and nobody can quite say why.

I was not above the cascade. I added to it. There is a particular kind of guilt to opening a notebook six months later and realising that the augmentation policy you wrote in a hurry has been silently shaping every retrain since. The Kreuzberger, Kühl, and Hirschl MLOps survey (IEEE Access, 2023) gave me language for what we had been improvising — terms like *workflow orchestration*, *model registry*, *feature store* — and a way to look at our own diagram and see the gaps without shame. Breck et al.'s *ML Test Score* rubric (2017) was even more useful as a mirror. We were not failing on cleverness. We were failing on tests we had simply never written: schema tests, training-serving skew tests, slice tests on the documents that customers actually sent.

A quiet rule formed for me around month twelve: if a part of the system would embarrass us in front of a customer at 11pm on a Friday, it deserved tests, dashboards, and an owner — even if it never deserved a demo.

Underspecification, and the limits of "it worked on our split"

The other paper that kept ambushing me was D'Amour et al.'s *Underspecification Presents Challenges for Credibility in Modern Machine Learning* (2020). Their argument is simple and unkind: many ML pipelines admit a wide set of solutions that perform identically on the held-out set but behave very differently in the real world. Two checkpoints with the same validation accuracy can disagree wildly on a Vietnamese ID card scanned through a scratched phone screen.

I felt this every week at GMO. Two backbones, indistinguishable on our internal eval, would split apart the moment they met faded ink, low-light photos, or a form layout we had only seen twice. Underspecification is not a model bug. It is a reminder that an evaluation set is a hypothesis about the world, and the world keeps voting against the hypothesis. The practical response was boring and good: more slice metrics, more deliberate hard-example mining, more evaluation as a *continuing argument* rather than an end-of-quarter ritual.

What Vietnam's AI scene rewards

Radar chart of Vietnam AI engineering context
chart ·Vietnam AI engineering, by feel: high deadline pressure, real customer constraints, scrappy infra.

Vietnam's AI engineering context is not unique, but it is distinctive. Budgets are tight enough that every GPU hour gets argued. Deadlines are real enough that pilots ship under quarter-end pressure with the customer's logo already on the deck. Customer expectations are shaped by enterprise contracts — banks, BPOs, regulated workflows — that do not award points for novelty. The talent pool is deepening fast but still uneven; there are excellent juniors everywhere and a thinner layer of senior production hands. Infra access is scrappy: cloud credits stitched to on-prem boxes, regional zones chosen for latency to a specific bank, distillation and quantization treated as default tools rather than exotic optimizations.

Industry reporting from MIC and adjacent sources sketches the macro picture, but the micro picture is what shaped me: the constraints made trade-offs *legible*. You could not hide behind cleverness because there was no budget for cleverness that did not pay rent. That environment rewards engineers who can defend a decision in three sentences to a sceptical product owner who has been on a phone call since seven in the morning. It rewards systems that degrade gracefully, not systems that perform brilliantly only in good weather.

It also rewards a particular kind of agent design — the kind Anthropic wrote about in *Building Effective Agents* (2024). The post argues, gently, that most production value comes not from the most autonomous agent but from well-composed workflows with clear escape hatches. That matched my experience exactly. Every time we built something autonomous and brittle, we paid for it. Every time we built something modular, observable, and easy to interrupt, the customer trusted it more — and so did we.

A reading list I kept revisiting

References

  1. [1]Sculley, Holt, Golovin, Davydov, Phillips, Ebner, Chaudhary, Young, Crespo, Dennison (2015). Hidden Technical Debt in Machine Learning Systems · NeurIPS 2015The paper that named the scaffolding nobody applauds — most ML cost lives outside the model code itself.
  2. [2]Breck, Cai, Nielsen, Salib, Sculley (2017). What's Your ML Test Score? A Rubric for ML Production Systems · NeurIPS ML Systems WorkshopA practical rubric — schema tests, training-serving skew tests, slice tests — that doubles as a mirror for any production team.
  3. [3]D'Amour et al. (2020). Underspecification Presents Challenges for Credibility in Modern Machine Learning · JMLR 2022Two checkpoints with the same validation accuracy can disagree wildly in the wild. The held-out set is a hypothesis, not a verdict.
  4. [4]Anthropic (2024). Building Effective Agents · Anthropic Engineering BlogProduction value tends to come from well-composed workflows with clear escape hatches, not from maximally autonomous agents.
  5. [5]Kreuzberger, Kühl, Hirschl (2023). Machine Learning Operations (MLOps): Overview, Definition, and Architecture · IEEE AccessA vocabulary for the architecture most teams keep reinventing — workflow orchestration, model registry, feature store, monitoring.
  6. [6]Sambasivan, Kapania, Highfill, Akrong, Paritosh, Aroyo (2021). 'Everyone wants to do the model work, not the data work': Data Cascades in High-Stakes AI · CHI 2021Fieldwork on the small upstream compromises that quietly compound into large downstream failures.
  7. [7]MIC / industry reports (2023). Vietnam's AI Strategy and Industry Landscape (overview) · Public industry reportingMacro context for the budget, deadline, and customer-expectation pressures described in the post.

I was not reading these papers to look smart. I was reading them because they put names on things I was already paying for in tired evenings and Saturday hotfixes. Sculley named the scaffolding. Sambasivan named the cascades. D'Amour named the gap between a clean validation curve and a customer's actual document. Breck and colleagues named the tests we had not written. Kreuzberger and colleagues named the architecture we kept reinventing. Anthropic named the design discipline that made our agent products less embarrassing.

Named things are easier to fix. That is most of what these eighteen months have given me — not certainty, but better names for what is happening when a system feels honest, and better names for what is happening when it does not. I am still not the fastest engineer in any room. I am, slowly, the one most willing to live with the system on a tired Tuesday. That is the version of myself I am trying to build next.

Drag to move · tap to chat · double-click for terminal