Back to Partners
Guide

Governance by design: How review gates, QC scoring, and BYOK turn AI localization into an auditable system

Your legal team does not need a translation vendor. They need to know, six months after a contract addendum ships in Portuguese, who approved it, what changed between the AI draft and the published version, and whether the model that produced it ever had your data. If you can't answer those three questions on...

Governance by design: How review gates, QC scoring, and BYOK turn AI localization into an auditable system

Your legal team does not need a translation vendor. They need to know, six months after a contract addendum ships in Portuguese, who approved it, what changed between the AI draft and the published version, and whether the model that produced it ever had your data. If you can't answer those three questions on demand, "AI-powered localization" is a liability wearing a productivity headline.

That is the problem most localization tooling ignores. Vendors sell throughput, more languages, faster turnaround, lower per-word cost, and treat governance as a dashboard bolted on afterward, a report you can pull if someone asks. That approach fails when legal, brand, or compliance actually asks, because a dashboard summarizing outcomes is different from a system that recorded the decisions.

The argument of this piece is narrower and more structural: governance can't be a reporting layer on top of AI localization. It has to be built into the object model the execution layer runs on, into how orders, review gates, QC evaluations, and model selection are represented as callable, queryable entities. If review status, quality scores, and edit history are not first-class objects with IDs, timestamps, and API endpoints, there is nothing for a compliance function to audit, only a claim to trust.

Why "trust me" doesn't survive procurement

Enterprises publishing under their own brand name in a market they don't have native fluency in are making a bet they can't verify by reading the output. A Head of Content approving German product copy, Japanese contract language, or Portuguese support content is relying on someone else's quality signal. When that signal is "the AI seemed fine" or "our vendor says it's good," the bet is unhedged.

What changes the calculus is a record of quality. Ollang's position is that AI localization becomes defensible to legal, brand, and compliance stakeholders only when every stage of the pipeline, the AI draft, the QC evaluation, the human edit, the final approval, produces a queryable artifact tied to an order ID. That's the difference between a vendor telling you localization is good and a system letting you prove it.

AI QC scoring: structured output, not a spot check

Traditional QA in localization is sampling. Someone with language skills reads a percentage of segments, flags what looks wrong, and the rest ships on faith. It doesn't scale, and it produces an opinion, not a record.

Ollang's execution layer runs QC as a discrete, callable operation against a specific order, returning structured scores rather than a narrative impression. The documentation describes the evaluation as one that triggers an AI-powered assessment of the translation quality, evaluating criteria such as accuracy, fluency, tone, and cultural fit. The output is an array of evaluation scores for each criterion, plus segment-by-segment evaluation results, meaning a reviewer or an integrated system can see exactly which segment triggered a lower tone score, not just that "quality" was rated 82 out of 100 somewhere in the file.

This matters for a Head of Content because it turns "is this good?" into a question with a stored, inspectable answer rather than a subjective judgment that evaporates when the reviewer moves to the next file. The evaluation can be steered: teams can pass a custom prompt to guide the QC evaluation, instructions or focus areas for the AI evaluator such as "Please focus on technical terminology accuracy" or "Pay attention to formal register consistency". A pharmaceutical labeling team and a marketing team asking the same execution layer to check the same four dimensions can weight the evaluation toward what matters for their content, and that instruction is itself part of the record.

Because this runs as an API call rather than a manual review step, it can be triggered automatically, after every order, on a schedule, or as a gate before publication, and the results delivered via callback rather than requiring someone to log in and check. The API reference for running a QC evaluation documents the request and callback payload shape for teams that want to wire this into an existing content pipeline rather than operate it by hand.

QC score progression and human-edit-percentage: the trend line legal actually wants

A single QC score answers "was this file good." It doesn't answer the question a compliance stakeholder actually cares about: is quality in this market improving, holding steady, or degrading as volume scales? That requires a second layer, scores tracked over time, per market, alongside a measure of how much human intervention each output actually needed.

This is where the object model does work a dashboard can't replicate after the fact. Ollang's platform concepts documentation groups AI QC across accuracy, fluency, tone, and cultural fit; human QC annotations; QC score progression and human-edit-percentage analytics as connected primitives, not separate features. QC score progression turns individual evaluations into a series, the same market, the same content type, scored consistently order over order, so a reviewing stakeholder can see the trend line rather than a single snapshot. Human-edit-percentage analytics answers the companion question: when a linguist touched the AI output, how much did they actually change? A market where editors are rewriting 40% of segments is telling you something different than a market where edits are cosmetic and shrinking release over release.

Together, these two metrics are what a risk sign-off process actually runs on. "Trust the AI" is not an auditable statement. "QC scores for this market have held above a defined threshold for the last six release cycles, and human-edit-percentage has dropped by half over the same period" is, because it's a claim built from stored, timestamped data rather than a one-time assurance. The analytics concepts in the API documentation describe this as order analytics, QC progression, payment, credit, and edit-percentage insights, explicitly positioned as data a program can pull, not just a screen someone glances at.

For a Head of Content building a case to expand AI-first localization into a new regulated market, this is the evidence base. It's also the evidence base for the opposite decision, pulling back on a market where the trend line is flat or reversing, before a bad translation reaches a customer or a regulator.

Ready to see Ollang in action?

Talk to our team about your localization goals and see how the Ollang platform fits your workflow.

Book a Demo

The Level 1 review gate: routing control, not a single fixed workflow

The other governance question every stakeholder asks is some version of "who actually looked at this before it went out." A workflow where every order automatically ships AI-only, with human review as an occasional afterthought, does not answer that question well. Neither does a workflow that forces every single order through a human linguist regardless of risk, because that just recreates the throughput bottleneck AI localization was supposed to solve.

Ollang's answer is a review gate that can be applied selectively, per order, rather than fixed at the account level. The documentation is direct about this: enterprises can add a Level 1 review gate to any order to route output to Ollang-managed linguists or your own LSPs and editors. That routing decision is exposed as an API operation, requests or cancels a professional linguist review for an order, meaning it can be triggered programmatically based on content type, market risk, or a QC score that fell below threshold, not just clicked manually in a UI.

The mechanism supports both operating models a Head of Content is likely to be juggling: content that goes to Ollang's own linguist network, and content that has to route to the enterprise's existing LSP relationships or in-house editors for contractual or domain-expertise reasons. Legal content acquired under an existing outside-counsel translation arrangement does not have to leave that relationship to get the benefit of the execution layer's QC scoring and audit trail, the gate routes the work, the platform still records what happened to it.

The routing can also be automated based on quality thresholds rather than requiring a human to decide case by case. Ollang's troubleshooting documentation notes that review gates respect qcThreshold routing rules, orders may be re-routed automatically, and reversing the decision is equally explicit: teams can use Cancel Human Review to revert to the AI-only state and refund the review credits if priority changes. That's a governance control with an undo button, which is what makes it usable in practice rather than a one-way ratchet nobody wants to trigger for fear of getting stuck.

This is also where legal document localization, website and software localization, and video or dubbing workflows converge on the same control surface. A contract clause, a UI string with regulatory disclosure language, and a compliance disclaimer baked into a dubbed training video all carry different risk profiles, but they can all be routed through the same Level 1 gate logic, high-risk content forced through human review, lower-risk high-volume content left to run AI-only with QC scoring as the safety net.

BYOK: keeping model choice inside your governance boundary

The fourth piece is the one legal and security teams usually raise first, because it's the one with the most direct data-handling implication: which model is actually processing our content, and where does our data go when it does.

Ollang's documentation lists BYOK alongside the platform's other core governance concepts, Folders, Projects, Orders, Levels, Workflows, Review Gates, LSPs, QC, BYOK, and more, treated as a structural concept of the platform rather than a checkbox feature. In practice, this connects to something the platform already exposes at the workflow level: the ability to change the AI provider in the Folder workflow and rerun. Model choice is not fixed platform-wide; it is a property of the folder-level workflow, which means an enterprise can decide, at the level of a specific content stream, which underlying provider handles its data.

That distinction matters more than it might first appear. A generic AI localization tool that routes everything through a single provider relationship puts the enterprise's data handling posture entirely inside the vendor's contract with that provider. BYOK moves the decision, and, by extension, the data governance boundary, back inside the enterprise's own control. Security and legal do not have to trust the vendor's provider selection; they can specify it, and that specification is recorded the same way orders, review status, and QC scores are, as part of the workflow configuration, not a side conversation with an account manager.

This is a control an enterprise buyer should verify directly rather than take on faith from marketing language, and it's worth walking through with your own security team against the specifics in Ollang's API documentation before treating it as satisfying a particular compliance requirement.

How this sits in the stack

None of this requires an enterprise to abandon the systems they already run content through. Ollang is the execution layer that content passes through, invoked from wherever the enterprise's own pipeline already lives, a CMS publish hook, a CI/CD step, or an agent workflow, rather than a destination content gets manually shipped to. The platform is API-key authenticated, with a hosted Model Context Protocol server using OAuth 2.0 plus PKCE for agent-native access, and it is built to drop into Claude, Cursor, Claude Code, Devin, Replit, Windsurf, and more. For teams that do not want to run a hosted MCP connection, file-based Agent Skills give natural-language ops with no server to run, and a TypeScript SDK covers asset scanning and CMS capture for engineering-led integrations.

What that means operationally: an order gets created programmatically, QC evaluation and review-gate routing happen as calls against that order rather than manual steps in a separate tool, and the score, edit-percentage, and approval history accumulate against the order ID automatically. A Head of Content does not need to chase down what happened to a piece of content across three different systems, upload, translation memory tool, vendor portal, because the record lives in one place, addressable through the same API surface used to create the order in the first place. The production guidance in the API documentation covers the callback, retry, and error-handling patterns that make this reliable enough to run unattended in a pipeline rather than requiring someone to babysit it.

Ready to see Ollang in action?

Talk to our team about your localization goals and see how the Ollang platform fits your workflow.

Book a Demo

The closing argument

The category shift underlying all of this is that localization stops being a project you commission and becomes infrastructure you call, and infrastructure has to answer to a different standard than a service does. A service earns trust through relationship and track record you take on faith. Infrastructure earns trust through a queryable record: this order, this score, this edit, this approval, this model. Governance built after the fact, as a report generated from logs nobody designed to be audited, will always be reconstructive and partial. Governance built into the object model, where review gates, QC scores, edit-percentage, and model selection are objects with IDs before they're ever summarized in a dashboard, is what lets a Head of Content say, with the same confidence they'd apply to any other regulated business process, exactly what happened to a piece of content before it went out under the company's name.

Published on September 1, 2026