Translation Agency Software: Fixing the Routing, Word Count and QA Leaks That Cost You Margin
If you are running a translation agency doing more than roughly $4M a year, routing 200+ jobs a month across 300+ linguists in 20+ language pairs, the honest answer is that you should build the layer your TMS refuses to own: vendor selection, margin per job, and QA evidence. Expect $60k to $130k for a focused first release shipping in 12 to 16 weeks, and $150k to $400k for a full platform phased across 6 to 12 months. Below that volume, Plunet or XTRF plus memoQ will hurt but they will hold. Above it, the spreadsheet your vendor manager maintains is already the real system, and it does not scale past the person who built it.
Why translation agency software makes or breaks a localization operator
A translation agency is a routing business wearing a language business costume. The value you add is not the words: the linguists produce those. The value is picking the right linguist for the right content at the right price inside the right hour, and proving the output was checked. Everything that goes wrong in a 40 person agency goes wrong in that routing layer.
Here is the scene most operators recognize. It is 4:40pm and a client sends 18 files: an InDesign package, two XLIFF exports from their CMS, a PDF of a regulatory annex, and a zip of subtitle SRTs. Your project manager runs the analysis in memoQ or Trados Studio, gets a weighted word count, then opens Plunet or XTRF to build the quote. Then she opens the spreadsheet. The spreadsheet is where the actual business lives: which of the 340 linguists in your pool handle Japanese medical devices, which one blew a deadline in March, which one raised rates to 0.14 EUR per word in January but is still priced at 0.11 in the TMS, who is on holiday, and who the client's reviewer in Osaka has quietly rejected twice. She sends nine availability emails. Six answer by morning. Two say yes then go quiet. The job goes out 26 hours after it arrived, and the margin on it is a guess until month end.
Across an agency running 200 jobs a month, that is where the money leaks. It is not one catastrophe, it is 200 small ones: a job priced on last year's linguist rate, a fuzzy match discount applied to the client but not recovered from the vendor, a QA step logged as an email saying "looks fine", a rush fee never invoiced because nobody tracked the hour the file arrived. Digital Heroes has rebuilt this layer for agencies where a single vendor manager was carrying 300+ linguist relationships in her head, and the first thing the build surfaces is always the same: nobody knew, per job, whether they made money.
Problem: vendor selection lives in one person's head, not in your system
Your TMS has a vendor database. It has language pair, rate, and maybe a specialization dropdown. What it does not have is the thing your PM actually decides on: this linguist is excellent on EN to DE technical but slow on marketing, has never missed a deadline on jobs under 5,000 words but missed two above that, is 40 percent cheaper than your fallback, and the end client's in-country reviewer likes her. Plunet and XTRF let you filter a list. They do not rank.
The consequence is that when your senior vendor manager takes two weeks off, throughput drops and quality complaints rise, and everyone blames "a bad month". It was not a bad month. It was the routing intelligence walking out of the building.
A custom build turns that into a scored assignment engine. Every completed job writes back structured outcomes: on-time or late and by how many hours, QA error count by category from your LQA form, client reviewer changes accepted versus rejected, actual cost versus quoted cost, and rework hours. The engine scores each candidate linguist against the incoming job on content type, subject domain, volume band, deadline pressure, and historical performance on that specific client. Your PM opens a job and sees a ranked shortlist with the reasoning visible: first choice because 14 prior jobs for this client, zero late, LQA average 0.4 errors per thousand words; second choice cheaper by 0.03 per word but two late deliveries above 8,000 words. This is the part of the job a model is actually good at: not translating, but reading the free-text reviewer comments and QA notes that nobody ever coded into categories, and turning five years of them into structured performance signals. Then the offer goes out automatically to the top three in sequence with an acceptance window, and if nobody takes it in 45 minutes it escalates to a human. The 26 hour turnaround becomes 90 minutes, at 5am, with no PM awake.
Problem: word counts do not equal what you get paid or what you pay out
Your CAT tool produces a weighted analysis: 101 percent, repetitions, 100 percent matches, 95 to 99, 85 to 94, and no match. Your client contract applies one grid. Your linguist agreement applies a different one. Your PM reconciles them by hand, or does not, and takes the CAT tool's number as gospel for both sides.
This is the single most expensive silent leak in the category. When a client negotiates a better fuzzy grid, the agency almost never renegotiates the vendor grid to match, so the discount comes straight out of margin, invisibly, on every job for two years. Plunet and XTRF can hold rate cards, but they cannot reconcile them per job against actual CAT analysis output, per client, per vendor, per file, and tell you the delta.
What we build: an analysis ingestion service that pulls the raw analysis directly from memoQ, Trados, or Phrase via their APIs on job creation, then applies the client grid and the vendor grid to the same numbers independently. The job screen shows quoted revenue, projected vendor cost, and gross margin percent before anyone accepts the job. A configurable floor blocks or flags assignment below your margin threshold. A dashboard shows margin by client, by language pair, by content type, and by linguist, so you can see that your third biggest account is running at 11 percent while you think it is at 34. Agencies routinely find one or two accounts they have been subsidizing for years. That single view has paid for the build inside a quarter more than once.
Problem: QA is an opinion until you can prove it
The reviewer sends back "quality was poor on this batch." You ask for specifics. You get a Word doc with tracked changes and three comments. Now you are in a dispute you cannot win because you have no counter-evidence, and you either credit the invoice or lose the account.
Your CAT tools have QA checks: terminology, numbers, tags, consistency. Those catch mechanics. They do not produce an auditable quality record tied to a linguist, a job, and a client standard. Off-the-shelf LQA modules exist but they are generic MQM or DQF forms bolted on, and nobody fills them in because they add 20 minutes per job with no payoff to the person doing the work.
The build makes QA a byproduct instead of a chore. The reviewer works in the same interface, and every change they make is captured at segment level and auto-classified against your error typology: accuracy, terminology, style, locale convention, or preferential. The classification is the model's job and it is a good one, because comparing source, translated segment, revised segment, and your termbase reliably separates a real terminology error from a reviewer's stylistic preference. That distinction is worth real money: preferential changes are not linguist errors and should never count against a vendor's score or justify a client credit. You end up with an error rate per thousand words per linguist per client, a defensible record when a client escalates, and the input the routing engine needs. Attach the client's own reviewer to the same interface and their complaints arrive as structured data instead of a Word doc.
Problem: intake is manual and clients want a portal that is actually yours
A client's marketing manager emails files at 6pm Friday. Nothing happens until Monday at 9am. Meanwhile your competitor with a portal has already quoted.
Plunet and XTRF ship customer portals. They work, and they look like 2011, and they cannot express your specific intake logic: this client's legal content always goes to the sworn translator pool, files over 15,000 words always trigger a human quote review, anything tagged "product launch" gets the rush grid.
A custom intake does file drop, automatic format detection and file prep across InDesign, XLIFF, SRT, PDF, and DOCX, CAT analysis fired automatically, and a quote returned in under 10 minutes at 6pm Friday, with a checkout that takes a PO or a card. Document extraction is where a model does load-bearing work: a scanned regulatory PDF that used to take a DTP specialist 90 minutes to prepare becomes a clean, segmented, count-accurate source file in minutes. For jobs above your threshold, the client gets an instant estimate and a human confirmation by 9am, which is still a competitive answer and does not commit you to a number you regret. Push status back into their world too: a Slack notification, a webhook into their CMS, an API endpoint their developers call from their own release pipeline. Localization buyers who can trigger jobs from their own build process do not leave.
Problem: nobody can forecast capacity, so you say yes and then panic
Sales closes a 400,000 word EN to 12 languages contract starting in six weeks. Do you have the linguists? Nobody knows. You find out in week two.
Your TMS knows about jobs that exist. It knows nothing about jobs that are 70 percent likely to land, and it has no model of your pool's realistic throughput: this linguist does 2,500 words a day sustainably, not the 3,000 in their profile, and already has 40 percent of next month booked from your other client.
The build joins your pipeline to your pool. Weighted forecast volume by language pair against modeled available capacity, per week, showing you in week minus six that EN to Finnish medical is the constraint and you need two more qualified linguists. This is the fourth job for a model here, and the least glamorous: your own history predicts client volume per month better than sales does, because clients have patterns and sales has hope. The output is a recruiting trigger, not a report.
What this actually costs and how long it takes
Framed only as Digital Heroes delivery experience across 2,000+ projects: a focused first release typically runs $60k to $130k and ships in 12 to 16 weeks. For a translation agency that usually means the assignment engine plus margin-per-job plus linguist portal, running alongside your existing TMS rather than replacing it. Full platforms, meaning intake portal, CAT integrations, LQA, vendor payments, and finance sync, land at $150k to $400k phased over 6 to 12 months.
What drives price up specifically in this category: the number of CAT tool integrations, because memoQ, Trados, Phrase, and XTM each have a different API personality and Trados in particular varies by deployment. File format handling, because clean InDesign and scanned PDF prep is real engineering, not a library call. Multi-currency vendor payments across 40 countries with tax handling. Data residency, if you serve EU clients under GDPR or handle patient or legal content, since your linguists are contractors in a dozen jurisdictions and content routed to them is a processing question your enterprise clients will audit. ISO 17100 process evidence, if you are certified, because the workflow must produce the audit trail as it runs. And migration: pulling 8 years of job history out of Plunet cleanly is where projects quietly go over, and it is also the data your scoring engine depends on, so it cannot be skipped.
Build versus buy: take the honest position
Buy if you are under roughly $3M in revenue, or under 100 jobs a month, or working in fewer than 8 language pairs. Plunet and XTRF are mature, they cost a fraction of a build, and at that scale the routing intelligence genuinely does fit in one competent vendor manager's head. Building at that stage is ego, not economics. Same answer if your work is single-vertical and highly repetitive: the spreadsheet is fine because there are only 30 linguists that matter.
Build when these signals appear, and they usually appear together. Your vendor manager is the single point of failure and everyone knows it. You cannot state gross margin per job without a month-end exercise. You have lost or credited an account over a quality dispute you could not evidence. Your biggest client is asking for API access into their release pipeline and your TMS cannot give it. And the tell that settles it: you are paying for Plunet or XTRF and your team still runs the business out of a spreadsheet next to it. That is not a training problem. That is the tool refusing to model your business, and every month you wait, the gap compounds into the pricing you cannot defend and the linguists you cannot replace.
The pragmatic path is rarely rip and replace. Keep the TMS for what it does adequately, invoicing, basic job records, client contacts, and build the intelligence layer on top with a two-way sync. That is why the 12 to 16 week first release is realistic: you are not rebuilding accounting, you are building the part nobody sells you.
How to choose a developer for translation agency software
Make them explain your data model back to you. Ask how they would model a job that splits into 12 target languages, where 3 have a separate review step, 1 goes to a sworn translator, and the client changes the source file after 4 are delivered. If they reach for a flat jobs table, they have never built this. The right answer involves job, language task, and step as separate entities with independent state, and versioned source assets. Getting this wrong is not a refactor, it is a rewrite.
Ask what they have integrated, by name and by version. "We can integrate anything" means they have integrated nothing. You want to hear specifics about memoQ Server's API versus its web service, about how Trados behaves on-prem versus Language Cloud, about parsing analysis logs when the API will not give you what you need. Ask what broke.
Push on compliance before the proposal, not after. Where does content sit, who processes it, how do you handle a linguist in a country your client's DPA does not cover, and how does the system produce ISO 17100 evidence without a human assembling it. If they treat GDPR as a checkbox, your first enterprise security review will be a disaster.
Confirm you own the code and the data outright. Full repository ownership from day one, your cloud account, exportable schema, and no per-seat licence on your own build. If a developer wants to host your linguist pool and job history on their infrastructure under their terms, you have just swapped one lock-in for a smaller, less accountable one.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- In the Flexera 2025 State of ITAM report, respondents reported roughly 33% of SaaS spend is wasted, underscoring how paying for off-the-shelf seats and tiers that go unused erodes the supposed cost advantage of generic SaaS. Source: Flexera (2025) →
- PMI's Pulse of the Profession research found organizations waste an average of roughly 9.9% of every dollar invested in projects due to poor performance - equivalent to about $1 million wasted every 20 seconds collectively worldwide. Source: Project Management Institute (PMI) (2018) →
- One in four US employees report lacking career advancement opportunities; 48% of employees who participated in mentorship programs report high job satisfaction versus 29% of non-participants, and access to advancement opportunities ranges from 33% at organizations under 10 employees to 74% at those with 1,000+. Source: Gallup (2025) →
- 48% of private companies cite integration with legacy systems or technical debt as a top obstacle to realizing the full value of their digital and AI investments (behind data quality/availability at 72% and gaps in AI fluency or technology talent/leadership at 53%). Source: Deloitte (2026) →
Rohan advises mid-market and enterprise teams on ERP, CRM and custom software, and has led delivery on dozens of business-software builds.
Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.