Leading Technical Translation Company in the US
Home/Resources/Translation Memory Explained

Guide

Translation Memory Explained: How TM Saves Money on Technical Documents

A translation memory is a database of segments your documents have already been through, each one translated by a linguist and approved by a reviewer. It is the reason the second edition of a 300-page manual costs a fraction of the first, and the reason two chapters written three years apart still sound like the same document.

Translator reviewing segment matches in a CAT tool editor
Approved segments only

A sentence enters the memory after the second linguist has signed off on it, never before. Drafts and unreviewed output stay out.

Discounts you can audit

The match analysis is a file you can read before anyone starts typing: new words, fuzzy words, repeated words, line by line.

Revisions that match

A warning phrased one way in revision A comes back phrased the same way in revision F, even when a different translator handles the update.

What a translation memory actually is

A translation memory, usually shortened to TM, is a bilingual database of segments. Each record holds one source segment, its approved target, and a set of metadata: the date, the linguist, the project, the file it came from, sometimes the product line. A segment is normally one sentence, though a table cell, a list item, a figure caption or a heading also counts as one.

When a translator opens a file in a CAT tool such as Trados Studio or memoQ, the tool cuts the file into segments and queries the memory for each one. Whatever comes back close enough to the new sentence is offered as a suggestion, with a similarity percentage attached. The translator accepts it, edits it, or ignores it. Nothing is inserted into your deliverable without a human deciding it belongs there.

Two things a TM is not. It is not a dictionary: it stores whole sentences in context, not word pairs. And it is not machine translation. A memory generates nothing. If a sentence has never been translated and stored, the memory returns silence for it, whereas an engine will produce fluent output for any input, including input it has never seen. The two technologies are often used in the same file, which is why they get confused; the difference matters when you are reading a quote. We treat them separately in our overview of machine translation in technical documentation.

Segmentation decides how much you can reuse

Reuse happens at segment level, so the way your source text is cut up governs how much of it comes back. Segmentation rules split on sentence-final punctuation, with exception lists for abbreviations, decimal points and part numbers. A sentence that ends with "approx." will not be split if the rule set knows that abbreviation. If it does not, you get two half-segments that will never match anything cleanly again.

This has a practical consequence for the way technical documents are written. A manual built from short, self-contained instructions reuses beautifully. "Torque the four mounting bolts to 45 ft-lb." is a unit that can reappear in three chapters and two service bulletins. A manual built from long chained sentences, each carrying three conditions and a cross-reference, produces segments that are unique by accident and never come back. Nothing about the meaning changed. The reuse rate did.

One more thing worth knowing: segmentation rules belong to the project, not to the file. Changing the rule set halfway through a program fragments the memory, and matches that used to come back at 100 percent start arriving as fuzzy hits. Vendors who take over a memory without asking which rules produced it often deliver a first project with a suspiciously poor analysis.

Match types, and what each one does to your invoice

The analysis a CAT tool produces before a project sorts every word of the new file into a band. The bands below are the standard ones. Discount grids vary between agencies and language pairs, so read the ranges as typical market practice rather than a price list; your own quote may sit outside them.

BandWhat the tool foundTypical share of the full rate
Context match (also 101% or ICE)Identical source segment, and the segments before and after it are identical too15–25%
100% matchIdentical source segment, different surroundings20–35%
RepetitionSegment appears more than once inside the current file set; the first occurrence is priced normally20–35%
95–99% fuzzyA word or a number differs, or a formatting tag moved40–60%
85–94% fuzzyA clause differs; the sentence still has to be reworked60–80%
75–84% fuzzyEnough overlap to be a starting point, little more80–100%
No match (below 75%)Nothing usable in the memory100%

Two remarks about the bottom of that table. A 75 percent match is frequently slower to fix than a blank segment, because the translator has to read it, diagnose what changed, and resist the pull of wording that no longer applies. Agencies that offer deep discounts down to 70 percent are either not reading those segments or absorbing the cost somewhere else. And repetitions inside a single file set are real savings only when the repeated sentence carries the same meaning everywhere it appears, which is not guaranteed in a spare parts catalog full of one-word cells.

A worked example: the second edition of a manual

Take a 40,000-word operator manual translated into German for the first release of a machine. Eighteen months later the manufacturer issues a second edition: three new chapters on an optional feed module, a rewritten safety section, revised part numbers throughout, and light editing everywhere else. The raw word count is now 42,600.

A typical analysis on that file looks like this: 27,900 words come back as 100 percent or context matches, 7,400 land in the fuzzy bands, 6,100 are genuinely new, and 1,200 are internal repetitions. About 70 percent of the document is recycled. Applying a normal grid, the weighted word count lands somewhere near 17,000, which is roughly 40 percent of the raw count. That is the number your quote is built on.

Two caveats keep that figure honest. Independent review is priced on the full document, because the second linguist reads the manual end to end, including the recycled parts, to confirm the new chapters sit correctly against the old ones. And desktop publishing does not shrink with the match rate: a German second edition still expands, still reflows, still needs its table of contents regenerated. Reuse compresses translation cost, not every cost. Our page on how technical translation rates are built breaks down where the rest goes.

On safety content we do not take the maximum discount. Hazard statements, warning labels and lockout procedures are read in full context on every revision, even when the memory returns a context match. A sentence that was correct against the 2022 revision of a standard can be quietly wrong against the 2026 one, and the tool has no way of knowing that. The saving on those pages is smaller on purpose.

Where a memory quietly costs you

Match percentages describe string similarity. They say nothing about whether the stored translation is still the right one. These are the failure modes worth knowing before you sign a grid.

  • False 100 percent matches. "Check the level." is identical in a hydraulic chapter and an electronics chapter and means two different things. Short segments, table cells and UI strings are where this bites: "Open" as a verb and "Open" as a valve state are one entry in the memory and two different words in French or Korean.
  • Superseded content. A memory built over six years holds phrasing from three revisions of a standard. Without a date filter or a cleanup pass, the translator sees a confident match from 2019 sitting on top of the 2025 wording.
  • Polluted memories. Unreviewed drafts, raw engine output stored as if it were human work, or content merged in from another product line with a different house style. Once it is in, every future project inherits it, and the person who finds the problem is usually the end reader.
  • Tag noise. A bolded word, a cross-reference field or an index marker in a different position drops an otherwise identical sentence to 97 percent. You pay a fuzzy rate for a formatting difference, which is fair enough for the translator and irritating on an invoice.
  • Rewritten sources. When a technical writer restyles a chapter without changing what it says, reuse collapses. The cheapest editorial decision your documentation team can make is to change source sentences only when the content changes.

Who owns the memory

Your memory belongs to you. It is built from your source content and paid for through your projects, and no agency has a defensible claim on it. At Techniwords the memory for each client sits in its own container, is exported as a standard TMX file on request at no charge, and travels with you if you change vendors. TMX imports into every serious CAT tool, so portability is real rather than theoretical.

This is a useful question to put to any supplier before signing. Ask who holds the memory, in what format it can be exported, whether the export costs anything, and how quickly it is delivered. Hesitation on any of those four points tells you what the commercial relationship is going to feel like in year three. The practical side of this, hosting, versioning and handover, is what our translation memory services cover.

Memory, termbase, engine: three tools, three jobs

These three get bundled together in sales conversations and they do genuinely different work.

Translation memoryTermbaseMachine translation
Unit storedWhole segments, source and targetConcepts and their approved designationsNothing; a model produces output on demand
Question it answersHave we translated this sentence before?What do we call this thing, and what are we forbidden to call it?What is a plausible rendering of this sentence?
Who builds itTranslators and reviewers, project after projectTerminologists with your engineers validatingA vendor, on generic and domain corpora
Main failureStale or context-blind matchesNobody maintains it after year oneFluent output that is factually wrong

The memory and the termbase are complementary rather than interchangeable. The memory carries sentences you have already approved; the termbase governs the words inside sentences nobody has written yet. Building and validating that second asset is a discipline of its own, covered in our guide to building a termbase, and delivered as part of our ongoing terminology work.

Keeping a memory in good condition

A memory is an asset that depreciates if nobody looks after it. Six habits keep it worth what you paid for it.

  1. One memory per client and language pair. Split further by product line when the writing styles genuinely diverge, for example a consumer-facing quick start guide against a field service manual.
  2. Update after review, never before. The memory is written back to once the second linguist has signed off. Anything else means you are storing drafts.
  3. Align legacy documents with judgment. Old PDFs and their published translations can be aligned into a starter memory, which is worth doing when the pairs genuinely correspond. When the old translation was poor, alignment imports the problem at scale.
  4. Clean up on a schedule. Once a year, hunt duplicate sources with divergent targets, obsolete part numbers, and phrasing tied to a standard that has since been revised.
  5. Apply an age penalty. Most tools can flag or downgrade segments older than a set date, which puts the translator on alert instead of on autopilot.
  6. Never merge blind. A memory inherited from a previous vendor is reviewed as a batch before it joins yours, or kept as a read-only reference alongside it.

None of this requires anything from you beyond keeping your source files editable and telling your agency when a product line changes name. The rest is housekeeping that belongs to the vendor, and it is fair to ask what their cleanup routine actually is.

What to ask before your next revision

Three questions get you most of the way. Can I see the analysis before the project starts, band by band? What discount applies to each band, and is review priced on the weighted count or the raw count? And can I get the TMX export today, without a reason. If you already have translated material sitting in PDFs with no memory behind it, send a sample with the source files and we will tell you whether alignment is worth the effort, which is part of what goes into an itemized quote.

Translation memory: common questions

Is a translation memory the same thing as machine translation?

No. A translation memory only returns sentences that a human translated and a reviewer approved at some point in the past. If your sentence is new, it returns nothing. Machine translation produces output for any input, including sentences it has never encountered, and that output has not been checked by anyone until a post-editor reads it. The two can be used in the same project, but they are billed differently and carry very different risk.

Do I pay full price for sentences that repeat?

No. Repeated segments inside the same file set, and segments already stored from earlier projects, are billed at a reduced share of the full rate. The first occurrence of a repeated sentence is priced normally because someone has to translate it. On the projects we quote, the analysis showing how many words fall into each band is attached to the quote itself, so you can check the arithmetic before you approve anything.

Can I take my translation memory to another agency?

Yes. The memory is yours. We export it as a TMX file, the standard interchange format, and any professional CAT tool will import it. There is no charge for the export and no waiting period. If you are evaluating a supplier, ask this question early: a vendor who treats your memory as their property is holding your future revision costs hostage.

How much does a memory save on a first project?

On a genuinely first project, usually very little beyond internal repetitions, which on a manual with recurring procedures can still reach five to fifteen percent of the word count. The saving arrives on the second document and grows from there. Clients who translate one document a year see modest returns; clients maintaining a document family across revisions often see half their word count come back as matches by the third year.

What happens if a mistake is stored in the memory?

It gets corrected at the source, in the memory, not only in the delivered file. A fix applied to a single PDF reappears on the next revision. When a client reports an error or a preferred wording changes, we run a search-and-replace pass across the affected memory, log what changed, and note it against the relevant terminology entry so the same decision holds on future projects.

Wondering what your revision would actually cost?

Send the new version and the previous translation. We run the analysis, show you the match bands, and come back with an itemized quote within one business hour.

Get a Free Quote

info@techniwords.us · (346) 296-6516

Request a free technical translation quote from Techniwords