Anonymization services

Identities out.
Meaning intact.

Your documents carry your clients' identities. Ours is the technology that removes them: before AI, before sharing, before risk. Built by Lexmetra for legal content, in production today, and available with or without the QA platform.

What it does

De-identification built for legal text

Legal documents resist naive anonymization: identities hide in abbreviations, case references, account formats and the thousand ways a name mutates across a filing. Lexmetra's pipeline combines machine-learning entity recognition with legal-domain recognizers, including country-specific identifier formats, and propagates variants and acronyms across whole document sets, with layered overrides tuned on real legal text.

The result: bilingual legal corpora with the identities removed and the linguistic value intact, still usable for translation memories, AI workflows and knowledge bases.

What gets removed

  • Names of people and organisations
  • Addresses
  • Financial identifiers
  • Jurisdiction-specific identifier formats
  • Variants and acronyms, across the whole corpus

The standard

Validated like everything we ship

Anonymization is only as good as its verification, so ours is a methodology, not a script: iterative evaluation loops pairing expert human review with LLM-based auditing over stratified samples, engineered to drive residual identification risk toward zero and rigorously validated on real legal corpora.

We publish the method, not marketing percentages. Ask us for the validation story on a corpus like yours.

Two ways to use it

Built in, or on its own

Built into the platform

Every translation-memory asset is de-identified before it enters the AI layer. Your clients' identities never ride along with your firm's linguistic assets: protection is the default, not an option.

As a standalone service

De-identification of legal document sets, translation memories, corpora and datasets, with or without the QA platform: for AI-adoption data preparation, vendor and cross-border sharing, and building knowledge bases on sensitive material, safely.

Designed for a world of GDPR and CCPA: de-identification as an architectural habit, not an afterthought.