Stilisierte Browseroberfläche mit Profilkarten, umgeben von Begrüßungen in vielen Sprachen.

Question in English, technical expertise in German

Why traditional chatbots fail on multilingual TYPO3 websites – and what the solution is

An international user asks the chatbot on your TYPO3 platform for regulatory details in English – and receives the reply: “I cannot find any relevant information.” Yet the answer has long been available: it’s contained in a 30-page PDF in the TYPO3 FAL, albeit exclusively in German. This is precisely where standard RAG systems reach their limits.

The pitfall: asymmetric multilingualism


In established multi-site landscapes, there is rarely a uniform picture:

  • Subportal A offers German and English.
  • Subportal B exists only in German.
  • Sub-portal C covers German, French and English.
  • Over 90 per cent of in-depth specialist PDFs, statutes and data sheets are available exclusively in German.

If a conventional language filter is set at database level (sys_language_uid = query_language), search queries in other languages will inevitably return no results. The knowledge is there, but remains trapped in monolingual silos.

The solution: three-stage decoupling


Instead of expensively translating millions of documents in advance, a robust architecture decouples the language layers from one another:

  1. Front-end level: Detects the user’s language and controls the interface and dialogue flow.
  2. Cross-lingual retrieval: Multilingual embeddings map semantic meaning independently of language within the same vector space. An English query thus matches directly with German text and PDF passages. Bilingual metadata and synthetic queries support the retrieval of purely German documents.
  3. LLM synthesis: The language model summarises the content found strictly in the user’s language and transparently cites the source (“According to the official German document…”).

Enterprise governance and access control

Intelligent retrieval must never undermine 
TYPO3’s security boundaries:

  • Pre- and post-ACL filters: Front-end user groups (fe_groups), rootline inheritance and time-based access permissions (starttime/endtime) are checked deterministically before and after the vector search.
  • Data protection in accordance with the GDPR: The vector index and language models run on isolated European infrastructure, with no data transfer to third parties.
  • No context leakage: Conversations remain strictly session-bound to prevent unauthorised access via the chat history.

Conclusion


Multilingualism does not require a massive translation effort for static archives. With a cross-lingual retrieval approach, you can make existing German specialist PDFs and sub-portals directly accessible to international users – accurately, efficiently and in an audit-proof manner.