For more than 20 years, data extraction in digital onboarding relied on traditional OCR. Today, the emergence of multimodal generative models is disrupting this established architecture. Should OCR engines be put out to pasture just yet? Does GenAI mark the end of OCR in digital onboarding?

OCR Facing the GenAI Wave: Disruption or Evolution?
For two decades, automatically reading an ID document or proof of address required a dedicated set of tools. Each processing step demands its own algorithm, and each block requires custom code, integration testing, and ongoing maintenance. To add a new document to the catalog, teams must go through lengthy data annotation phases again.
“OCR-based systems perform well, but they prove complex to maintain and heavy to evolve at the pace of business requirements.” — Julien Lerouge, Staff AI Engineer at QuickSign
As early as 2019, multimodal transformer architectures began beating benchmarks by jointly analyzing text and visuals. The real turning point came in 2021 with the emergence of the Donut model, capable of interpreting a document image without going through a textual OCR step.
Generative Models Facing Real-World Realities
The arrival of multimodal models changes the technical approach. These architectures combine visual perception and semantic understanding within a single model.
In terms of development, the promise is straightforward. A functional prototype can be obtained in a few hours using few-shot in-context learning techniques, simply by providing a few guidelines and concrete examples to the model. There is no longer a need to retrain a specific block to extract a field on an unseen document.
“GenAI offers unmatched flexibility to handle complex formats or new alphabets. But in production, the dedicated hardware infrastructure represents a significant operating cost that must be managed carefully.”
— Julien Lerouge, Staff AI Engineer at QuickSign
Traditional OCR remains cost-effective, fast, and proven for perfectly standardized formats. GenAI brings the necessary flexibility to tackle the diversity of European documents.
QuickSign’s Approach: Innovation Under Control
Abruptly replacing a proven infrastructure with a generative model makes no economic or operational sense. The real challenge lies in building a hybrid architecture, capable of activating the right technology block depending on the context.
For financial institutions, selecting models does not rely solely on raw extraction performance. It also depends on:
- Data security and sovereignty: guaranteeing client data protection (GDPR, eIDAS compliance) by avoiding exposing data to third parties and favoring hosting on a controlled sovereign infrastructure.
- The trade-off between quality and speed: balancing model precision with execution time, particularly by fine-tuning small models or distilling into lighter architectures.
- Cost control: preventing unpredictable costs by maintaining infrastructure control, which also helps reduce the environmental footprint.
This is where QuickSign delivers value. We integrate advances in Computer Vision and GenAI within a secure technical core designed to absorb regulatory, economic, and technological complexity. Our teams orchestrate these models, via fine-tuning or distillation, whether deployed on sovereign infrastructures or hosted in Europe, to ensure maximum automation without compromising data security. Our approach enables financial institutions to combine accuracy and execution speed while controlling operational costs and energy consumption.
Discover more in video in “Expert Voices”
This article is drawn directly from the first episode of our Season 2 of Expert Voices. In this two-minute video, Julien Lerouge, Staff AI Engineer at QuickSign, details this technological transition and explains how to balance agility and security in document extraction.
Discover how we secure your onboarding processes on our KYC and compliance page.
Looking to adapt your onboarding journeys to the latest AI advancements?
Contact our experts.
Redacted by Marilou T.