Residual Decoder Adapter Boosting AR Text Rendering without Retraining the Tokenizer

1Central South University   2University of Oxford   3Microsoft Research
* Equal contribution    † Corresponding author
TL;DR RDA is a lightweight plug-and-play decoder adapter that can be plugged into existing visual tokenizers and removed without changing the token space, tokenizer, or AR model. It broadly improves text rendering across AR-based unified generation models with minimal extra cost.

Abstract

Visual Autoregressive (AR) models generate images by predicting discrete tokens that are decoded by a visual tokenizer. Despite demonstrating strong overall image generation ability, they still underperform on text rendering—producing blur strokes and disrupted letter shapes. We trace this limitation to the visual tokenizer, which struggles to reconstruct fine-grained detail.

Improving the tokenizer is straightforward but expensive, as it necessitates retraining both the tokenizer and the AR model. Can we improve text rendering performance of AR models without retraining the existing tokenizer and AR model?

To achieve this, we propose the Residual Decoder Adapter (RDA) that upgrades an existing tokenizer post-hoc without changing its token space. Specifically, it refines the decoder output of the visual tokenizer by introducing two novel components: (i) a paired Hint Codebook that shares the token distribution with the original one; and (ii) a Residual Decoder that learns the tiny differences (residual) between the reconstructed image and the ground-truth images in pixel space. This residual design allows us to enhance the tokenizer non-invasively while preserving compatibility with prior AR models.

RDA substantially improves text rendering by a large margin. For instance, we boost finetuned Janus-Pro OCR accuracy from 24.52% → 58.26% (TextVisionBlend) and from 12.75% → 36.81% (StyledTextSynth) on the competitive TextAtlas benchmark.

Motivation

Intuition: (a) original ecosystem, (b) retrain tokenizer approach, (c) RDA plug-in approach

Method

Overview of the proposed Residual Decoder Adapter (RDA) architecture

RDA keeps the original tokenizer frozen and attaches a lightweight Shared-ID Hint Codebook and Residual Decoder to recover fine-grained details lost during quantization—without changing the token space or retraining the AR model.

Results

Benchmark Results

A unified benchmark board with dataset buttons: switch between the overview table and focused benchmark views without leaving the Results section.

w/o before RDA w/ after RDA Δ RDA gain CER ↓: lower is better
Dataset view Click a benchmark to focus the table

Tokenizer Reconstruction & Ablation Studies

A compact analysis board for Table 3–6. Switch between reconstruction results and each ablation setting while keeping the same polished dataset-board layout.

★ best value in the current table ✓ enabled / AR-free ✗ disabled / not AR-free LPIPS ↓: lower is better
Analysis view Click to switch result tables

AR Model Generation

Qualitative generation results of AR models before and after applying RDA

Tokenizer Reconstruction

Reconstruction performance of image tokenizer equipped with RDA

BibTeX

@inproceedings{mao2026rda,
  title     = {Residual Decoder Adapter: Boosting AR Text Rendering without Retraining the Tokenizer},
  author    = {Mao, Dongxing and Wang, Alex Jinpeng and Tang, Jiahao and Lin, Kevin Qinghong and
               Li, Linjie and Yang, Zhengyuan and Wang, Lijuan and Li, Min and Tan, Jingru},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year      = {2026}
}