The year/Topics/Multimodal

Topic area

Multimodal

Every collection across multimodal.

Papers
144
Research labs
9
Official code
109

150 of 144 papers in this topic area

01

Independent research

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

SwanTale is a unified model for multi-speaker speech and audio generation supporting both zero-shot and instruct tasks. It introduces SwanData-Caption, a data pipeline that cleans raw audio, adds targeted synthetic coverage (elderly speech, short utterances, challenging pronunciations), and annotates multi-level captions (environment, speakers, content).…

Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, et al.
Published
Aug 2026
Citations
0
Code
Not linked
02

Independent research

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

ReDesign is an agentic framework that recovers editable design structures (e.g., Figma files) from raster images by growing a layer hierarchy through tool composition. A VLM controller selects actions from a fixed set (text extraction, multi-layer decomposition, connected component labeling, detection/segmentation, vectorization) to expand nodes, with…

Jooyeol Yun, Jintae Park, Hyesu Lim, Junha Hyung, et al.
Published
Jul 2026
Citations
1
Code
177 stars
03

Moonshot AI

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

PerceptionBench is a benchmark introduced by Moonshot AI to evaluate atomic visual perception in Multimodal Large Language Models (MLLMs). It addresses limitations of existing benchmarks that conflate perception with reasoning or knowledge. The benchmark was constructed bottom-up: failures of frontier MLLMs on 42 existing benchmarks were attributed to…

Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, et al.
Published
Jul 2026
Citations
0
Code
170 stars
04

Independent research

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

This paper analyzes on-policy distillation (OPD) for diffusion models under classifier-free guidance (CFG). The authors show that the naive objective of matching CFG-composed velocities is under-identified at the branch level, allowing positive- and negative-branch errors to cancel. They identify a failure mode, Negative Branch Asymmetry (NBA), which…

Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, et al.
Published
Jul 2026
Citations
0
Code
8 stars
05

Independent research

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

Mage-Flow is a compact 4B-parameter generative stack from Microsoft for efficient text-to-image generation and instruction-based image editing. It comprises Mage-VAE, a lightweight latent tokenizer using one-step diffusion-style encoding/decoding with anchor-latent regularization, and a Native-Resolution Multimodal Diffusion Transformer (NR-MMDiT) trained…

Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, et al.
Published
Jul 2026
Citations
1
Code
Not linked
06

Independent research

OvisOCR2 Technical Report

OvisOCR2 is a 0.8B end-to-end document parsing model that converts document page images into Markdown, covering text, formulas, tables, and visual regions. It uses a data engine combining filtered real-document annotations with synthetic pages generated from HTML sources. Training includes supervised fine-tuning, reinforcement learning (GRPO) on a 4B…

Shiyin Lu, Yinglun Li, Yu Xia, Yuhui Chen, et al.
Published
Jul 2026
Citations
0
Code
Not linked
07

Independent research

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

Boogu-Image-0.1 is an open-source family of unified multimodal understanding and generation models (Base, Turbo, Edit, Edit-Turbo) that achieves competitive performance in text-to-image generation, fast inference, instruction-based editing, and bilingual text rendering. The authors argue that strengthening understanding—via a stronger multimodal encoder…

Guoxuan Chen, Chufeng Xiao, Haoran Yang, Siyue Xie, et al.
Published
Jul 2026
Citations
0
Code
Not linked
08

Independent research

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

MonkeyOCRv2 is a visual-text foundation model for document AI, addressing the mismatch between natural-image encoders and document images. The authors construct MonkeyDoc v2, a 113-million-image pretraining corpus spanning 17 languages, and propose a dual-objective pretraining strategy combining image-to-text generation with pixel-level reconstruction to…

Yuliang Liu, Zhang Li, Ziyang Zhang, Shuo Zhang, et al.
Published
Jul 2026
Citations
0
Code
598 stars
09

Independent research

SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding

SynthDocBench is a fully synthetic benchmark for long-context visual document understanding, designed to systematically control factors like document length, layout, modality, and question type. It comprises 200 reports (avg. 51.1 pages, 16.7 charts) and 1,788 questions across three subsets: chart-reading, complex multi-hop, and cross-modal. Documents are…

Abhigya Verma, Khyati Mahajan, Amit Kumar Saha, Shruthan Radhakrishna, et al.
Published
Jul 2026
Citations
0
Code
8 stars
10

Independent research

Scalable Visual Pretraining for Language Intelligence

This paper introduces Visual Pretraining (VP), a framework that trains foundation models directly on raw document images without text extraction or image-text pairing, using a next-visual-latent prediction objective. VP consistently outperforms text-only pretraining (TP) on scientific reasoning benchmarks across multiple backbones (Qwen3.5, Qwen3, Llama3.2…

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, et al.
Published
Jul 2026
Citations
0
Code
Not linked
11

Independent research

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

ResearchStudio-Reel is a native-editable dissemination workspace that automates the last mile of research communication, turning a single paper PDF into a print-ready conference poster, a narration-aligned talk video, and a bilingual blog. Implemented as five composable skills in Claude Code and Codex, it uses a shared Paper2Assets extractor to ground all…

Lingao Xiao, Yalun Dai, Yangyu Huang, Qihao Zhao, et al.
Published
Jul 2026
Citations
1
Code
Not linked
12

Independent research

DanceOPD: On-Policy Generative Field Distillation

DanceOPD is an on-policy generative field distillation framework for flow-matching image generation models, designed to compose multiple capabilities (e.g., text-to-image, local editing, global editing) into a single student model. The method treats each frozen capability source as a velocity field over a shared state space and addresses three key…

Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, et al.
Published
Jun 2026
Citations
2
Code
395 stars
13

Independent research

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Wan-Streamer is a native-streaming, end-to-end interactive foundation model from Alibaba Group designed for real-time, low-latency, full-duplex audio-visual interaction. It models language, audio, and video as both input and output within a single Transformer, using block-causal attention for incremental streaming. Unlike cascaded systems, it does not rely…

Lianghua Huang, Zhi-Fan Wu, Wei Wang, Yupeng Shi, et al.
Published
Jun 2026
Citations
4
Code
Not linked
14

Independent research

Unlimited OCR Works

Baidu's Unlimited OCR introduces Reference Sliding Window Attention (R-SWA) to enable one-shot long-horizon document parsing. R-SWA replaces all attention layers in the decoder of DeepSeek OCR, allowing each generated token to attend to all reference tokens (visual and prompt) and a causal sliding window of the previous 128 output tokens. This maintains a…

Youyang Yin, Huanhuan Liu, YY, Qunyi Xie, et al.
Published
Jun 2026
Citations
4
Code
22K stars
15

Independent research

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

Moebius is a highly efficient lightweight image inpainting framework that rivals the generation quality of 10B-level industrial models like FLUX.1-Fill-Dev while using only 0.22B parameters (less than 2% of FLUX's 11.9B) and delivering over 15× faster total inference time. To overcome the representation bottleneck from extreme structural compression, the…

Kangsheng Duan, Ziyang Xu, Wenyu Liu, Xiaohu Ruan, et al.
Published
Jun 2026
Citations
0
Code
508 stars
16

Independent research

InterleaveThinker: Reinforcing Agentic Interleaved Generation

InterleaveThinker is a multi-agent framework that endows existing image generators with interleaved text-image generation capabilities, addressing visual over-reliance and step-wise error accumulation in Unified Multimodal Models (UMMs). It uses a Planner agent to pre-plan the full instruction sequence, a Generator (e.g., FLUX.2-klein-9B) to execute steps,…

Dian Zheng, Harry Lee, Manyuan Zhang, Kaituo Feng, et al.
Published
Jun 2026
Citations
0
Code
208 stars
17

Independent research

Kwai Keye-VL-2.0 Technical Report

The report introduces Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model with 30B total parameters and 3B active, designed for long-video understanding and agentic intelligence. It is the first to adapt DeepSeek Sparse Attention (DSA) to GQA-based architectures, enabling lossless 256K context processing. The model…

Kwai Keye Team, Bin Wen, Changyi Liu, Chengru Song, et al.
Published
Jun 2026
Citations
0
Code
809 stars
18

Independent research

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

Multimodal Large Language Models (MLLMs) degrade under real-world visual corruptions. Existing robustness methods are limited: black-box feature alignment lacks interpretability, and text-based reasoning cannot restore pixel-level details. This paper proposes Robust-U1, a framework that equips MLLMs with explicit visual self-recovery capability. It uses a…

Jiaqi Tang, Jianmin Chen, Youyang Zhai, Wei Wei, et al.
Published
Jun 2026
Citations
1
Code
363 stars
19

Independent research

Audio Interaction Model

The paper introduces Audio-Interaction, a unified streaming audio language model that operates via an always-on perceive–decide–respond loop, listening to continuous audio and deciding when to respond or remain silent. It addresses limitations of offline LALMs and task-specific streaming models by unifying capabilities like real-time ASR, translation,…

Zhifei Xie, Zihang Liu, Ze An, Xiaobin Hu, et al.
Published
Jun 2026
Citations
0
Code
575 stars
20

Independent research

Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models

The paper introduces Imaginative Perception Tokens (IPTs), intermediate visual representations that externalize what a VLM would perceive under an alternative spatial configuration, to improve spatial reasoning. Three tasks requiring imaginative perception are formulated: Perspective Taking (PET), Path Tracing (PT), and Multiview Counting (MVC), with…

Mahtab Bigverdi, Linjie Li, Weikai Huang, Yiming Liu, et al.
Published
Jun 2026
Citations
0
Code
102 stars
21

arXiv.org

Representation Forcing for Bottleneck-Free Unified Multimodal Models

The paper introduces Representation Forcing (RF), a technique for unified multimodal models (UMMs) that eliminates the need for a separately pretrained VAE in image generation. RF trains the decoder to autoregressively predict discrete visual representation tokens, derived from the model's own understanding encoder via online vector quantization, before…

Yuqing Wang, Zhijie Lin, Ceyuan Yang, Yang Zhao, et al.
Published
May 2026
Citations
3
Code
Not linked
22

Independent research

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

The paper introduces SwanVoice, a zero-shot text-to-speech (TTS) model for expressive long-form monologue and dialogue synthesis with 1–4 speakers. It addresses limitations of stitching monologue outputs for dialogue, which breaks acoustic consistency and affective continuity. The authors build SwanData-Speech, a data pipeline that processes 2.59 million…

Ruiqi Li, Yu Zhang, Changhao Pan, Ke Lei, et al.
Published
May 2026
Citations
2
Code
Not linked
23

arXiv.org

Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs

The paper introduces CRAFTER, a multi-agent harness for generating scientific figures from diverse inputs, and CRAFTEDITOR, which converts raster outputs into editable SVGs. CRAFTER uses five agents (intent reasoner, plan generator, critic, specification refiner, convergence judge) sharing an evolving specification, with mechanisms for diversity-driven…

Haozhe Zhao, Shuzheng Si, Zhenhailong Wang, Zheng Wang, et al.
Published
May 2026
Citations
1
Code
149 stars
24

arXiv.org

From Pixels to Words -- Towards Native One-Vision Models at Scale

NEO-ov is a native vision-language foundation model that unifies single-image, multi-image, video understanding, and spatial intelligence in a single monolithic backbone, eliminating external visual encoders, adapters, and post-hoc fusion. It uses a unified serialization scheme with spatiotemporal attention (THW-decoupled) and Native-RoPE to enable…

Haiwen Diao, Jiahao Wang, Penghao Wu, Yuhao Dong, et al.
Published
May 2026
Citations
1
Code
881 stars
25

NVIDIA

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

LocateAnything is a unified vision-language model for visual grounding and detection that introduces Parallel Box Decoding (PBD). Unlike standard next-token prediction (NTP) which serializes bounding box coordinates into 1D token streams, PBD treats each bounding box as an atomic unit, predicting all its coordinates in a single forward pass. This…

Shihao Wang, Shilong Liu, Yuanguo Kuang, Xinyu Wei, et al.
Published
May 2026
Citations
7
Code
Not linked
26

arXiv.org

CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation

CollectionLoRA is a multi-teacher on-policy distillation framework that consolidates up to 50 diverse visual effects and few-step generation capabilities into a single LoRA, addressing storage overhead, routing latency, and parameter conflicts in conventional multi-LoRA pipelines. The method introduces three key components: Probabilistic Dual-Stream…

Fangtai Wu, Hailong Guo, Shijie Huang, Jiayi Song, et al.
Published
May 2026
Citations
2
Code
29 stars
27

arXiv.org

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

Lens is a 3.8B-parameter text-to-image model that achieves performance competitive with larger state-of-the-art models while using significantly less training compute, requiring only about 19.3% of the compute used by Z-Image. Its efficiency stems from maximizing data information density per batch via the Lens-800M dataset of densely captioned image-text…

Dong Chen, Fangyun Wei, Ziyu Wan, Dongdong Chen, et al.
Published
May 2026
Citations
1
Code
Not linked
28

arXiv.org

Rethinking Cross-Layer Information Routing in Diffusion Transformers

This paper investigates cross-layer information routing in Diffusion Transformers (DiTs), identifying three symptoms of standard residual connections: forward magnitude inflation, backward gradient decay, and block-wise redundancy. The authors propose Diffusion-Adaptive Routing (DAR), a drop-in replacement that uses learnable, timestep-adaptive,…

Chao Xu, Maohua Li, Qirui Li, Yixuan Xu, et al.
Published
May 2026
Citations
2
Code
Not linked
29

arXiv.org

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools

IndusAgent is a tool-augmented agentic framework for open-vocabulary industrial anomaly detection (IAD). It addresses limitations of multimodal large language models (MLLMs), such as domain-misaligned reasoning and structural hallucinations, by combining supervised fine-tuning (SFT) with reinforcement learning (RL). The framework constructs Indus-CoT, a…

Rongbin Tan, Fangfang Lin, Zhenlong Yuan, Min Qiu, et al.
Published
May 2026
Citations
0
Code
Not linked
30

arXiv.org

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

MEGA-ASR is a framework for automatic speech recognition (ASR) in real-world environments, addressing the 'acoustic robustness bottleneck' where models fail under severe, compositional distortions. The authors introduce VOICES-IN-THE-WILD-2M, a large-scale dataset with 7 atomic acoustic effects (noise, far-field, obstructed, echo&reverb, recording,…

Zhifei Xie, Kaiyu Pang, Haobin Zhang, Deheng Ye, et al.
Published
May 2026
Citations
1
Code
1.1K stars
31

arXiv.org

Lance: Unified Multimodal Modeling by Multi-Task Synergy

Lance is a lightweight native unified multimodal model from ByteDance that supports understanding, generation, and editing for both images and videos. It uses a dual-stream mixture-of-experts architecture on shared interleaved multimodal sequences, combining autoregressive language modeling for understanding with flow matching for generation. A…

Fengyi Fu, Mengqi Huang, Shaojin Wu, Yunsheng Jiang, et al.
Published
May 2026
Citations
3
Code
1.3K stars
32

Qwen

Qwen-Image-VAE-2.0 Technical Report

Qwen-Image-VAE-2.0 is a suite of high-compression image VAEs (f16 and f32) designed to overcome the trade-off between compression ratio, reconstruction fidelity, and diffusability. The architecture uses Global Skip Connections (GSC) to preserve fine details, expanded latent channels, and an attention-free, asymmetric encoder-decoder backbone for…

Zekai Zhang, Deqing Li, Kuan Cao, Yujia Wu, et al.
Published
May 2026
Citations
1
Code
69 stars
33

arXiv.org

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

CiteVQA is a benchmark for evaluating Multimodal Large Language Models (MLLMs) on document visual question answering, requiring both correct answers and element-level bounding-box citations. It includes 1,897 questions from 711 PDFs across seven domains and two languages, with an average of 40.6 pages per document. Ground-truth citations are generated via…

Dongsheng Ma, Jiayu Li, Zhengren Wang, Yijie Wang, et al.
Published
May 2026
Citations
3
Code
69 stars
34

arXiv.org

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

SenseNova-U1 is a native unified multimodal model built on the NEO-unify architecture, designed to overcome the traditional divide between understanding and generation. It operates directly on raw pixels and text, eliminating the need for pretrained vision encoders (VEs) and variational autoencoders (VAEs). The model uses a near-lossless visual interface…

Haiwen Diao, Penghao Wu, Hanming Deng, Jiahao Wang, et al.
Published
May 2026
Citations
12
Code
4.5K stars
35

Qwen

Qwen-Image-2.0 Technical Report

Qwen-Image-2.0 is an image generation foundation model that unifies text-to-image (T2I) generation and instruction-based image editing in a single framework. It addresses challenges in ultra-long text rendering (up to 1K tokens), multilingual typography, high-resolution photorealism (native 2K), complex instruction following, and inference efficiency. The…

Bing Zhao, Chenfei Wu, Deqing Li, Hao Meng, et al.
Published
May 2026
Citations
8
Code
Not linked
36

arXiv.org

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

MiniCPM-o 4.5 is a 9B-parameter open-source multimodal large language model (MLLM) designed for real-time full-duplex omni-modal interaction, enabling simultaneous perception and response. It introduces Omni-Flow, a unified streaming framework that aligns multimodal inputs and outputs along a shared temporal axis, converting turn-based interaction into a…

Junbo Cui, Bokai Xu, Chongyi Wang, Tianyu Yu, et al.
Published
Apr 2026
Citations
16
Code
26K stars
37

Z.ai / GLM

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

GLM-5V-Turbo is a native multimodal foundation model for agentic tasks, integrating perception, reasoning, planning, and execution. It introduces CogViT, a vision encoder trained via distillation and contrastive learning, and Multimodal Multi-Token Prediction (MMTP) using a shared <|image|> token for efficiency. The model undergoes joint RL over 30+ task…

GLM-V Team, :, Wenyi Hong, Xiaotao Gu, et al.
Published
Apr 2026
Citations
11
Code
Not linked
38

Meta AI

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Tuna-2 is a native unified multimodal model that performs visual understanding and generation directly from raw pixel embeddings, eliminating pretrained vision encoders such as VAEs and representation encoders. It uses simple patch embedding layers to encode images and a single transformer decoder for joint processing, with pixel-space flow matching for…

Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen, et al.
Published
Apr 2026
Citations
10
Code
744 stars
39

arXiv.org

Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items

Tstars-Tryon 1.0, developed by the Pailitao Team at Alibaba Group, is a commercial-scale virtual try-on system designed for robustness, realism, versatility, and efficiency. It handles challenging real-world cases like extreme poses, lighting variations, and motion blur, while preserving garment details and avoiding synthetic artifacts. The system supports…

Mengting Chen, Zhengrui Chen, Yongchao Du, Zuan Gao, et al.
Published
Apr 2026
Citations
2
Code
Not linked
40

arXiv.org

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation

This paper extends MeanFlow, a one-step generation framework originally designed for class-label conditioning, to flexible text-to-image (T2I) generation. The authors find that directly integrating LLM-based text encoders into MeanFlow with standard training fails, and they identify that high-quality text representations must possess strong semantic…

Chenxi Zhao, Chen Zhu, Xiaokun Feng, Aiming Hao, et al.
Published
Apr 2026
Citations
2
Code
109 stars
41

arXiv.org

Elucidating the SNR-t Bias of Diffusion Probabilistic Models

This paper identifies a Signal-to-Noise Ratio-timestep (SNR-t) bias in Diffusion Probabilistic Models (DPMs), where during inference the SNR of denoised samples becomes misaligned with their timestep due to accumulated prediction and discretization errors. The authors provide empirical evidence and theoretical proof showing that reverse-process samples…

Meng Yu, Lei Sun, Jianhao Zeng, Xiangxiang Chu, et al.
Published
Apr 2026
Citations
2
Code
120 stars
42

arXiv.org

Qwen3.5-Omni Technical Report

Qwen3.5-Omni is a fully omnimodal large language model that scales to hundreds of billions of parameters and supports a 256k context length. It is pretrained on a massive dataset including over 100 million hours of audio-visual content. The model uses a Thinker-Talker architecture with Hybrid-Attention Mixture-of-Experts (MoE) for both components, enabling…

Qwen Team
Published
Apr 2026
Citations
87
Code
Not linked
43

arXiv.org

RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time

The paper introduces RationalRewards, a reasoning-based reward model for visual generation that produces structured, multi-dimensional critiques before assigning scores, unlike traditional scalar reward models. It is trained using Preference-Anchored Rationalization (PARROT), a variational framework that recovers rationales from preference data via…

Haozhe Wang, Cong Wei, Weiming Ren, Jiaming Liu, et al.
Published
Apr 2026
Citations
7
Code
56 stars
44

Independent research

EXAONE 4.5 Technical Report

EXAONE 4.5 is LG AI Research's first open-weight vision-language model, integrating a custom 1.2B-parameter vision encoder with the EXAONE 4.0 32B language backbone. It supports six languages and a 256K token context, achieved via context extension during SFT. The model uses hybrid attention, GQA, 2D RoPE, and MTP for efficiency. Pre-training includes two…

Eunbi Choi, Kibong Choi, Sehyun Chun, Seokhee Hong, et al.
Published
Apr 2026
Citations
0
Code
45 stars
45

arXiv.org

HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Tencent's HY-Embodied-0.5 is a family of vision-language foundation models designed for real-world embodied agents, bridging the gap between general VLMs and physical-world tasks. The suite includes an efficient 2B-activated-parameter model (MoT-2B) for edge deployment and a powerful 32B-activated-parameter model (MoE-A32B) for complex reasoning. Key…

Tencent Robotics X, HY Vision Team, :, Xumin Yu, et al.
Published
Apr 2026
Citations
12
Code
843 stars
46

arXiv.org

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

MinerU2.5-Pro improves document parsing purely through data engineering and training strategy, keeping the 1.2B-parameter architecture of MinerU2.5 unchanged. The authors identify that state-of-the-art models share failure patterns on hard samples, indicating a data bottleneck rather than an architectural one. They build a Data Engine with three…

Bin Wang, Tianyao He, Linke Ouyang, Fan Wu, et al.
Published
Apr 2026
Citations
14
Code
Not linked
47

Meta AI

Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

This paper introduces process-driven image generation, a multi-step paradigm that decomposes text-to-image synthesis into an interleaved reasoning trajectory of textual planning and visual generation. The method uses a recurring four-stage cycle: Plan, Sketch, Inspect, and Refine, where the model generates incremental instructions and scene descriptions,…

Lei Zhang, Junjiao Tian, Zhipeng Fan, Kunpeng Li, et al.
Published
Apr 2026
Citations
3
Code
Not linked
48

arXiv.org

Steerable Visual Representations

The paper introduces Steerable Visual Representations (SteerViT), a method to make pretrained Vision Transformers (ViTs) steerable by natural language. SteerViT injects text into the visual encoder via lightweight gated cross-attention layers (early fusion), unlike late-fusion models like CLIP. It is trained on a referential segmentation pretext task using…

Jona Ruthardt, Manu Gaur, Deva Ramanan, Makarand Tapaswi, et al.
Published
Apr 2026
Citations
2
Code
119 stars
49

arXiv.org

GEMS: Agent-Native Multimodal Generation with Memory and Skills

GEMS is a framework for multimodal generation that uses an agent-native approach to improve performance on complex instructions and specialized tasks. It has three main parts: an Agent Loop that iteratively refines generation through planning, decomposition, generation, verification, and refinement; Agent Memory that stores a persistent, hierarchically…

Zefeng He, Siyuan Huang, Xiaoye Qu, Yafu Li, et al.
Published
Mar 2026
Citations
8
Code
142 stars
50

arXiv.org

LongCat-Next: Lexicalizing Modalities as Discrete Tokens

LongCat-Next, developed by Meituan's LongCat team, introduces the Discrete Native Autoregressive (DiNA) paradigm, which unifies text, vision, and audio into a shared discrete token space, enabling a single autoregressive model to handle all modalities. A key innovation is the Discrete Native Any-resolution Visual Transformer (dNaViT), which uses…

Meituan LongCat Team, Bin Xiao, Chao Wang, Chengjiang Li, et al.
Published
Mar 2026
Citations
18
Code
467 stars