The year/Topics/Multimodal

Topic area

Multimodal

Every collection across multimodal.

Papers
144
Research labs
9
Official code
109

51100 of 144 papers in this topic area

51

arXiv.org

PixelSmile: Toward Fine-Grained Facial Expression Editing

PixelSmile is a diffusion-based framework for fine-grained facial expression editing, addressing semantic overlap between expressions like fear-surprise and anger-disgust. The authors construct the Flex Facial Expression (FFE) dataset with 60,000 images (real and anime) annotated with continuous 12-dimensional affective scores, and establish FFE-Bench to…

Jiabin Hua, Hengyuan Xu, Aojie Li, Wei Cheng, et al.
Published
Mar 2026
Citations
0
Code
Not linked
52

Mistral AI

Voxtral TTS

Voxtral TTS is a multilingual zero-shot text-to-speech model that generates natural speech from as little as 3 seconds of reference audio. It uses a hybrid architecture: an autoregressive decoder backbone (based on Ministral 3B) predicts semantic speech tokens, while a flow-matching transformer predicts acoustic tokens. The tokens are produced by Voxtral…

Mistral-AI, :, Alexander H. Liu, Alexis Tacnet, et al.
Published
Mar 2026
Citations
0
Code
Not linked
53

arXiv.org

RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models

RealRestorer is an open-source image restoration model designed to handle diverse real-world degradations, including blur, rain, noise, low-light, moiré patterns, haze, compression artifacts, reflection, and flare. The authors construct a large-scale dataset with a synthesis pipeline that combines synthetic and real-world degradation data, and they…

Yufeng Yang, Xianfang Zeng, Zhangqi Jiang, Fukun Yin, et al.
Published
Mar 2026
Citations
3
Code
Not linked
54

arXiv.org

Calibri: Enhancing Diffusion Transformers via Parameter-Efficient Calibration

The paper introduces Calibri, a parameter-efficient method to enhance Diffusion Transformers (DiTs) by calibrating block outputs with learned scaling parameters. The authors show that selectively disabling or re-weighting DiT blocks can improve generation quality, leading to a black-box optimization problem solved via CMA-ES, modifying only ~10^2…

Danil Tokhchukov, Aysel Mirzoeva, Andrey Kuznetsov, Konstantin Sobolev
Published
Mar 2026
Citations
0
Code
59 stars
55

arXiv.org

MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding

MinerU-Diffusion is a 2.5B-parameter diffusion-based framework for document OCR that replaces autoregressive (AR) decoding with block-wise parallel diffusion denoising under visual conditioning. The authors argue that left-to-right causal generation is an artifact of serialization, not intrinsic to OCR, and propose inverse rendering via diffusion. The…

Hejun Dong, Junbo Niu, Bin Wang, Weijun Zeng, et al.
Published
Mar 2026
Citations
7
Code
628 stars
56

arXiv.org

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

VEGA-3D is a plug-and-play framework that repurposes pre-trained video generation models as Latent World Simulators to provide implicit 3D priors for Multimodal Large Language Models (MLLMs), addressing their spatial blindness. The method extracts spatiotemporal features from intermediate noise levels of a frozen video diffusion model (e.g., Wan2.1-T2V)…

Xianjin Wu, Dingkang Liang, Tianrui Feng, Kui Xia, et al.
Published
Mar 2026
Citations
10
Code
421 stars
57

arXiv.org

Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs

The paper presents AwaRes, a framework for efficient vision-language model (VLM) inference that processes a low-resolution global image and uses tool-calling to retrieve only the high-resolution crops needed for a query. AwaRes trains a coupled-decision policy (CDP) that jointly decides whether to escalate resolution and which crops to request. Supervision…

Nimrod Shabtay, Moshe Kimhi, Artem Spector, Sivan Haray, et al.
Published
Mar 2026
Citations
1
Code
11 stars
58

arXiv.org

Qianfan-OCR: A Unified End-to-End Model for Document Intelligence

Qianfan-OCR is a 4B-parameter end-to-end vision-language model that unifies document parsing, layout analysis, and understanding, outperforming all end-to-end models on OmniDocBench v1.5 (93.12) and OlmOCR Bench (79.8). It introduces Layout-as-Thought, an optional thinking phase triggered by ⟨think⟩ tokens that generates structured layout representations…

Daxiang Dong, Mingming Zheng, Dong Xu, Chunhua Luo, et al.
Published
Mar 2026
Citations
11
Code
421 stars
59

arXiv.org

Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Penguin-VL, developed by Tencent AI Lab, introduces compact 2B and 8B vision-language models that challenge the reliance on contrastive pretraining (e.g., CLIP/SigLIP) for vision encoders. The authors argue that contrastive learning suppresses fine-grained visual cues needed for reasoning. Instead, they propose Penguin-Encoder, initialized from a text-only…

Boqiang Zhang, Lei Ke, Ruihan Yang, Qi Gao, et al.
Published
Mar 2026
Citations
4
Code
206 stars
60

Meta AI

Beyond Language Modeling: An Exploration of Multimodal Pretraining

This paper presents controlled, from-scratch experiments to clarify the design space of unified multimodal pretraining, using the Transfusion framework (next-token prediction for language, diffusion for vision) on text, video, image-text pairs, and action-conditioned video. Key findings: (1) Representation Autoencoders (RAE), e.g., SigLIP 2, provide a…

Shengbang Tong, David Fan, John Nguyen, Ellis Brown, et al.
Published
Mar 2026
Citations
21
Code
Not linked
61

arXiv.org

UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

UniG2U-Bench is a new benchmark for evaluating whether unified multimodal models (UMMs) benefit from generation when performing understanding tasks. It includes 3,000 samples across 7 categories and 30 subtasks, and evaluates over 30 models, including base VLMs, unified models, and agentic models. The study finds that unified models generally underperform…

Zimo Wen, Boxiu Li, Wanbo Zhang, Junxiang Lei, et al.
Published
Mar 2026
Citations
3
Code
Not linked
62

arXiv.org

Enhancing Spatial Understanding in Image Generation via Reward Modeling

The paper introduces a method to improve spatial understanding in text-to-image generation using reward modeling. The authors construct the SpatialReward-Dataset, containing over 80,000 adversarial preference pairs, where each pair consists of an image correctly depicting complex spatial relationships and a perturbed image violating some relationships,…

Zhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu, et al.
Published
Feb 2026
Citations
0
Code
86 stars
63

Google DeepMind

Unified Latents (UL): How to train your latents

Unified Latents (UL) is a framework for learning latent representations regularized by a diffusion prior and decoded by a diffusion model. The encoder outputs a deterministic latent, which is noised to a fixed minimum noise level (log-SNR of 5), linking the encoder's output noise to the prior's precision. The training objective combines a diffusion prior…

Jonathan Heek, Emiel Hoogeboom, Thomas Mensink, Tim Salimans
Published
Feb 2026
Citations
15
Code
Not linked
64

arXiv.org

BitDance: Scaling Autoregressive Generative Models with Binary Tokens

BitDance is a scalable autoregressive image generation model that predicts binary visual tokens instead of codebook indices, scaling the vocabulary to 2^256 states. It introduces a binary diffusion head to sample from this large discrete space, and a next-patch diffusion method for parallel multi-token prediction. On ImageNet 256×256, BitDance achieves an…

Yuang Ai, Jiaming Han, Shaobin Zhuang, Weijia Mao, et al.
Published
Feb 2026
Citations
5
Code
481 stars
65

arXiv.org

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs

MedXIAOHE is a medical vision-language foundation model from ByteDance that achieves state-of-the-art performance across 30+ medical benchmarks, surpassing leading closed-source systems like GPT-5.2 Thinking and Gemini 3.0 Pro. The model uses a Seed-ViT vision encoder and a large language model, trained via a three-stage pipeline. Continual pretraining…

Baorong Shi, Bo Cui, Boyuan Jiang, Deli Yu, et al.
Published
Feb 2026
Citations
3
Code
Not linked
66

arXiv.org

DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing

DeepGen 1.0 is a lightweight 5B-parameter unified multimodal model (3B VLM + 2B DiT) for image generation and editing, achieving performance competitive with or surpassing much larger models. It introduces Stacked Channel Bridging (SCB), which fuses features from six uniformly distributed VLM layers with learnable 'think tokens' to provide the DiT with…

Dianyi Wang, Ruihang Li, Feng Han, Chaofan Ma, et al.
Published
Feb 2026
Citations
15
Code
584 stars
67

arXiv.org

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

The paper introduces Region-to-Image Distillation (R2I), a method that internalizes the benefits of inference-time zooming into a single forward pass of a multimodal large language model (MLLM). R2I synthesizes fine-grained VQA data by zooming into micro-cropped regions, using strong teacher models to generate high-consensus question-answer pairs, and then…

Lai Wei, Liangbo He, Jun Lan, Lingzhong Dong, et al.
Published
Feb 2026
Citations
23
Code
184 stars
68

arXiv.org

GENIUS: Generative Fluid Intelligence Evaluation Suite

The paper introduces GENIUS, the first benchmark for evaluating Generative Fluid Intelligence (GFI) in Unified Multimodal Models (UMMs), distinguishing it from Crystallized Intelligence (CI). GFI is formalized into three primitives: Inducing Implicit Patterns, Executing Ad-hoc Constraints, and Adapting to Contextual Knowledge. The benchmark comprises 510…

Ruichuan An, Sihan Yang, Ziyu Guo, Wei Dai, et al.
Published
Feb 2026
Citations
12
Code
43 stars
69

arXiv.org

P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads

The P1-VL technical report introduces a family of open-source vision-language models (VLMs) designed for advanced scientific reasoning, specifically targeting physics Olympiad problems. The models are trained exclusively via reinforcement learning (RL) using a curriculum that progressively increases problem difficulty and expands exploration space,…

Yun Luo, Futing Wang, Qianjia Cheng, Fangchen Yu, et al.
Published
Feb 2026
Citations
5
Code
15 stars
70

arXiv.org

ERNIE 5.0 Technical Report

ERNIE 5.0 is a natively autoregressive foundation model from Baidu for unified multimodal understanding and generation across text, image, video, and audio. All modalities are trained from scratch under a unified next-group-of-tokens prediction objective, using an ultra-sparse mixture-of-experts (MoE) architecture with modality-agnostic expert routing. A…

Haifeng Wang, Hua Wu, Tian Wu, Yu Sun, et al.
Published
Feb 2026
Citations
8
Code
Not linked
71

arXiv.org

UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing

UniReason is a unified framework for text-to-image (T2I) generation and image editing, built on the Bagel architecture. It addresses limitations of existing methods by incorporating world knowledge-enhanced textual reasoning before synthesis and fine-grained editing-like visual refinement after initial generation. The framework uses two complementary…

Dianyi Wang, Chaofan Ma, Feng Han, Size Wu, et al.
Published
Feb 2026
Citations
6
Code
144 stars
72

arXiv.org

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

The paper addresses the Modality Gap in multimodal contrastive learning, where embeddings of different modalities for the same semantics occupy offset regions. Prior methods rely on isotropic assumptions, which are flawed. The authors propose the Fixed-frame Modality Gap Theory, decomposing the gap into stable biases (PMB, POB) and anisotropic residuals,…

Xiaomin Yu, Yi Xin, Yuhui Zhang, Wenjie Zhang, et al.
Published
Feb 2026
Citations
9
Code
76 stars
73

arXiv.org

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding

This paper presents the first systematic study on the effectiveness of Vision Language Models (MLLMs) for code understanding by representing source code as images. The authors evaluate seven MLLMs across four tasks (code summarization, completion, clone detection, and question answering) with compression ratios from 1x to 8x and rendering strategies…

Yuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen, et al.
Published
Feb 2026
Citations
16
Code
30 stars
74

DeepSeek

DeepSeek-OCR 2: Visual Causal Flow

DeepSeek-OCR 2 introduces DeepEncoder V2, a novel vision encoder that replaces the CLIP component with a compact LLM (Qwen2-0.5B) to enable causal reordering of visual tokens, mimicking human visual scanning. The encoder uses a dual attention mask: bidirectional for visual tokens and causal for learnable query tokens, allowing queries to attend to all…

Haoran Wei, Yaofeng Sun, Yukun Li
Published
Jan 2026
Citations
53
Code
3.2K stars
75

Moonshot AI

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

WorldVQA is a benchmark introduced to evaluate the atomic visual world knowledge of Multimodal Large Language Models (MLLMs), decoupling visual knowledge retrieval from reasoning. It comprises 3,500 VQA pairs across nine semantic categories, from common head-class entities to long-tail rarities. The benchmark follows four design principles: atomic…

Runjie Zhou, Youbo Shao, Haoyu Lu, Bowei Xing, et al.
Published
Jan 2026
Citations
4
Code
121 stars
76

arXiv.org

Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

The paper introduces SpatialGenEval, a benchmark for evaluating the spatial intelligence of text-to-image (T2I) models. It uses 1,230 long, information-dense prompts across 25 real-world scenes, each integrating 10 spatial sub-domains (object, attribute, position, orientation, layout, comparison, proximity, occlusion, motion, causal) and paired with 10…

Zengbin Wang, Xuecai Hu, Yong Wang, Feng Xiong, et al.
Published
Jan 2026
Citations
8
Code
132 stars
77

arXiv.org

Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders

This paper investigates scaling Representation Autoencoders (RAEs) for large-scale text-to-image (T2I) generation. The authors train RAE decoders on a frozen SigLIP-2 encoder using web, synthetic, and text-rendering data, finding that data composition is crucial for text reconstruction. They show that dimension-dependent noise scheduling remains essential,…

Shengbang Tong, Boyang Zheng, Ziteng Wang, Bingda Tang, et al.
Published
Jan 2026
Citations
40
Code
255 stars
78

Qwen

Qwen3-TTS Technical Report

The Qwen3-TTS technical report introduces a family of multilingual, controllable, robust, and streaming text-to-speech models. Trained on over 5 million hours of speech data across 10 languages, Qwen3-TTS supports 3-second voice cloning, description-based voice design, and fine-grained control. It uses a dual-track LM architecture with two tokenizers:…

Hangrui Hu, Xinfa Zhu, Ting He, Dake Guo, et al.
Published
Jan 2026
Citations
88
Code
13K stars
79

arXiv.org

Urban Socio-Semantic Segmentation with Vision-Language Reasoning

The paper introduces urban socio-semantic segmentation, targeting entities defined by social attributes (e.g., schools, parks) rather than physical ones. The authors present SocioSeg, a benchmark with over 13,000 samples, organizing labels into three hierarchical tasks: socio-name, socio-class, and socio-function. A key innovation is representing…

Yu Wang, Yi Wang, Rui Dai, Yujie Wang, et al.
Published
Jan 2026
Citations
2
Code
175 stars
80

arXiv.org

STEP3-VL-10B Technical Report

Step3-VL-10B is a 10B-parameter open-source multimodal foundation model that rivals or surpasses models 10-20x larger, such as GLM-4.6V-106B and Qwen3-VL-235B, and proprietary systems like Gemini 2.5 Pro. It achieves 92.2% on MMBench, 80.11% on MMMU, 94.43% on AIME2025, and 75.95% on MathVision. The model uses a unified pre-training strategy on 1.2T…

Ailin Huang, Chengyuan Yao, Chunrui Han, Fanqi Wan, et al.
Published
Jan 2026
Citations
26
Code
412 stars
81

arXiv.org

BabyVision: Visual Reasoning Beyond Language

The paper introduces BABYVISION, a benchmark to evaluate core visual abilities in Multimodal LLMs (MLLMs) that are independent of linguistic knowledge, targeting skills humans develop before language. It contains 388 questions across 22 subtypes in four categories: Fine-grained Discrimination, Visual Tracking, Spatial Perception, and Visual Pattern…

Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He, et al.
Published
Jan 2026
Citations
24
Code
238 stars
82

arXiv.org

VIBE: Visual Instruction Based Editor

VIBE is a compact, high-throughput instruction-based image editing pipeline that combines a 2B-parameter Qwen3-VL model for instruction interpretation with a 1.6B-parameter Sana1.5 diffusion model for image generation. The architecture uses channel-wise concatenation for reference image guidance and learnable meta-tokens processed by a lightweight…

Grigorii Alekseenko, Aleksandr Gordeev, Irina Tolstykh, Bulat Suleimanov, et al.
Published
Jan 2026
Citations
2
Code
55 stars
83

arXiv.org

NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation

NextFlow is a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image tokens. It uses a dual-codebook tokenizer for semantic and pixel-level features, and adopts next-scale prediction for visual generation instead of raster-scan, enabling 1024x1024 image generation in 5 seconds. The model retains next-token prediction…

Huichao Zhang, Liao Qu, Yiheng Liu, Hang Chen, et al.
Published
Jan 2026
Citations
12
Code
331 stars
84

arXiv.org

MOSS Transcribe Diarize Technical Report

MOSS Transcribe Diarize is a unified multimodal large language model for Speaker-Attributed, Time-Stamped Transcription (SATS), jointly performing word recognition, speaker attribution, and timestamp prediction in a single end-to-end pass. It uses a 128k-token context window to process up to 90 minutes of audio without chunking, preserving long-range…

MOSI. AI, :, Donghua Yu, Zhengyuan Lin, et al.
Published
Jan 2026
Citations
5
Code
Not linked
85

arXiv.org

The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding

The paper introduces the Prism Hypothesis, which posits that multimodal data can be understood through a shared frequency spectrum: semantic encoders capture low-frequency components (abstract meaning), while pixel encoders retain high-frequency details (fine texture). This is supported by experiments showing that text-image retrieval relies on low…

Weichen Fan, Haiwen Diao, Quan Wang, Dahua Lin, et al.
Published
Dec 2025
Citations
15
Code
209 stars
86

AAAI Conference on Artificial Intelligence

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

Robust-R1 is a framework designed to improve the robustness of Multimodal Large Language Models (MLLMs) against real-world visual degradations. Unlike existing methods that rely on implicit training or adaptation of visual encoders, Robust-R1 explicitly models degradations through structured reasoning chains. The approach consists of three stages:…

Jiaqi Tang, Jianmin Chen, Wei Wei, Xiaogang Xu, et al.
Published
Dec 2025
Citations
8
Code
530 stars
87

Qwen

Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition

Qwen-Image-Layered is an end-to-end diffusion model that decomposes a single RGB image into multiple semantically disentangled RGBA layers, enabling consistent image editing where each layer can be independently manipulated. The model introduces three key components: an RGBA-VAE that unifies latent representations for RGB and RGBA images, a VLD-MMDiT…

Shengming Yin, Zekai Zhang, Zecheng Tang, Kaiyuan Gao, et al.
Published
Dec 2025
Citations
24
Code
2.1K stars
88

MiniMax

Towards Scalable Pre-training of Visual Tokenizers for Generation

The paper introduces VTP, a visual tokenizer pre-training framework that integrates image-text contrastive learning, self-supervised learning (MIM and self-distillation), and reconstruction losses to address the 'pre-training scaling problem' in latent diffusion models. The authors argue that reconstruction-only training biases the latent space toward…

Jingfeng Yao, Yuda Song, Yucong Zhou, Xinggang Wang
Published
Dec 2025
Citations
23
Code
497 stars
89

arXiv.org

TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows

TWINFLOW is a framework for training one-step generative models on large multimodal models, addressing the inefficiency of multi-step diffusion and flow matching models that require 40-100 NFEs. Existing few-step methods either rely on auxiliary trained models (e.g., GAN discriminators) or frozen teachers, causing instability and memory overhead, or…

Zhenglin Cheng, Peng Sun, Jianguo Li, Tao Lin
Published
Dec 2025
Citations
16
Code
537 stars
90

Meta AI

TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models

Tuna is a native unified multimodal model (UMM) that creates a unified continuous visual representation by cascading a VAE encoder with a representation encoder (SigLIP 2). This design avoids the representation format mismatches of decoupled models, improving both understanding and generation. The model uses an LLM decoder (Qwen2.5) for autoregressive text…

Zhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou, et al.
Published
Dec 2025
Citations
31
Code
94 stars
91

arXiv.org

Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights

The paper introduces Envision, a benchmark for evaluating text-to-image (T2I) and unified multimodal models (UMMs) on causal, multi-image event generation. It addresses the limitation of static single-image benchmarks by proposing chained text-to-multi-image generation with 1,000 four-stage prompts across six scientific and humanities domains.…

Juanxi Tian, Siyuan Li, Conghui He, Lijun Wu, et al.
Published
Dec 2025
Citations
1
Code
32 stars
92

arXiv.org

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Z-Image is a 6B-parameter image generation foundation model from Alibaba Group, built on a Scalable Single-Stream Diffusion Transformer (S3-DiT). It challenges the 'scale-at-all-costs' paradigm by optimizing data infrastructure, architecture, training, and inference. The full training workflow costs 314K H800 GPU hours (~$628K). Z-Image-Turbo, a distilled…

Z-Image Team, Huanqia Cai, Sihan Cao, Ruoyi Du, et al.
Published
Nov 2025
Citations
206
Code
12K stars
93

Qwen

Qwen3-VL Technical Report

Qwen3-VL is a state-of-the-art vision-language model family from the Qwen team, released on December 1, 2025. It supports interleaved contexts up to 256K tokens and comes in dense (2B/4B/8B/32B) and MoE (30B-A3B/235B-A22B) variants. Key architectural innovations include interleaved-MRoPE for balanced spatial-temporal encoding, DeepStack for multi-level ViT…

Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, et al.
Published
Nov 2025
Citations
1.8K
Code
20K stars
94

arXiv.org

DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation

DeCo is a frequency-decoupled pixel diffusion framework for end-to-end image generation. It addresses the challenge of pixel diffusion models jointly modeling high-frequency signals and low-frequency semantics in a single diffusion transformer (DiT), which slows training and inference. DeCo uses a DiT to model low-frequency semantics from downsampled…

Zehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang, et al.
Published
Nov 2025
Citations
37
Code
238 stars
95

arXiv.org

MedSAM3: Delving into Segment Anything with Medical Concepts

MedSAM-3 adapts the SAM 3 architecture for medical image and video segmentation, enabling Promptable Concept Segmentation (PCS) via open-vocabulary text prompts. The model is fine-tuned on medical images paired with concise concept phrases (≤3 words) across modalities like X-ray, MRI, Ultrasound, CT, and video. The MedSAM-3 Agent integrates a Multimodal…

Anglin Liu, Rundong Xue, Xu R. Cao, Yifan Shen, et al.
Published
Nov 2025
Citations
28
Code
315 stars
96

Meta AI

SAM 3: Segment Anything with Concepts

SAM 3 is a unified model for promptable concept segmentation (PCS) in images and videos, accepting noun phrases, image exemplars, or both as prompts to detect, segment, and track all matching instances. It decouples recognition and localization via a presence head, improving detection accuracy. A data engine with human and AI verifiers produced 4M unique…

Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, et al.
Published
Nov 2025
Citations
711
Code
11K stars
97

arXiv.org

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

Kandinsky 5.0 is a family of foundation models for high-resolution image and video synthesis, comprising three model line-ups: Image Lite (6B parameters), Video Lite (2B), and Video Pro (19B). The models are built on a unified latent diffusion architecture with a CrossDiT backbone and trained using flow matching. Key innovations include the NABLA sparse…

Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko, Denis Parkhomenko, et al.
Published
Nov 2025
Citations
9
Code
805 stars
98

arXiv (Cornell University)

Back to Basics: Let Denoising Generative Models Denoise

The paper argues that denoising diffusion models should directly predict clean images (x-prediction) rather than noise (epsilon-prediction) or velocity (v-prediction), as natural data lies on a low-dimensional manifold while noised quantities do not. The authors propose 'Just image Transformers' (JiT), a plain Vision Transformer applied to large pixel…

Tianhong Li, Kaiming He
Published
Nov 2025
Citations
4
Code
2.5K stars
99

arXiv.org

Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data

Uni-MoE-2.0-Omni is a fully open-source omnimodal large model (OLM) built from the dense Qwen2.5-7B LLM, designed for unified understanding, reasoning, and generation across text, image, audio, and video. Its architecture introduces a dynamic-capacity Mixture-of-Experts (MoE) with shared, routed, and null experts for efficient computation and modality…

Yunxin Li, Xinyu Chen, Shenyuan Jiang, Haoyuan Shi, et al.
Published
Nov 2025
Citations
21
Code
1.1K stars
100

arXiv.org

One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter for Your Diffusion Models

The paper introduces the Latent Upscaler Adapter (LUA), a lightweight module that performs super-resolution directly on a diffusion model's latent code before VAE decoding, enabling high-resolution image synthesis without retraining the generator or adding diffusion stages. LUA uses a shared SwinIR-style backbone with scale-specific pixel-shuffle heads for…

Aleksandr Razin, Danil Kazantsev, Ilya Makarov
Published
Nov 2025
Citations
1
Code
34 stars