Independent research
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
The paper introduces PerceptionDLM, a multimodal diffusion language model for efficient parallel region perception. It builds on PerceptionDLM-Base, a strong diffusion-based vision-language model, and adds region prompting, RoI-aligned feature replay, and structured attention masking to generate captions for multiple image regions simultaneously in a…
Yueyi Sun, Yuhao Wang, Jason Li, Ye Tian, et al.- Published
- Jun 2026
- Citations
- 0
- Code
- 77 stars
