Paper 2511.15552
Multimodal Evaluation of Russian-language Architectures
- Published
- Nov 2025
- Research lab
- Independent
- Citations
- 3
- GitHub
- Not linked
01 In brief
Summary
The paper introduces MERA Multi, the first open multimodal evaluation benchmark for Russian-language architectures, addressing the lack of such benchmarks for Slavic languages.
It comprises 18 instruction-based tasks across text, image, audio, and video modalities, built on a unified taxonomy of multimodal abilities.
The benchmark includes 11 private datasets created from scratch with Russian cultural and linguistic specificity, and 7 public datasets.
Evaluation uses block-prompting, dual metrics (Exact Match and JudgeScore via a trained LLM-as-a-judge), and a scoring system combining Attempted Score and Coverage.
Data protection includes watermarking and a Multimodal Semantic Membership Inference Attack (MSMIA) for leakage detection.
Baselines for over 50 open-source and proprietary models (e.g., GPT-4.1) are provided.
Results show Qwen3-Omni-30B-A3B-Instruct leads with a Total Score of 0.5, while GPT-4.1 excels in image tasks but has low coverage.
The benchmark offers a reproducible methodology for culturally aware multimodal evaluation in non-English languages.
02 From the paper
Abstract
Multimodal large language models (MLLMs) are currently at the center of research attention, showing rapid progress in scale and capabilities, yet their intelligence, limitations, and risks remain insufficiently understood. To address these issues, particularly in the context of the Russian language, where no multimodal benchmarks currently exist, we introduce MERA Multi, an open multimodal evaluation framework for Russian-spoken architectures. The benchmark is instruction-based and encompasses default text, image, audio, and video modalities, comprising 18 newly constructed evaluation tasks for both general-purpose models and modality-specific architectures (imageto-text, video-to-text, and audio-to-text). Our contributions include: (i) a universal taxonomy of multimodal abilities; (ii) 18 datasets created entirely from scratch with attention to Russian cultural and linguistic specificity, unified prompts, and metrics; (iii) baseline results for both closed-source and open-source models; (iv) a methodology for preventing benchmark leakage, including watermarking for private sets. While our current focus is on Russian, the proposed benchmark provides a replicable methodology for constructing multimodal benchmarks in typologically diverse languages, particularly within the Slavic language family.