The year/Independent research

Paper 2607.20145

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Published
Jul 2026
Research lab
Independent
Citations
0
GitHub
32 stars

01 In brief

Summary

This technical report presents SLAI T-Rex, a full-stack framework for post-training the DeepSeek-V4 model family on Ascend SuperPOD.

System-level optimizations (parallelism, communication, memory, kernels) increased Model FLOPs Utilization (MFU) from 11.67% to 34.22%, a 2.93x improvement.

For Operations Research (OR) specialization, a solver-grounded Continued Pre-Training (CPT) and Supervised Fine-Tuning (SFT) pipeline was developed.

On DeepSeek-V4-Flash, the final model achieved an average OR score of 71.81%, outperforming GPT-5.4-Mini by 3.98 points and the base model by 11.27 points.

Applying the workflow to DeepSeek-V4-Pro improved its average OR score from 70.16% to 77.33%.

Key findings include: CPT initialization provides additional gains over SFT alone, especially on structural equivalence (B4O-ORGEval); data cleaning and chain-of-thought enhancement are more effective than simply scaling synthetic data; and a balanced data mixture is essential to avoid catastrophic forgetting.

02 From the paper

Abstract

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.