01
Logics-STEM: Empowering LLM Reasoning via Failure-Driven Post-Training and Document Knowledge Enhancement
We introduce Logics-STEM, a state-of-the-art reasoning model fine-tuned on Logics-STEM-SFT-Dataset, a 10M-scale, high-quality open-source long chain-of-thought corpus for STEM reasoning. Logics-STEM achieves strong performance across STEM benchmarks, outperforming the next-best 8B model by 4.68% on average. These gains stem from a data–algorithm co-design framework that jointly optimizes data curation and post-training. The dataset is built via a five-stage pipeline—annotation, deduplication, decontamination, distillation, and stratified sampling—while the training framework adopts a failure-driven post-training strategy with targeted knowledge retrieval and data synthesis to refine SFT and RL. We release Logics-STEM models (8B, 32B) and datasets (10M, 2.2M) to support open-source research.
Large Language ModelReinforcement LearningData Processing