arxiv:2510.01141

Apriel-1.5-15b-Thinker

Published on Oct 1

· Submitted by

Aman Tiwari on Oct 6

#1 Paper of the day

ServiceNow-AI

Upvote

109

Authors:

Aman Tiwari ,

Akintunde Oladipo ,

Sai Rajeswar Mudumba ,

Torsten Scholak ,

Abstract

A 15-billion parameter multimodal reasoning model achieves competitive performance through a progressive training methodology without reinforcement learning, demonstrating efficient use of computational resources.

AI-generated summary

We present Apriel-1.5-15B-Thinker, a 15-billion parameter open-weights multimodal reasoning model that achieves frontier-level performance through training design rather than sheer scale. Starting from Pixtral-12B, we apply a progressive three-stage methodology: (1) depth upscaling to expand reasoning capacity without pretraining from scratch, (2) staged continual pre-training that first develops foundational text and vision understanding, then enhances visual reasoning through targeted synthetic data generation addressing spatial structure, compositional understanding, and fine-grained perception, and (3) high-quality text-only supervised fine-tuning on curated instruction-response pairs with explicit reasoning traces spanning mathematics, coding, science, and tool use. Notably, our model achieves competitive results without reinforcement learning or preference optimization, isolating the contribution of our data-centric continual pre-training approach. On the Artificial Analysis Intelligence Index, Apriel-1.5-15B-Thinker attains a score of 52, matching DeepSeek-R1-0528 despite requiring significantly fewer computational resources. Across ten image benchmarks, its performance is on average within five points of Gemini-2.5-Flash and Claude Sonnet-3.7, a key achievement for a model operating within single-GPU deployment constraints. Our results demonstrate that thoughtful mid-training 2 design can close substantial capability gaps without massive scale, making frontier-level multimodal reasoning accessible to organizations with limited infrastructure. We release the model checkpoint, all training recipes, and evaluation protocols under the MIT license to to advance open-source research.

View arXiv page View PDF Add to collection

Community

amant555

Paper author Paper submitter 14 days ago

Introducing ServiceNow’s 15B-parameter model that matches 𝗗𝗲𝗲𝗽𝗦𝗲𝗲𝗸–𝗥𝟭–𝟬𝟱𝟮𝟴, 𝗠𝗶𝘀𝘁𝗿𝗮𝗹–𝗺𝗲𝗱𝗶𝘂𝗺–𝟭.𝟮 and 𝗚𝗲𝗺𝗶𝗻𝗶 𝗙𝗹𝗮𝘀𝗵 𝟮.𝟱 on the Artificial Analysis Index (𝗔𝗔𝗜 𝟱𝟮) — delivering comparable results at a 𝗳𝗿𝗮𝗰𝘁𝗶𝗼𝗻 𝗼𝗳 𝘁𝗵𝗲 𝘀𝗶𝘇𝗲 (at least 8-10 times smaller)

𝗙𝗿𝗼𝗻𝘁𝗶𝗲𝗿-𝗹𝗲𝘃𝗲𝗹 𝗿𝗲𝗮𝘀𝗼𝗻𝗶𝗻𝗴 on a single GPU
𝗡𝗼 𝗥𝗟 𝗽𝗵𝗮𝘀𝗲 — the step-change comes from mid-training
𝗥𝗲𝗮𝘀𝗼𝗻𝘀 𝗼𝘃𝗲𝗿 𝗶𝗺𝗮𝗴𝗲𝘀 - Image + Text mid training enables model to reason over images without additional training
𝗚𝗿𝗲𝗮𝘁 𝗮𝘁 𝗿𝗲𝗮𝘀𝗼𝗻𝗶𝗻𝗴 — AIME2025: 88, GPQA: 71, LCB: 73
𝗙𝗼𝗹𝗹𝗼𝘄𝘀 𝗶𝗻𝘀𝘁𝗿𝘂𝗰𝘁𝗶𝗼𝗻𝘀 reliably — IFBench: 62
T𝗮𝘂𝟮 𝗕𝗲𝗻𝗰𝗵 (Telecom): 68 → ready for real-world workflows
𝗢𝗽𝗲𝗻 𝘄𝗲𝗶𝗴𝗵𝘁𝘀 model to further research and reproducibility (MIT license)

acjulius

13 days ago

will the data be released?

librarian-bot

13 days ago

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

WaltonFuture

13 days ago

Thanks for your work! I have a quick question: how do you organize the data formats for tasks like Image Reconstruction and Visual Matching in CPT Stage 2? I think this synthetic augmentation approach is particularly interesting.

Thank you!