arxiv:2510.14919

Predicting Task Performance with Context-aware Scaling Laws

Published on Oct 16

· Submitted by

Kyle Montgomery on Oct 17

Upvote

Authors:

Kyle Montgomery ,

Abstract

A framework models downstream performance of large language models as a function of training compute and context, offering insights into efficient design for long-context tasks.

AI-generated summary

Scaling laws have transformed our understanding of large language models by linking upstream metrics like cross-entropy loss to design factors such as model size, training data, and compute. However, these conventional laws fail to capture downstream task performance, where context plays a critical role. In this work, we propose a straightforward, interpretable framework that jointly models downstream performance as a function of the training compute and the provided context. We empirically validate our framework by fitting it on the observed downstream performance of extended-context variants of Llama-2-7B and Llama-2-13B across 65,500 unique instances spanning three tasks: arithmetic reasoning, common sense reasoning, and machine translation. Our results demonstrate that our framework accurately models in-distribution downstream performance, generalizes across three orders of magnitude in training compute, and reliably extrapolates performance as the amount of context increases. These findings offer valuable insights into the interplay between training compute and context utilization, providing guidance for designing more efficient long-context LLMs for diverse downstream tasks. Our code is available at https://github.com/wang-research-lab/context-scaling.

View arXiv page View PDF GitHub 1 Add to collection

Community

kylemontgomery

Paper author Paper submitter about 19 hours ago

The paper extends traditional scaling laws by jointly modeling downstream task performance as a function of both training compute and context length (e.g., number of in-context demonstrations). Empirical evaluation on extended-context Llama-2 variants across arithmetic reasoning, common sense reasoning, and machine translation tasks shows that the model fits observed behavior and generalizes across orders of magnitude in compute and context length.

librarian-bot

about 14 hours ago

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment

Upvote

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2510.14919 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2510.14919 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2510.14919 in a Space README.md to link it from this page.