Skip to main content
arXiv is now an independent nonprofit! Learn more
archive
Search Submit Donate Log in

Computer Science > Computation and Language

(cs)
[Submitted on 3 May 2025 (v1), last revised 14 Nov 2025 (this version, v3)]

Title:New News: System-2 Fine-tuning for Robust Integration of New Knowledge

Authors:Core Francisco Park, Zechen Zhang, Hidenori Tanaka
View a PDF of the paper titled $\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge, by Core Francisco Park and 2 other authors
View PDF HTML (experimental)
Abstract:Humans and intelligent animals can internalize new information and accurately internalize their implications to perform downstream tasks. While large language models (LLMs) can achieve this through in-context learning (ICL) when the information (news) is explicitly given as context, adequately integrating the information into model weights via fine-tuning remains challenging. In this paper, we introduce New News, a dataset composed of hypothetical yet plausible news spanning multiple domains (mathematics, coding, discoveries, leaderboards, events), accompanied by downstream evaluation questions whose correct answers critically depend on understanding and internalizing the news. First, we demonstrate a substantial gap between naive fine-tuning and in-context learning (FT-ICL gap) on our dataset. To address this gap, we explore a suite of self-play data generation protocols -- paraphrases, implications, and Self-QA -- designed to distill the knowledge processed by the model with context into the weights of the model, which we term System-2 Fine-tuning (Sys2-FT). We systematically evaluate ICL and Sys2-FT performance across data domains and model scales with the Qwen 2.5 family of models. Our results demonstrate that the Self-QA protocol of Sys2-FT significantly improves models' in-weight learning of the news while preserving general capabilities. Furthermore, we discover the contextual shadowing effect, where training with the news in context followed by its rephrases or QAs catastrophically degrades learning of the news. Finally, we show preliminary evidence of an emerging scaling law of Sys2-FT.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2505.01812 [cs.CL]
  (or arXiv:2505.01812v3 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2505.01812
arXiv-issued DOI via DataCite

Submission history

From: Core Francisco Park [view email]
[v1] Sat, 3 May 2025 12:49:35 UTC (22,471 KB)
[v2] Sat, 27 Sep 2025 21:44:18 UTC (18,558 KB)
[v3] Fri, 14 Nov 2025 15:58:22 UTC (18,550 KB)
Full-text links:

Access Paper:

    View a PDF of the paper titled $\textit{New News}$: System-2 Fine-tuning for Robust Integration of New Knowledge, by Core Francisco Park and 2 other authors
  • View PDF
  • HTML (experimental)
  • TeX Source
license icon view license

Current browse context:

cs.CL
< prev   |   next >
new | recent | 2025-05
Change to browse by:
cs
cs.AI
cs.LG

References & Citations

  • NASA ADS
  • Google Scholar
  • Semantic Scholar
Loading...

Bookmark

BibSonomy Reddit

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)
We gratefully acknowledge support from our major funders, member institutions, , and all contributors.
About · Help · Contact · Subscribe · Copyright · Privacy · Accessibility · Operational Status (opens in new tab)
Major funding support from
Simons Foundation Simons Foundation International Schmidt Sciences