From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
Abstract
Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.
Community
let's discuss!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- ForkPilot: Self-Evolving Policy for Retrospective Search in Long-Horizon Agents (2026)
- Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories (2026)
- MOAE: Multi-Objective Agent Evolution with Pareto-Preserving Search (2026)
- JET: Judge-Guided Evolution at Test Time for Agent Programs (2026)
- HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents (2026)
- LatentSift: Policy-State Filtering for Token-Efficient Verification of Software Engineering Agents (2026)
- How Much Can Language Models Gain from Test-Time Computation? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.09684 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper