Each one is read from its source and summarised: what it is, the problem it tackles, and what you could use it for.
Paper2026-09-30
This paper introduces a method to adaptively combine forecasts from multiple Time Series Foundation Models (TSFMs) using a time-dependent latent space.
ProblemTime Series Foundation Models are sensitive to user-selected parameters (lookback, covariates, horizon), leading to variable and inconsistent forecast quality that requires manual tuning or selection of the best context.
Use it forImproving the accuracy of time series forecasting by ensembling multiple foundation models; Mitigating the sensitivity of TSFM performance to user-selected lookback and covariates; Creating robust forecasting pipelines that leverage complementary strengths of different TSFMs
time-seriesfoundation-modelsensemblingforecastinglatent-space
arxiv.org ↗
Paper2026-09-30
This paper introduces Hessian Null Space Continuation (HNC), a method that uses local curvature to traverse weight space regions that preserve network function.
ProblemStandard gradient-based optimization fails to reveal the full diversity of internal mechanisms and representations that exist within low-loss regions of weight space, limiting mechanistic understanding and model manipula
Use it forMechanistic interpretability of neural network solution spaces; Model merging and editing by navigating between functionally equivalent solutions; Identifying reward hacking in reinforcement learning agents
neural-networksoptimizationinterpretabilityhessianweight-space
arxiv.org ↗
Paper2026-09-30
ReCIRC is a method that improves conformal risk control by inverting estimated local risk curves to create a common target conditional risk budget.
ProblemStandard conformal risk control uses a single threshold for all inputs, which overprotects easy cases and underprotects hard ones due to varying conditional risk.
Use it forMedical image segmentation with controlled missed lesion rates; Multilabel classification with controlled missed label rates; Multiclass classification with controlled error rates
conformal predictionrisk controlmachine learningstatistical guaranteescalibration
arxiv.org ↗
Paper2026-09-30
This paper investigates the gap between the planning modes LLM agents declare and how they actually execute them.
ProblemExisting planner-executor systems often fail because generic agents (like Plan+ReAct) do not faithfully preserve the declared planning structure during execution, and final success metrics cannot distinguish between poor
Use it forImproving the reliability of LLM agents on long-horizon tasks like software engineering (SWE-bench) and household simula; Designing agent architectures that enforce specific planning structures to prevent structural drift; Evaluating the effectiveness of different planning strategies (Search vs. Hierarchical) across different environments
LLM AgentsPlanningExecutionRoutingSWE-bench
arxiv.org ↗
Paper2026-09-30
SelfSearch is a reward-free search procedure that allows LLM agents to modify their own instructions and tools using records of previous self-improvement episodes.
ProblemExisting self-improvement methods for LLM agents rely on repeated downstream evaluations, which are costly and tie the search process to specific evaluated tasks.
Use it forImproving agent performance on coding benchmarks like Terminal-Bench and SWE-bench; Reducing execution costs for complex software engineering tasks; Creating self-optimizing agent harnesses for autonomous coding tasks
LLM agentsself-improvementsearch algorithmsautonomous agentscoding agents
arxiv.org ↗
Paper2026-09-30
HorizonFlow is a hierarchical planner for offline goal-conditioned reinforcement learning that treats plan length as a generated output rather than a fixed input.
ProblemExisting generative planning methods for offline RL require specifying the planning horizon before generating plan content, which can lead to infeasible transitions if too short or redundant motion if too long.
Use it forOffline goal-conditioned reinforcement learning in navigation tasks; Visual manipulation tasks with variable task horizons; Trajectory inpainting where the appropriate planning horizon depends on the specific route
reinforcement learningoffline RLplanningflow matchinggoal-conditioned
arxiv.org ↗
Paper2026-09-30
KUPAS MASTER is an experience engineering platform that converts heterogeneous work records and practitioner interviews into structured, traceable experience corpora for LLM agents.
ProblemRoutine work records and standard RAG systems fail to capture the tacit knowledge of experts, such as which cues matter, why a judgment is reasonable, and specific action boundaries, limiting the effectiveness of LLM age
Use it forConverting unstructured expert interviews into reusable agent skills; Building domain-specific knowledge bases for professional LLM agents; Organizing organizational knowledge from individual practitioner records
LLM agentsknowledge distillationtacit knowledgeexperience engineeringRAG
arxiv.org ↗
Paper2026-09-30
SAKI is a method for on-policy distillation that uses maximal coupling to route token-level supervision based on whether the student's sampled token matches the teacher's highest-probability token.
ProblemWeak students in on-policy distillation often visit teacher-misaligned prefixes where standard supervision is less representative, leading to suboptimal learning and train-test state mismatch.
Use it forTraining smaller language models to match larger teacher models on mathematical reasoning tasks; Improving the efficiency of on-policy distillation pipelines by reducing train-test state mismatch; Implementing adaptive supervision signals that adjust based on student-teacher alignment
distillationon-policy-learninglanguage-modelsmathematical-reasoningmachine-learning
arxiv.org ↗
Paper2026-09-30
This paper introduces AdaLCPI, an adaptive long-context prompt injection attack that splits malicious instructions into fragments distributed across retrieved content.
ProblemExisting prompt injection defenses and safety evaluations typically assume malicious instructions are contiguous or complete, failing to account for attacks where the objective is reconstructed from distributed fragments
Use it forEvaluating the robustness of agentic systems against distributed prompt injection attacks; Benchmarking agent safety mechanisms against adaptive, long-context threats; Researching the limits of LLM reasoning when reconstructing instructions from incomplete data
prompt-injectionagent-securityllm-safetyred-teaminglong-context
arxiv.org ↗