AI papers & tools · read and explained

Reduce costs.
Boost quality.
Get inspired.

Nowness collects AI research papers and developer tools and explains each one in plain terms — the problem it tackles and what you could use it for.

Latest finds

What the lab found.

Each one is read from its source and summarised: what it is, the problem it tackles, and what you could use it for.

Paper2026-09-30

Latent Inference-Time Guidance of Time Series Foundation Models

This paper introduces a method to adaptively combine forecasts from multiple Time Series Foundation Models (TSFMs) using a time-dependent latent space.

ProblemTime Series Foundation Models are sensitive to user-selected parameters (lookback, covariates, horizon), leading to variable and inconsistent forecast quality that requires manual tuning or selection of the best context.

Use it forImproving the accuracy of time series forecasting by ensembling multiple foundation models; Mitigating the sensitivity of TSFM performance to user-selected lookback and covariates; Creating robust forecasting pipelines that leverage complementary strengths of different TSFMs

time-seriesfoundation-modelsensemblingforecastinglatent-space
arxiv.org ↗
Paper2026-09-30

Traversing the solution space of neural networks with Hessian Null Space Continuation

This paper introduces Hessian Null Space Continuation (HNC), a method that uses local curvature to traverse weight space regions that preserve network function.

ProblemStandard gradient-based optimization fails to reveal the full diversity of internal mechanisms and representations that exist within low-loss regions of weight space, limiting mechanistic understanding and model manipula

Use it forMechanistic interpretability of neural network solution spaces; Model merging and editing by navigating between functionally equivalent solutions; Identifying reward hacking in reinforcement learning agents

neural-networksoptimizationinterpretabilityhessianweight-space
arxiv.org ↗
Paper2026-09-30

ReCIRC: Rectified Conformal Risk Control

ReCIRC is a method that improves conformal risk control by inverting estimated local risk curves to create a common target conditional risk budget.

ProblemStandard conformal risk control uses a single threshold for all inputs, which overprotects easy cases and underprotects hard ones due to varying conditional risk.

Use it forMedical image segmentation with controlled missed lesion rates; Multilabel classification with controlled missed label rates; Multiclass classification with controlled error rates

conformal predictionrisk controlmachine learningstatistical guaranteescalibration
arxiv.org ↗
Paper2026-09-30

Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution

This paper investigates the gap between the planning modes LLM agents declare and how they actually execute them.

ProblemExisting planner-executor systems often fail because generic agents (like Plan+ReAct) do not faithfully preserve the declared planning structure during execution, and final success metrics cannot distinguish between poor

Use it forImproving the reliability of LLM agents on long-horizon tasks like software engineering (SWE-bench) and household simula; Designing agent architectures that enforce specific planning structures to prevent structural drift; Evaluating the effectiveness of different planning strategies (Search vs. Hierarchical) across different environments

LLM AgentsPlanningExecutionRoutingSWE-bench
arxiv.org ↗
Paper2026-09-30

SelfSearch: Reward-Free Search for Self-Improving Agents

SelfSearch is a reward-free search procedure that allows LLM agents to modify their own instructions and tools using records of previous self-improvement episodes.

ProblemExisting self-improvement methods for LLM agents rely on repeated downstream evaluations, which are costly and tie the search process to specific evaluated tasks.

Use it forImproving agent performance on coding benchmarks like Terminal-Bench and SWE-bench; Reducing execution costs for complex software engineering tasks; Creating self-optimizing agent harnesses for autonomous coding tasks

LLM agentsself-improvementsearch algorithmsautonomous agentscoding agents
arxiv.org ↗
Paper2026-09-30

HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL

HorizonFlow is a hierarchical planner for offline goal-conditioned reinforcement learning that treats plan length as a generated output rather than a fixed input.

ProblemExisting generative planning methods for offline RL require specifying the planning horizon before generating plan content, which can lead to infeasible transitions if too short or redundant motion if too long.

Use it forOffline goal-conditioned reinforcement learning in navigation tasks; Visual manipulation tasks with variable task horizons; Trajectory inpainting where the appropriate planning horizon depends on the specific route

reinforcement learningoffline RLplanningflow matchinggoal-conditioned
arxiv.org ↗
Paper2026-09-30

KUPAS MASTER: Distilling Tacit Expertise into Agent-Ready Corpora

KUPAS MASTER is an experience engineering platform that converts heterogeneous work records and practitioner interviews into structured, traceable experience corpora for LLM agents.

ProblemRoutine work records and standard RAG systems fail to capture the tacit knowledge of experts, such as which cues matter, why a judgment is reasonable, and specific action boundaries, limiting the effectiveness of LLM age

Use it forConverting unstructured expert interviews into reusable agent skills; Building domain-specific knowledge bases for professional LLM agents; Organizing organizational knowledge from individual practitioner records

LLM agentsknowledge distillationtacit knowledgeexperience engineeringRAG
arxiv.org ↗
Paper2026-09-30

SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation

SAKI is a method for on-policy distillation that uses maximal coupling to route token-level supervision based on whether the student's sampled token matches the teacher's highest-probability token.

ProblemWeak students in on-policy distillation often visit teacher-misaligned prefixes where standard supervision is less representative, leading to suboptimal learning and train-test state mismatch.

Use it forTraining smaller language models to match larger teacher models on mathematical reasoning tasks; Improving the efficiency of on-policy distillation pipelines by reducing train-test state mismatch; Implementing adaptive supervision signals that adjust based on student-teacher alignment

distillationon-policy-learninglanguage-modelsmathematical-reasoningmachine-learning
arxiv.org ↗
Paper2026-09-30

Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

This paper introduces AdaLCPI, an adaptive long-context prompt injection attack that splits malicious instructions into fragments distributed across retrieved content.

ProblemExisting prompt injection defenses and safety evaluations typically assume malicious instructions are contiguous or complete, failing to account for attacks where the objective is reconstructed from distributed fragments

Use it forEvaluating the robustness of agentic systems against distributed prompt injection attacks; Benchmarking agent safety mechanisms against adaptive, long-context threats; Researching the limits of LLM reasoning when reconstructing instructions from incomplete data

prompt-injectionagent-securityllm-safetyred-teaminglong-context
arxiv.org ↗
Browse finds →