<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI Pulse · Research (daily)</title><link>https://projectaipulse.com/#research</link><description>New AI papers, from arXiv, the big labs and leading scholars. One post a day.</description><language>en</language><lastBuildDate>Sun, 04 Oct 2026 20:03:36 +0000</lastBuildDate><atom:link href="https://projectaipulse.com/feeds/research.xml" rel="self" type="application/rss+xml"/><item><title>Research · Fri 2 Oct 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/limits-confidence-diffusion"&gt;Limits of Confidence in Diffusion&lt;/a&gt;&lt;br&gt;Discrete diffusion, including remasking and uniform-state samplers, generate a sequence by writing multiple token positions per step… &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/language-discrimination-multilingual-learning"&gt;Language Discrimination Improves Linguistic Learning in Multilingual Speech Models&lt;/a&gt;&lt;br&gt;Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.00395"&gt;Specificity-Aware Diffusion Steering via Variance-Reduced Sequential Monte Carlo&lt;/a&gt;&lt;br&gt;Addresses this problem by formulating specificity-aware steering as a target-design problem and deriving a target distribution from an overlap-based objective. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.00492"&gt;EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights&lt;/a&gt;&lt;br&gt;Introduces EurekaBench, a cross-domain benchmark that tests AI agents' ability to conduct long-horizon experiments and discover mechanisms that explain observations. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.00724"&gt;Reason in Style: Discovering and Controlling Style in Language Models&lt;/a&gt;&lt;br&gt;Studies whether recurring styles in model responses can be discovered without supervision and explicitly controlled. We design an algorithm that learns to separate representations of content and style from language models' outputs and validate its effectiveness on math questions in a controlled setting. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.00820"&gt;On-the-fly Weight Generation: A Hypernetwork Proof of Concept on ARC-1D&lt;/a&gt;&lt;br&gt;Shows that individual transformations can be represented by tiny specialist models, and that a hypernetwork can generate their parameters from context. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.01037"&gt;SLIM: Simplex-Lattice Interpolation Merging&lt;/a&gt;&lt;br&gt;Proposes Simplex-Lattice Interpolation Merging (SLIM), which constructs a quadratic surrogate of aggregate performance on the coefficient simplex using a classical mixture design. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.01092"&gt;Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation&lt;/a&gt;&lt;br&gt;Introduces Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.01618"&gt;Agents Are Systems, Not Models: Rethinking Agentic Evaluation&lt;/a&gt;&lt;br&gt;Studies these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.01742"&gt;World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories&lt;/a&gt;&lt;br&gt;Proposes World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.02191"&gt;The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models&lt;/a&gt;&lt;br&gt;Introduces the notion of Mathematical Primitive to probe structural mathematical understanding and propose, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.02200"&gt;VISTA: A Visual Harness for Reasoning in an Interactive World&lt;/a&gt;&lt;br&gt;Shows that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.02202"&gt;ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research&lt;/a&gt;&lt;br&gt;Builds ScholarCatalyst by having 184 lead authors of 207 recent computer science papers label which candidates did or could have advanced their project, each with a detailed rationale. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.02204"&gt;Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents&lt;/a&gt;&lt;br&gt;Presents Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.01823"&gt;Generalized Engression Models&lt;/a&gt;&lt;br&gt;Considers estimating the conditional distribution of a multivariate outcome given covariates when its coordinates may be continuous, binary, categorical, ordinal or rankings, and are conditionally dependent on one another. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Sat, 03 Oct 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-10-02</guid></item><item><title>Research · Thu 1 Oct 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.02185"&gt;Decoding Looped Transformers Better for (Almost) Free&lt;/a&gt;&lt;br&gt;Introduces LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.01509"&gt;Sharpening Tax in Post-Training&lt;/a&gt;&lt;br&gt;Analyzes the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.00964"&gt;RPTune: Learned Context Curation for LLM Catalog Search&lt;/a&gt;&lt;br&gt;Studies in-context catalog search through two complementary questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on curated contexts. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.01415"&gt;Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States&lt;/a&gt;&lt;br&gt;Introduces PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.00906"&gt;ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization&lt;/a&gt;&lt;br&gt;Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.01215"&gt;AutoGUIWorld: Image Generators as Visual World Models for GUI Agent&lt;/a&gt;&lt;br&gt;Introduces AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/harness-autonomous-ml-engineering"&gt;How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?&lt;/a&gt;&lt;br&gt;Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/rltl-dr-self-improvement"&gt;RLTL;DR: Self-Improvement by Internalizing Self-Generated Feedback&lt;/a&gt;&lt;br&gt;The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38329"&gt;ExploreNet: Learning Where to Explore in Diffusion GRPO&lt;/a&gt;&lt;br&gt;Introduces EXPLORENET to learn an adaptive exploration distribution. EXPLORENET is a policy that predicts a noise scale for every latent element from the current latent, the denoising step, and the prompt, before any reward is observed; it is trained on the reward spread of each rollout group and discarded after training, leaving inference unchanged. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38597"&gt;PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation&lt;/a&gt;&lt;br&gt;Presents PixelUMM, an encoder-free model for unified image and video understanding and generation directly in pixel space. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38632"&gt;Proper Scoring Rule-based Diffusion for Probabilistic Weather Forecasting&lt;/a&gt;&lt;br&gt;Introduces auxiliary conditional denoising tasks that predict the same future state from the context and its corrupted version, which provides partial future information that can reduce prediction ambiguity. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38721"&gt;UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement&lt;/a&gt;&lt;br&gt;Introduces UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38764"&gt;Lasting Effects of Abstract Pretraining Beyond Perplexity&lt;/a&gt;&lt;br&gt;Shows that in small language models, such a warm-up improves specific capabilities that are not reflected in language-modeling perplexity. Our warm-up uses an abstract stack-manipulation task that requires compositional and state-tracking capabilities. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38797"&gt;Evaluating Persistent Calibration under Evolving Model Knowledge&lt;/a&gt;&lt;br&gt;Introduces the problem of persistent calibration, which requires a confidence estimator to faithfully reflect the knowledge contained in a model as that knowledge changes, without recurring supervision. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.39247"&gt;Trust the Critic More&lt;/a&gt;&lt;br&gt;Introduces Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory to completion. AC2 instead assigns credit to action chunks: short continuations of prefixes of past trajectories. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.39504"&gt;PartiCam: Camera Controlled Video Generation with Reward Guidance&lt;/a&gt;&lt;br&gt;Presents PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffusion models. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.39560"&gt;Self-Repulsive Sampling for Diffusion Language Models&lt;/a&gt;&lt;br&gt;Introduces Self-Repulsion (SR), a sampler for masked diffusion language models that uses peer commitments to diversify the pool. At each penalized denoising step, each path lowers a token's logit according to how many peers have committed that token at the same position. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.39632"&gt;Towards Better Exploration in Sequential Test-Time Scaling&lt;/a&gt;&lt;br&gt;Shows that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.39737"&gt;A library for differentiable signal processing and machine learning on the sphere&lt;/a&gt;&lt;br&gt;Presents torch-harmonics, a comprehensive library that offers efficient, differentiable implementations of advanced signal processing and machine learning (ML) methods for spherical data. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.39866"&gt;Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing&lt;/a&gt;&lt;br&gt;Proposes a gradient-sealing principle that blocks these pathways by pushing relevant pre-activations into the negative region, where ReLU-family activations exhibit zero or near-zero derivatives. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40090"&gt;PTNO: Training Neural Operators with Noisy Monte Carlo Estimates for Particle Transport Problems&lt;/a&gt;&lt;br&gt;Proposes the Particle Transport Neural Operator (PTNO), a neural operator that learns particle transport surrogates directly from noisy, low-cost MC labels. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40127"&gt;Learning Functional Subspaces for Neural Network Compression&lt;/a&gt;&lt;br&gt;Introduces Learnable Subspace Projections (LSP), which instead learns the subspaces to discard end-to-end. Each linear layer, or tied group of layers that read the same activations, is assigned an orthogonal projector. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40134"&gt;Tactile Curiosity Drives Robot Interaction&lt;/a&gt;&lt;br&gt;Argues that tactile feedback provides a natural signal for exploration, and introduce TacEx, a framework that incorporates touch into epistemic uncertainty-driven exploration by decomposing model uncertainty across sensory modalities and directing curiosity toward the tactile channel. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40284"&gt;cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents&lt;/a&gt;&lt;br&gt;Proposes cua-speedrun, which introduces standardized infrastructure and task sets, with a focus on evaluating the speed and efficiency of CUAs. cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40305"&gt;Looped Diffusion Transformer&lt;/a&gt;&lt;br&gt;Explores an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth while keeping the parameter count fixed. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40325"&gt;WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents&lt;/a&gt;&lt;br&gt;Introduces WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Fri, 02 Oct 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-10-01</guid></item><item><title>Research · Wed 30 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.39687"&gt;Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation&lt;/a&gt;&lt;br&gt;Finds that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2610.00574"&gt;Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL&lt;/a&gt;&lt;br&gt;Studies this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-group density: the fraction of rollout groups in which the reward provides nonzero relative advantages. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38792"&gt;Training LLM Judges from Language Feedback via Position-Selective Self-Distillation&lt;/a&gt;&lt;br&gt;Studies training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40195"&gt;MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories&lt;/a&gt;&lt;br&gt;Introduces MemLife, a multimodal memory system that constructs entity-grounded, first-person text episodes and retrieves them via a time-indexed agentic reader. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.39982"&gt;Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents&lt;/a&gt;&lt;br&gt;Investigates whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40316"&gt;Scaling Laws for Looped Mixture of Experts&lt;/a&gt;&lt;br&gt;Introduces Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40285"&gt;PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents&lt;/a&gt;&lt;br&gt;Finds that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.40358"&gt;Physis-Lang: Self-Evolving Language as a Physical Representation for Video World Model&lt;/a&gt;&lt;br&gt;Revisits this assumption and introduce Physis-Lang, a self-evolving framework that treats physical language as a shared and optimizable representation across data curation, model training, and video generation. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/sclate-agent-training-evaluation"&gt;SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation&lt;/a&gt;&lt;br&gt;Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/effectiveness-fluency-llm-conditioning"&gt;On the Effectiveness-Fluency Trade-Off in LLM Conditioning: A Systematic Study&lt;/a&gt;&lt;br&gt;Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36277"&gt;From Surfaces to Volumes: Registered Geometry for Protein Representation Learning&lt;/a&gt;&lt;br&gt;Introduces Protein-TetSphere, a registered residue-wise volumetric representation for proteins. Each protein chain is tetrahedralized to obtain local volumetric regions associated with individual residues, which are then registered to a shared fixed-topology tetrahedral reference and represented in a common Laplacian basis. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36337"&gt;Adapting Linear-Time Architectures for Tabular In-Context Learning&lt;/a&gt;&lt;br&gt;Shows that the best training setup for causal models resembles next-token prediction. Then, perhaps surprisingly, the most promising linear sequence mixer is causal: DeltaNet outperforms even non-causal linear attention. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36521"&gt;PDE-OBS: Controlled Evaluation Across Observation Patterns&lt;/a&gt;&lt;br&gt;Introduces PDE-OBS, an integrated benchmarking platform spanning numerical data generation, model training, and inference and evaluation under varying observation conditions. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36527"&gt;SCOPE: Observation-Conditioned Full-Target Prediction for Sparse PDE Inference&lt;/a&gt;&lt;br&gt;Proposes SCOPE (Sparse-Context Observability-aware Predictive Embeddings) to recover complete PDE fields from sparse observations by coupling full-field latent prediction with physical reconstruction. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36720"&gt;T$^2$Mem: Learning Test-Time Memory for Robotics&lt;/a&gt;&lt;br&gt;Introduces T^2Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37171"&gt;Bridging Semantic Gaps in RAG through Generated Context Knowledge Fusion&lt;/a&gt;&lt;br&gt;Proposes Knowledge-Aware Semantic Bridging (KASB), a novel framework that improves passage selection quality through semantic space alignment between queries and retrieved documents through intelligent knowledge fusion. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37226"&gt;Follow the Entities: A Corpus Map for Agentic Search&lt;/a&gt;&lt;br&gt;Introduces CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37311"&gt;ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents&lt;/a&gt;&lt;br&gt;Proposes a novel recommendation agent framework, termed as ReMem, that combines OCR-based multimodal perception with time-evolving dynamic memory. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37374"&gt;MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding&lt;/a&gt;&lt;br&gt;Presents MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarchical reasoning, supported by a curated 25K MRG dataset with task-adaptive Chain-of-Thought (CoT) annotations that elicit multi-perspective evidence before conclusion. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37539"&gt;SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation&lt;/a&gt;&lt;br&gt;Proposes SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37588"&gt;Rational Clarification by Assistive Agents via Value-of-Information Reasoning&lt;/a&gt;&lt;br&gt;Introduces Rational Enquiry via Value-of-Information Reasoning (REVOIR). REVOIR makes clarification decisions via inference-time reasoning about the value-of-information of a question, which captures the expected improvement in task reward due to the answer received. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37631"&gt;Procedural Core: A Compact Recurrent Initialization for Vision Transformers&lt;/a&gt;&lt;br&gt;Proposes Procedural Core, an initialization strategy that captures this generic structure into a compact set of weights that can be reused across models. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37709"&gt;VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation&lt;/a&gt;&lt;br&gt;Introduces VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose)… &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37725"&gt;Context Language Models&lt;/a&gt;&lt;br&gt;Introduces Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38147"&gt;Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning&lt;/a&gt;&lt;br&gt;Introduces agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38172"&gt;Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation&lt;/a&gt;&lt;br&gt;Proposes PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36054"&gt;What if automating AI R&amp;amp;D triggers an intelligence explosion?&lt;/a&gt;&lt;br&gt;Assesses this evidence, analyze an intelligence explosion's potential impacts, and propose policy responses. AI systems are on track to automate most AI R\&amp;amp;D work within a few years, and possibly all of it. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36056"&gt;GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning&lt;/a&gt;&lt;br&gt;Presents GeoWind2Plan, a geometry-to-wind-to-planning framework for mission-time 3D urban wind prediction and energy-efficient UAV planning. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-30</guid></item><item><title>Research · Tue 29 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37200"&gt;Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL&lt;/a&gt;&lt;br&gt;Proposes Adaptive Reward Routing to jointly adapt update locations and reward coordination during forward-process RL (i.e., DiffusionNFT) of joint audio-video diffusion models. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38349"&gt;MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution&lt;/a&gt;&lt;br&gt;Introduces MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38660"&gt;Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation&lt;/a&gt;&lt;br&gt;Proposes SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38142"&gt;AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation&lt;/a&gt;&lt;br&gt;A small trainable advisor can steer a frozen language-model executor using natural-language advice. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38445"&gt;AIM: Agentic Idea Management for Automated Research&lt;/a&gt;&lt;br&gt;Introduces the Agentic Idea Manager (AIM), a fully autonomous framework for managing and exploring research directions in idea-driven automated research. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38360"&gt;On the Off-Policy Teacher in On-Policy Distillation&lt;/a&gt;&lt;br&gt;Finds that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2608.19669"&gt;Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning&lt;/a&gt;&lt;br&gt;Identifies two key limitations of this framework, one in each stage. First, the SFT stage typically relies on an off-the-shelf vision encoder to encode the helper image, yielding suboptimal latent representations that may not be well aligned with the downstream reasoning task. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37236"&gt;Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents&lt;/a&gt;&lt;br&gt;Studies a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38140"&gt;Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE&lt;/a&gt;&lt;br&gt;Shows that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38157"&gt;EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation&lt;/a&gt;&lt;br&gt;Studies vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37989"&gt;TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models&lt;/a&gt;&lt;br&gt;Introduces TabFM-Auto, which pairs a tabular foundation model, TabFM, with a language model agent that evolves the data pipeline around it. Guided by dataset metadata and validation feedback, TabFM-Auto iteratively refines data cleaning, feature engineering, context selection, and post-processing to reduce TabFM's error. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.37959"&gt;TabFM: A Zero-Shot Foundation Model for Tabular Data&lt;/a&gt;&lt;br&gt;Presents TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38059"&gt;WorldLine: Action-Driven Visual Simulation for Robotic Manipulation&lt;/a&gt;&lt;br&gt;Introduces WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38170"&gt;Adversarial Training for Pixel Diffusion&lt;/a&gt;&lt;br&gt;Shows that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38154"&gt;LongLive-Plug: Once-for-All Distillation for Video Generation&lt;/a&gt;&lt;br&gt;Introduces LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38155"&gt;Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies&lt;/a&gt;&lt;br&gt;Introduces Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36864"&gt;Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards&lt;/a&gt;&lt;br&gt;Introduces Hindsight-Divergence Localization (HDL), which uses hindsight-induced changes in token log-likelihoods to select branch points. HDL generates a small number of complete root trajectories and fills each training group with continuations from the selected positions under the original task context. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36601"&gt;SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation&lt;/a&gt;&lt;br&gt;Introduces SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.38079"&gt;OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?&lt;/a&gt;&lt;br&gt;Studies controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34965"&gt;Cyclostationary Phase Conditioning for Medical Time Series Diffusion&lt;/a&gt;&lt;br&gt;Introduces a training-free cyclostationarity index that quantifies phase structure and predicts when phase conditioning will help. Finally, we propose antithetic coupling of reverse trajectories to reduce sampling variance while achieving comparable performance with fivefold fewer network evaluations. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/communication-bottleneck-serialization"&gt;The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models&lt;/a&gt;&lt;br&gt;When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.31784"&gt;Witeness Overlap: Directional Provenance Inside Open-Weight Model Families&lt;/a&gt;&lt;br&gt;Proposes Witness Overlap, a prompt-free, training-free white-box test for directional provenance. On 176 LLM checkpoints from 16 families, our one-witness test orients 95.3\% of parent-child decisions using Frobenius cosine. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.32013"&gt;TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception&lt;/a&gt;&lt;br&gt;Presents TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to prior work, TriO utilizes three distinct sensor modalities (camera, LiDAR, and RADAR) as both inputs and sources of self-supervision, eliminating the need for additional human annotations. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.32189"&gt;ScopeIF: Improving Scope-Aware Precise Instruction-Following in Large Language Models via Graded Reward Modeling&lt;/a&gt;&lt;br&gt;Proposes ScopeIF, a novel training framework for scope-aware precise instruction-following. We first introduce a unified schema that factorizes objective constraints into three decoupled dimensions: Scope, Target, and Range. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.32750"&gt;CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning&lt;/a&gt;&lt;br&gt;Introduces CUA-Sandbox, which separates private state capsules from shared runtimes through state-scoped execution and transactional lifecycle operations, including resets and branches, while retaining the original software interfaces and task evaluators. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.33441"&gt;MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation&lt;/a&gt;&lt;br&gt;Introduces MIC (Multimodal Inconsistency Checking), an AFC framework that assists human fact-checkers by detecting AI-generated multimodal misinformation and explaining inconsistencies using world knowledge. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34085"&gt;AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving&lt;/a&gt;&lt;br&gt;Finds that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34218"&gt;Loop Dropout: Regularizing Shared Updates in Looped Language Models&lt;/a&gt;&lt;br&gt;Introduces Loop Dropout, which couples stochastic masking of adapter applications with inverse-survival rescaling to preserve expected update strength and promote effective adaptation across loops. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34272"&gt;Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training&lt;/a&gt;&lt;br&gt;Introduces GProj (gauge projection), which restores the zero sum after the cast with two rank-one corrections per row. It cuts the remaining median query/key gradient errors from 219%/13% to 0.34%/0.37%, on par with FP32 attention, for 4.7% more time per training step. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34276"&gt;NavHarness: Towards Lifelong Embodied Navigation&lt;/a&gt;&lt;br&gt;Presents NavHarness, a training-free embodied harness towards lifelong navigation that makes memory processing part of the navigation loop. During navigation, its multi-round agentic session draws on maps, task records, and house knowledge, checking them against observations and recording corrections to guide its actions. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34327"&gt;Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge&lt;/a&gt;&lt;br&gt;Finds that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34423"&gt;On the Relation Between Interval Regret and Dynamic Regret&lt;/a&gt;&lt;br&gt;Establishes a negative result that refutes this intuition of a metric-level implication. Specifically, for both convex and curved functions (including exp-concave and strongly convex functions), we show that there exist instances in which an algorithm with optimal interval regret nevertheless fails to achieve optimal dynamic regret. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34426"&gt;Q-learning Penalized Transformer for Safe Offline Reinforcement Learning&lt;/a&gt;&lt;br&gt;Addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34467"&gt;Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning&lt;/a&gt;&lt;br&gt;Presents Alignment-Guided Flow Transformer (AGFT), a novel framework that explicitly enforces tri-modal alignment through a dedicated alignment loss, bridging the representational gap across modalities and enhancing task adaptation. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34587"&gt;Reinforcement Learning from Intermediate Renders for Image-to-Code Generation&lt;/a&gt;&lt;br&gt;Introduces IR4RL, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34826"&gt;WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning&lt;/a&gt;&lt;br&gt;Investigates whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35025"&gt;AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?&lt;/a&gt;&lt;br&gt;Recent gains in language model capability have come more from data than from architecture. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35318"&gt;DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library&lt;/a&gt;&lt;br&gt;Introduces DexAgent, an agentic Human2Sim2Robot framework that converts a single egocentric human video and a task prompt into physically grounded robot trajectories for policy training. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Wed, 30 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-29</guid></item><item><title>Research · Mon 28 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34621"&gt;Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?&lt;/a&gt;&lt;br&gt;Proposes Tex-Zero, demonstrating that a high-fidelity native 3D texture generation framework can be trained without 3D assets. Our key observation is that only high-quality and fine-grained color information is essential for 3D texture training, while the required geometric information is less critical and can be manually constructed rather than obtained from real 3D assets. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35646"&gt;Rubric Rewards from Item Response Theory&lt;/a&gt;&lt;br&gt;Many language tasks have no single answer that can be checked automatically. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34563"&gt;Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence&lt;/a&gt;&lt;br&gt;Proposes ReaLVR, which brings visual-evidence supervision to the model's own free-running latent trajectories. ReaLVR contrasts correct and model-generated wrong answers to determine where stronger supervision is needed, and relevant and mismatched visual evidence to specify what to preserve. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36199"&gt;PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents&lt;/a&gt;&lt;br&gt;Introduces PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36322"&gt;Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression&lt;/a&gt;&lt;br&gt;Analyzes idealized retrieval models, showing how gradient flow dynamics may favor sharp phase specialization. Evaluating models with chunked KV-cache compression thus requires measuring across compression phases: high average accuracy can coexist with systematic positional failures. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34605"&gt;PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation&lt;/a&gt;&lt;br&gt;Finds that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35530"&gt;AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation&lt;/a&gt;&lt;br&gt;Proposes AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36352"&gt;StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks&lt;/a&gt;&lt;br&gt;Proposes StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36071"&gt;LongCat-DeepResearch Technical Report&lt;/a&gt;&lt;br&gt;Presents LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36139"&gt;Language Models Are "Insecure" Reporters&lt;/a&gt;&lt;br&gt;Introduces a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.36380"&gt;LEGO-Anything: Coding Agents for 3D Scene Reconstruction&lt;/a&gt;&lt;br&gt;Presents LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34378"&gt;Marathoner: Ultra-Long-Horizon Autonomous Intelligence&lt;/a&gt;&lt;br&gt;Proposes Marathoner, an autonomous agentic model possessing the ability of ultra-long-horizon execution. Specifically, we propose a comprehensive post-training pipeline to instill this critical capability into base model. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34606"&gt;WorldAttention: An Efficient Attention Architecture for Interactive Video World Models&lt;/a&gt;&lt;br&gt;Proposes WorldAttention, a system-oriented attention architecture that achieves high efficiency through the co-design of specialized attention kernels and hierarchical KV cache management. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35706"&gt;Reinforcing Agentic Creativity in Scientific Ideation with Night Science&lt;/a&gt;&lt;br&gt;Introduces AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35578"&gt;FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models&lt;/a&gt;&lt;br&gt;Proposes FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35673"&gt;FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching&lt;/a&gt;&lt;br&gt;Presents a novel approach to tool-based image editing by framing the task as a flow matching problem. We introduce FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35505"&gt;An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning&lt;/a&gt;&lt;br&gt;Studies on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.35457"&gt;How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining&lt;/a&gt;&lt;br&gt;Compares scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34754"&gt;Draft-KV: Learning Useful Latent Communication Between Language Models&lt;/a&gt;&lt;br&gt;Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.34972"&gt;Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models&lt;/a&gt;&lt;br&gt;Finds that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/federated-variational-inequalities"&gt;Faster Rates for Federated Variational Inequalities&lt;/a&gt;&lt;br&gt;In this paper, we study federated optimization for solving stochastic variational inequalities (VIs), a problem that has attracted growing attention in recent years. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.31562"&gt;Agentic Economies for Autonomous Scientific Discovery&lt;/a&gt;&lt;br&gt;Outlines an infrastructure for AI resource management by developing the necessary foundations of scientific agent economies, markets, and institutions. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.30935"&gt;Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models&lt;/a&gt;&lt;br&gt;Proposes EoupCT, a novel framework designed to Estimate and Orthogonalize Unknown Pre-training gradients for Continual LLM fine-Tuning. Specifically, EoupCT estimates pre-training gradients by dynamically generating pseudo data that is most susceptible to forgetting for new tasks through a learnable soft prompt equipped with Gumbel-Softmax relaxation. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.31128"&gt;Frame the adversary: a structure-aware attack methodology&lt;/a&gt;&lt;br&gt;Proposes a methodology for crafting principled frequency-based adversarial attacks, via a dedicated optimization framework. A cornerstone of our method hinges on the introduction of a perturbation constraint set, tied to highly structured non-orthogonal transforms, well-known for their flexible, non-predefined frequency handling. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.31318"&gt;AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents&lt;/a&gt;&lt;br&gt;Studies authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined attacker interface and be confirmed by an external verifier. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.31571"&gt;Strategically Diverse Sampling for Self-Training&lt;/a&gt;&lt;br&gt;Investigates strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-28</guid></item><item><title>Research · Sun 27 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.33589"&gt;TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs&lt;/a&gt;&lt;br&gt;Proposes Temperature-Grouped Reinforcement Learning (TGRL), which turns temperature-induced diversity into an explicit training signal. For each prompt, TGRL partitions its rollout group into low- and high-temperature subsets, estimates exploration gain through their reward contrast… &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.33780"&gt;Selecting Diverse SFT Traces Improves Post-RL Generalization&lt;/a&gt;&lt;br&gt;Presents a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.33848"&gt;QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents&lt;/a&gt;&lt;br&gt;Presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.33419"&gt;TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining&lt;/a&gt;&lt;br&gt;Addresses this with a matched 4 6 = 24 architecture-objective study at roughly 170M 190M encoder scale on 1.7M OpenVid and Moments-in-Time v2 clips for 8 epochs, and propose TT-VidT. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.33781"&gt;Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning&lt;/a&gt;&lt;br&gt;Introduces Entropic Advantage Policy Optimization (EAPO), an entropy-guided credit assignment method that treats success and failure asymmetrically. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.33295"&gt;TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces&lt;/a&gt;&lt;br&gt;Presents TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Mon, 28 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-27</guid></item><item><title>Research · Sat 26 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.32577"&gt;Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL&lt;/a&gt;&lt;br&gt;Introduces GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.32540"&gt;In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion&lt;/a&gt;&lt;br&gt;Introduces FlashForward, which directly reuses this cache to avoid the heavy cache-update-only model forwards. After the current chunk completes one denoising stage, its stage-specific cache is already available for the next chunk. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.32472"&gt;AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research&lt;/a&gt;&lt;br&gt;Proposes AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Sun, 27 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-26</guid></item><item><title>Research · Fri 25 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.30840"&gt;Aligning One-Step Generative Models with Reward-Weighted Transport Distillation&lt;/a&gt;&lt;br&gt;Introduces Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.31291"&gt;Softmax Reparameterization for Output-Head Quantization&lt;/a&gt;&lt;br&gt;Proposes softmax reparameterization, a post-training method that selects a functionally equivalent output head before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL separately for RTN, activation-weighted MSE, and full-Hessian GPTQ. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.31093"&gt;Block Sparse Attention with Log-Linear Complexity&lt;/a&gt;&lt;br&gt;Proposes PISA, a block-sparse attention mechanism that employs a pyramid Top-K selection strategy. The main idea is to gradually narrow down the candidates across different levels, making it more efficient to find the most relevant keys. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Sat, 26 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-25</guid></item><item><title>Research · Thu 24 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.29050"&gt;SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL&lt;/a&gt;&lt;br&gt;Proposes SLCA-GRPO, a framework incorporating Segment-Locked Credit Assignment (SLCA). To enable scalable exploration without costly real APIs and stable training, we first construct the Schema-Guided LLM Simulator (SGLS) as foundational training infrastructure. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.29265v1"&gt;A Stochastic Optimization Approach to Control-Affine Optimal Control Problems&lt;/a&gt;&lt;br&gt;Considers control-affine optimal control problems on the torus, where the dynamics and cost functions are only accessed through samples. Starting from a weak formulation of such problems, we derive a dual, a primal, and a primal-dual formulation, compatible with stochastic optimization. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.29421"&gt;Rufus-Air: An Open LLM Post-Training Recipe&lt;/a&gt;&lt;br&gt;Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.30199"&gt;ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds&lt;/a&gt;&lt;br&gt;Introduces ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.29892"&gt;Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents&lt;/a&gt;&lt;br&gt;Explores this question by building Qwen-Planner-Agent within a closed-loop AI-for-AI framework for scalable development and iterative improvement. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/latent-space-distillation"&gt;Compressing Streaming Neural Audio Encoders via Latent-Space Distillation&lt;/a&gt;&lt;br&gt;System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/practical-recipe-federated-asr"&gt;A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization&lt;/a&gt;&lt;br&gt;Semi-supervised federated learning (SSFL) trains models on clients’ unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.28921v1"&gt;PFArena: Benchmarking Language Models for Protein Modification&lt;/a&gt;&lt;br&gt;Introduces PFArena, a benchmark comprising four controlled task interfaces that cover single-mutant generation and multi-mutant ranking. By providing varying levels of mutation fitness data, PFArena reflects four representative research scenarios characterized by differing degrees of prior experimental context. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.30121v1"&gt;What, When, and How: Audio Description as Constrained Global Optimization&lt;/a&gt;&lt;br&gt;Formalizes AD generation as a constrained optimization problem over these three decisions. Our hybrid system uses large language models to propose and ground visual elements, estimate their salience to the narrative, and generate compressed realizations. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.30226v1"&gt;PoEM: Predicting RL Outcomes from Existing Policies&lt;/a&gt;&lt;br&gt;Shows that if the new reward function can be written as a linear combination of existing ones, then the new policy in log-space can be written as a linear combination of the existing log-policies. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-24</guid></item><item><title>Research · Wed 23 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.28730v1"&gt;AI-Accelerated Gyrokinetic Predictions of Turbulent Transport for Stellarator Design Optimization and Experimental Planning&lt;/a&gt;&lt;br&gt;Demonstrates the use of the AI-based turbulence surrogate in the optimization of stellarator magnetic equilibrium and to speed up stellarator transport solvers. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.27334"&gt;Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents&lt;/a&gt;&lt;br&gt;Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.27284"&gt;Hunyuan-A13B Technical Report&lt;/a&gt;&lt;br&gt;Presents Hunyuan-A13B, an open-source large language model based on a Mixture-of-Experts architecture. It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.27321"&gt;Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms&lt;/a&gt;&lt;br&gt;Compares written-out problems with stateful versions that reveal or hide their parameters. The comparison shows that most of the learnable gap lies in stateful interaction rather than underlying problem solving. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://machinelearning.apple.com/research/guide-language-flow"&gt;How to Guide Your Language Flow&lt;/a&gt;&lt;br&gt;We introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. &lt;i&gt;(Apple Machine Learning Research)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.28654v1"&gt;Training Object Permanence in World Models&lt;/a&gt;&lt;br&gt;Introduces WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.28660v1"&gt;Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy&lt;/a&gt;&lt;br&gt;Presents Morphometric Imitation, a three-stage framework that transforms reconstructed HOIs into zero-shot sim-to-real visuomotor policies. First, morphometric optimization (MMO) kinematically retargets human motion across hand morphologies while preserving demonstrated contacts. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Thu, 24 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-23</guid></item><item><title>Research · Tue 22 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.26781"&gt;Agensh: Scaling Organizational Intelligence to 1,024 Agents&lt;/a&gt;&lt;br&gt;Introduces Agensh, a scalable self-organized multi-agent harness without a central orchestrator: concurrent workers execute a multi-agent cooperation loop, continuously gathering context, claiming and self-assigning sub-tasks, taking action and sharing findings, verifying results, and merging progress in an asynchronous manner. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.25804"&gt;The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks&lt;/a&gt;&lt;br&gt;Builds Taste-Bench, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.26774"&gt;StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training&lt;/a&gt;&lt;br&gt;Argues that the root cause lies in the entanglement of the Encoder--Decoder and Codebook training: because neither module can reliably fulfill its own responsibility in isolation, the system can only function when the two subsystems happen to cooperate---a fragile condition that breaks down precisely when training is most stressed. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.27033v1"&gt;WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps&lt;/a&gt;&lt;br&gt;Introduces an optimal transport regularizer built directly from the pre-trained drift. Unlike KL reward tilting, the resulting objective transports individual samples toward higher reward rather than reweighting the base distribution. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Wed, 23 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-22</guid></item><item><title>Research · Mon 21 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24972"&gt;RRSI: Regularized Recursive Self-Improvement of Agent Harnesses&lt;/a&gt;&lt;br&gt;Introduces Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24984"&gt;WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory&lt;/a&gt;&lt;br&gt;Presents WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24308"&gt;HappyWorld-Bench&lt;/a&gt;&lt;br&gt;Introduces HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24981"&gt;GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation&lt;/a&gt;&lt;br&gt;Presents a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.25270"&gt;RULER: Instance-aware Rubric Rewards for SVG Generation&lt;/a&gt;&lt;br&gt;Addresses both limitations with rubric-based scoring. We first establish empirically that prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics, both across samples and within instructions. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24432"&gt;1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation&lt;/a&gt;&lt;br&gt;Studies this estimation problem at a fixed prefix in information geometry and propose an information-efficiency ratio (IER) based on a signal-to-noise decomposition. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24983"&gt;onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction&lt;/a&gt;&lt;br&gt;Presents onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and either picks a substitute from the model's candidate tokens or types the correct text via free-form editing. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.25001"&gt;GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay&lt;/a&gt;&lt;br&gt;Introduces GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24199v1"&gt;Vimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations&lt;/a&gt;&lt;br&gt;Introduces Vimarsha, a 100-hour benchmark spanning all 22 scheduled Indian languages, designed to address both distortions. Vimarsha combines demographically diverse on-field recordings with carefully mined in-the-wild audio selected for acoustic difficulty, alongside a lattice of variations framework that encodes multiple valid transcriptions per utterance. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24569v2"&gt;Poisson Exchange Beyond Submodularity: Effective Approximation Algorithms for Offline and Online Subset Selection over Matroids&lt;/a&gt;&lt;br&gt;Proposes a novel algorithm called, which repeatedly performs maximum-gain local exchanges through careful control of a non-homogeneous Poisson clock, and proves that this \ can attain an approximation ratio arbitrarily close to. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24352v1"&gt;Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs&lt;/a&gt;&lt;br&gt;Shows that extending this to few-shot settings, where each demonstration is generated from a different world with either the same or different graph topologies, enhances its prediction on 6 models from 4 model families. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24797v1"&gt;Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention&lt;/a&gt;&lt;br&gt;Shows that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24576v1"&gt;What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior&lt;/a&gt;&lt;br&gt;Studies the interpretability and steerability of VLN models. We use intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24745v1"&gt;Beyond Visual Quality: A Study of Test-Time Planning with World Action Models&lt;/a&gt;&lt;br&gt;Examines this planning potential empirically. First, we estimate an oracle upper bound on selection by choosing the sampled candidate whose realised outcome is best. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24976v1"&gt;DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation&lt;/a&gt;&lt;br&gt;Presents DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24487v1"&gt;AgentSTAR: Agentic Shape Tracking and Reconstruction from Monocular Videos&lt;/a&gt;&lt;br&gt;Presents a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.24967v1"&gt;Emergent Collusion in Long-Horizon LLM Agent Interaction&lt;/a&gt;&lt;br&gt;Studies the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.23980v1"&gt;MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes&lt;/a&gt;&lt;br&gt;Introduces a framework for evaluating vulnerability reports via probes, executable checks of security properties. A reported exploit is evaluated by replaying it against the application and running the probes: a triggered probe indicates both that the exploit succeeded and which security property it violated. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-21</guid></item><item><title>Research · Sun 20 Sep 2026</title><link>https://projectaipulse.com/#research</link><description>&lt;ul&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.23377"&gt;One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents&lt;/a&gt;&lt;br&gt;Develops a category-aware expert-training and policy-integration framework. Executable task construction and SWE Labeler, an evidence-grounded multi-axis labeling system, organize the training pools. &lt;i&gt;(Hugging Face Daily Papers)&lt;/i&gt;&lt;/li&gt;&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.23881v1"&gt;MotionJEPA: Preventing Temporal Feature Collapse by Capturing Visual Changes in Latent Space&lt;/a&gt;&lt;br&gt;Introduces Difference Image and Single image embedding Regularization (DISReg), a novel regularizer that builds on an inverse-dynamics-style module that predicts temporal difference image embeddings without any pixel reconstruction loss, encouraging balanced static and dynamic feature learning. &lt;i&gt;(arXiv)&lt;/i&gt;&lt;/li&gt;&lt;/ul&gt;</description><pubDate>Mon, 21 Sep 2026 00:00:00 +0000</pubDate><guid isPermaLink="false">aipulse-research-2026-09-20</guid></item></channel></rss>
