Agent Research Papers

Automatically Updated on 2026.09.11

Current Search Keywords: Agent,Multi-Agent,Tool Learning,Agent RL,Autonomous Agent,LLM Agent

If you have any other keywords, please feel free to let us know :)

Web Page (Scrape Code)

Agent

Publish Date Title Authors PDF Code
2026-09-09 IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications Yiling Ma et.al. 2609.10539 null
2026-09-09 Avatar: Toward Autonomous End-to-End Orchestration of Scientific Workflows using LLMs Suman Raj et.al. 2609.10509 null
2026-09-09 Glyph: A Multi-Strategy Agentic System for Column Description and Sensitivity-Ontology Tagging of Enterprise Data Catalogs Kostia Kudriavtsev et.al. 2609.10430 null
2026-09-09 TrajMark: Ownership Attribution and Segment-Level Tamper Localization for Coding-Agent Trajectories Bokang Zeng et.al. 2609.10416 null
2026-09-09 An Empirical Analysis of ReDoS Vulnerabilities and ReDoS Detection Tools N’Zolieh Ismaël Mahassadi et.al. 2609.10294 null
2026-09-09 A-JIT: Agentic Just-In-Time Software Construction Mark Marron et.al. 2609.10248 null
2026-09-09 Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries? Tianzhu Zhang et.al. 2609.10181 null
2026-09-09 RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases Yingqian Wu et.al. 2609.10092 null
2026-09-09 Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training Junwon Ko et.al. 2609.10052 null
2026-09-09 Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability Arnab Chattopadhayay et.al. 2609.10036 null
2026-09-09 Optimal Value Inference for Reinforcement Learning Nan Lu et.al. 2609.09981 null
2026-09-09 SpeechAnnotator: A Context-Aware Multi-Agent Framework and Benchmark for Multidimensional Speech Annotation Qirui Zhan et.al. 2609.09947 null
2026-09-09 Strangers to Themselves: What Language Models Say About Themselves Is Generic Phil Blandfort et.al. 2609.09899 null
2026-09-09 Can We Trust Video Hallucination Detectors? VidHalLoc for Evaluating the Evaluators Xinyu Chen et.al. 2609.09895 null
2026-09-09 AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents Shrey Nag et.al. 2609.09875 null
2026-09-09 With a Thermomix You Lose the Ability to Cook: A Kitchen Machine Analogy for Applications of Generative AI in Education Nikol Rummel et.al. 2609.09856 null
2026-09-09 The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents Benjamin Gruenbaum et.al. 2609.09853 null
2026-09-09 Can AI Agents Detect and Repair Artifact Drift in Network Experiments? Tianzhu Zhang et.al. 2609.09849 null
2026-09-09 InstantMimic: A High Performance System for Learning Physics-based Skills in Seconds Ikjun Choi et.al. 2609.09821 null
2026-09-09 Pairit: A Platform for Live Experiments on Human-AI Collaboration Harang Ju et.al. 2609.09789 null
2026-09-08 ReCite: Agentic Reasoning for Faithful Citation Yuyang Huang et.al. 2609.09156 null
2026-09-08 Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Yuxing Lu et.al. 2609.09153 null
2026-09-08 Copying explains the collective behavior of AI agents in the wild Giordano De Marzo et.al. 2609.09150 null
2026-09-08 ExecCritic: Learn to Test, Test to Improve for Coding Agents Leitian Tao et.al. 2609.09133 null
2026-09-08 MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents Boyu Yang et.al. 2609.09115 null
2026-09-08 SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? Yuqiao Tan et.al. 2609.09113 null
2026-09-08 PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation Yixuan Liu et.al. 2609.09087 null
2026-09-08 ToolLoop: Closed-Loop Tool-Use Data Synthesis via Decomposed Generation and Dynamic Self-Feedback Min Zeng et.al. 2609.09072 null
2026-09-08 Time-Varying Data as Sheaves: an Invitation to Narratives Wilmer Leal et.al. 2609.09056 null
2026-09-08 PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving Yuan Gao et.al. 2609.08965 null
2026-09-08 Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents Wenbo Gao et.al. 2609.08919 null
2026-09-08 Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course Evelyn Duesterwald et.al. 2609.08832 null
2026-09-08 A Controlled Comparison of Manual and Teleoperated Intraocular Instrument Motion for an Input Device Korab Hoxha et.al. 2609.08770 null
2026-09-08 MorphoOrgaAgent: A Foundation-Model-Based Multi-Agent System for Autonomous Organoid Analysis Hanyi Zhang et.al. 2609.08696 null
2026-09-08 Graph-Based Personalized Memory for LLM Agents: Representation, Evolution, Retrieval, and Evaluation Dac Duy Anh Nguyen et.al. 2609.08599 null
2026-09-08 A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation Rahul Khedar et.al. 2609.08592 null
2026-09-08 The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution? Boyang Wang et.al. 2609.08589 null
2026-09-08 BIO-MEMART: Biometric-Aware KV Cache Memory for Multi-User LLM Agents Yanhong Qian et.al. 2609.08566 null
2026-09-08 PLC-Bin2Src: Retrieving Corresponding Structured Text Source Files for PLC Binaries Ang Jia et.al. 2609.08563 null
2026-09-08 Personalizing LLM Agent Memory Using Biometrics Yanhong Qian et.al. 2609.08558 null
2026-09-04 Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe Dain Kim et.al. 2609.05395 null
2026-09-04 Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence Urja Pawar et.al. 2609.05385 null
2026-09-04 Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool Samuel Kushnir et.al. 2609.05364 null
2026-09-04 RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? Zhenxuan Fan et.al. 2609.05324 null
2026-09-04 Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness Alexander Neubauer et.al. 2609.05314 null
2026-09-04 Testing Interchangeability in LLM Agent Teams Jianxin Gao et.al. 2609.05279 null
2026-09-04 How to Speculate about Uncertainty in Agentic Coding? A Draft-Model Gate Method Konstantin Grotov et.al. 2609.05274 null
2026-09-04 CONTINUITY: Security-Context Contracts for Composable LLM Agent Controls Chris Zheng et.al. 2609.05269 null
2026-09-04 Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents Jiazheng Sun et.al. 2609.05261 null
2026-09-04 Governing Bring Your Own AI: A Parameterized Maturity Model Dare Bello et.al. 2609.05236 null
2026-09-04 Substrate-Aware AI Agents: Execution Context as a First-Class Input Manu Agrawal et.al. 2609.05232 null
2026-09-04 First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves Tianjie Ju et.al. 2609.05224 null
2026-09-04 LIBERO-RECOVER: Beyond Task Success Towards Failure Recovery in Robotic Manipulation Models Lin Liu et.al. 2609.05178 null
2026-09-04 TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents Zhibo Yang et.al. 2609.05079 null
2026-09-04 A Structured Debate-Mixture-of-Agents Framework for Complex Clinical Diagnostic Decision Support Chang Xia et.al. 2609.05069 null
2026-09-04 Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment Arno Libert et.al. 2609.05036 null
2026-09-04 Artificial Intelligence in Equity and Crypto Markets: Progress, Profitability Evidence, and the Limits of Automated Investing Linsen Zhu et.al. 2609.04917 null
2026-09-04 Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing Jiahe Geng et.al. 2609.04915 null
2026-09-04 ARIA - An Agentic Framework for Autonomous Testing of Infotainment Systems António Azevedo et.al. 2609.04913 null
2026-09-04 RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents Aziz Ben Amor et.al. 2609.04898 null
2026-09-03 A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms Davide Paglieri et.al. 2609.04170 null
2026-09-03 SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents Xin He et.al. 2609.04167 null
2026-09-03 SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center Uday Vallabhaneni et.al. 2609.04159 null
2026-09-03 Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Jie Wu et.al. 2609.04148 null
2026-09-03 Efficient Test-Time Adaptation through Human-AI Interaction Zora Zhiruo Wang et.al. 2609.04141 null
2026-09-03 The Natural Language Interaction Protocol and Standard for AI Agents Luyi Xing et.al. 2609.04135 null
2026-09-03 PatchBench: Evaluating AI Agents for Vulnerability Patching Chihao Shen et.al. 2609.04075 null
2026-09-03 Extending concurrent separation logic to the hardware level to verify the xv6 OS kernel on RISC-V with AI agents M. Frans Kaashoek et.al. 2609.04043 null
2026-09-03 Editable Visual Design Junyan Ye et.al. 2609.04034 null
2026-09-03 A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits Arslan Brömme et.al. 2609.04017 null
2026-09-03 Hierarchical automation of scanning probe microscopy through agentic orchestration and algorithmic control Boris N. Slautin et.al. 2609.04015 null
2026-09-03 Unlocking Lossless Speedups in LLMs via Discrete Diffusion Subham Sekhar Sahoo et.al. 2609.04010 null
2026-09-03 Shifting from Injection to Interaction: Rethinking Web Security in the Age of LLMs and Beyond Nivedita Singh et.al. 2609.03999 null
2026-09-03 FiMI Banking: A Sovereign Model for Indian Retail Banking NPCI AI Research Team et.al. 2609.03960 null
2026-09-03 Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting Muneeb Khan et.al. 2609.03923 null
2026-09-03 Value-Preserving Architectures for Agentic AI Systems Alessandro Pesare et.al. 2609.03920 null
2026-09-03 A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors Pengxun Li et.al. 2609.03884 null
2026-09-03 Bioinfoysis Technical Report Qingyang Shao et.al. 2609.03871 null
2026-09-03 Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations Lei Zheng et.al. 2609.03860 null
2026-09-03 Semantic Bayesian World Models Tommaso Soru et.al. 2609.03834 null
2026-09-02 Discriminative World Models for Web Agents Kelvin Li et.al. 2609.02885 null
2026-09-02 EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Yuling Shi et.al. 2609.02783 null
2026-09-02 Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Jianlyu Chen et.al. 2609.02749 null
2026-09-02 BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research Wooyoung Jung et.al. 2609.02729 null
2026-09-02 ACLE-MCP: Attested Capability Leases for Execution-Time Trust in Remote LLM Tool Use Zhiyang Ding et.al. 2609.02690 null
2026-09-02 HINT: Human-Intent Inception for Long-Horizon Robot Manipulation Mingyu Mei et.al. 2609.02653 null
2026-09-02 Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting Ron Begleiter et.al. 2609.02649 null
2026-09-02 PrimSynth: An Agentic Approach to Discover, Validate, and Synthesize Exploit Primitives for Linux Kernel Vulnerabilities Pengfei Wang et.al. 2609.02647 null
2026-09-02 Competitive Market Behavior of LLMs Pawel Struski et.al. 2609.02580 null
2026-09-02 A Finger on the Scale: Covert Policy Steering through Agentic Skills Jiarui Li et.al. 2609.02564 null
2026-09-02 Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment Chenyu Zhou et.al. 2609.02417 null
2026-09-02 Semantics-Guided Automatic Tensorization for Multiobjective Evolutionary Algorithms: A Multi-Agent Framework Zhenyu Liang et.al. 2609.02387 null
2026-09-02 Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions Jiayi Bi et.al. 2609.02371 null
2026-09-02 LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory Kun-Yang Yu et.al. 2609.02350 null
2026-09-02 SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology Ihor Stepanov et.al. 2609.02292 null
2026-09-02 PaperCompiler: Faithful Paper-to-Code Generation via Repository-Level Specification Compilation Yunhao Liu et.al. 2609.02272 null
2026-09-02 CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents S M Asif Hossain et.al. 2609.02265 null
2026-09-02 PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment Fan Yuxuan et.al. 2609.02231 null
2026-09-02 SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams Ao Yan et.al. 2609.02217 null
2026-09-02 Agentic Settlement Protocol: An Application Profile for Refundable, Delayed-Fulfilment Agent Commerce on Stablecoin Rails Behnam et.al. 2609.02208 null
2026-09-01 Mechanism Design for Alignment and Control Dirk Bergemann et.al. 2609.01595 null
2026-09-01 Designing Proactive Thought Partners for Writing Chao Zhang et.al. 2609.01588 null
2026-09-01 From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Olga Tsymboi et.al. 2609.01572 null
2026-09-01 Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories Nabira Rashid et.al. 2609.01556 null
2026-09-01 EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation Qing Zhao et.al. 2609.01526 null
2026-09-01 When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation Peiying Zhu et.al. 2609.01519 null
2026-09-01 GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions Elias Stengel-Eskin et.al. 2609.01491 null
2026-09-01 Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement Haoyang Yan et.al. 2609.01481 null
2026-09-01 Freemium Model for Information Provision Igal Milchtaich et.al. 2609.01468 null
2026-09-01 TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution Ruocan Wei et.al. 2609.01428 null
2026-09-01 Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching Jaewoo Park et.al. 2609.01404 null
2026-09-01 Autonomous robotic bridging using distributed swarm control without inter-agent communication Vishwaak C. Thamaraiselvan et.al. 2609.01394 null
2026-09-01 EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems Jun Hou et.al. 2609.01360 null
2026-09-01 LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting Yufei Chen et.al. 2609.01337 null
2026-09-01 Agentic Multimodal Models for Environmental Hyperspectral Unmixing Michał Cholewa et.al. 2609.01289 null
2026-09-01 Analog-DB: An Agent-First Analog Integrated Circuit Database, From Blocks to Systems Danial Noori Zadeh et.al. 2609.01286 null
2026-09-01 Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents Liming Pu et.al. 2609.01245 null
2026-09-01 Continuous Autonomous Refactoring: A Research Roadmap for AI-Driven Code Quality Maintenance Xin Sun et.al. 2609.01236 null
2026-09-01 What’s in Your Agent’s Context? Context Privilege Escalation Attacks against AI Agent Harness Zichuan Li et.al. 2609.01222 null
2026-09-01 Autonomous discovery of new structure-plausibility laws for explainable and rapid crystal diagnosis and screening Zhilong Song et.al. 2609.01209 null
2026-08-31 Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data Milad Rezaei Hajidehi et.al. 2608.31082 null
2026-08-31 Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents Le Chen et.al. 2608.31057 null
2026-08-31 Agentic Quantitative Trading: A Survey of Workflows, Systems, and Evaluation Fengrui Hua et.al. 2608.31041 null
2026-08-31 MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents Vernon Toh et.al. 2608.31022 null
2026-08-31 From Prompt to Prototype: Towards a Frontier LLM Driven RF Engineering Workflow Markus Heinrichs et.al. 2608.31006 null
2026-08-31 A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting Xiaoyu Tao et.al. 2608.30976 null
2026-08-31 The Hermon Moment: AI Self-Transcendence and Its Human Narration Alexei Grinbaum et.al. 2608.30971 null
2026-08-31 One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning Armin Dariani et.al. 2608.30952 null
2026-08-31 Detecting AI Impostors: How Do Middle Schoolers Identify LLM Agents in a Live Collaborative Setting? Dan Schumacher et.al. 2608.30948 null
2026-08-31 Evidence, Logic, and Compliance: Multi-Agent Structured Graph Reasoning with Expert Arbitration for Medical Referral Qi Peng et.al. 2608.30938 null
2026-08-31 Beacon: LLM Multi-Agent Driven Hardware Design Space Exploration for Heterogeneous Multi-Chiplet Deep Learning Accelerators Boyu Li et.al. 2608.30932 null
2026-08-31 TRIPPULSE: Multi-Agent Travel Planning with Review-Grounded Reasoning Priyanshu Karmakar et.al. 2608.30924 null
2026-08-31 Adaptive KV Retention for LLM Agents at Human-Approval Timescales Minseo Choi et.al. 2608.30830 null
2026-08-31 PRACTICE: From Experience to Expertise in Self-Evolving Embodied Agents Ziyi Bai et.al. 2608.30760 null
2026-08-31 E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation Wei Fan et.al. 2608.30730 null
2026-08-31 BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks Pradyumna Shyama Prasad et.al. 2608.30724 null
2026-08-31 VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs Jingyi He et.al. 2608.30705 null
2026-08-31 TUE-Detector: A Tool-Using Expert MLLM-Based Detector for AI-Generated Videos Yichen Wu et.al. 2608.30704 null
2026-08-31 A Phased Workflow for Operating LLM-Based Coding Agents Ante Kapetanovic et.al. 2608.30701 null
2026-08-31 Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning Fukang Zhu et.al. 2608.30686 null
2026-08-31 LLM-based Hardware Development with Hierarchical IRs and End-to-End Multi-Agent Workflow Chenyang Yin et.al. 2608.30659 null
2026-08-31 Geometry of Divergence: Tracking Hidden-State Trajectories for Adaptive Multi-Turn Reasoning Jie Liang et.al. 2608.30650 null
2026-08-31 Practical Implementation Report on Introducing Spec-Driven Development Using AI Agents in Software Development PBL Hidetake Tanaka et.al. 2608.30572 null
2026-08-31 Authority-Inference Separation in Agentic Finance: First-Line Control, Blockchain Enforcement, and Replayable Assurance Hui Gong et.al. 2608.30519 null
2026-08-31 CM2: Multimodal Cultural Reasoning via an Integrated Multi-Agent Framework Qi Li et.al. 2608.30498 null
2026-08-31 Agents in the Large: Perception-Centered Architecture for Persistent Agents Shihan Dou et.al. 2608.30478 null
2026-08-31 SeqAlign3DVG: A Sequence-Aligned Benchmark and Voxel Reasoning Framework for 3D Visual Grounding Yi Zhang et.al. 2608.30451 null
2026-08-31 ECLIPSE: Self-Evolving Stealthy Prompt Injection Attack against Long-Horizon Agentic Systems Shiqian Zhao et.al. 2608.30441 null
2026-08-31 Learning to Reason and Use Tools through Unsupervised Fine-Tuning in Task-Oriented Dialog Systems Markel Ferro et.al. 2608.30426 null
2026-08-31 Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents Yunseok Lee et.al. 2608.30362 null
2026-08-31 Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents Hanlin Tian et.al. 2608.30322 null
2026-08-31 Update from Hell: Can Coding Agents Survive Hidden Breakage in Dependency Upgrades? Zijian Luo et.al. 2608.30300 null
2026-08-31 Extracting Knowledge from Tools in LLM Agents Chuanchao Zang et.al. 2608.30288 null
2026-08-31 FABO: Agent-Guided Discovery of Joint Breakpoint Optimization for Timing-Driven Routing Trees Shang Liu et.al. 2608.30268 null
2026-08-31 Motus2: A Self-Evolving General World Model for Dexterous Manipulation Hongzhe Bi et.al. 2608.30237 null
2026-08-31 SIR: Self-improving Red-teaming for Compute Use Agents Chen Xiong et.al. 2608.30207 null
2026-08-31 FaVOR: LLM-Based Agentic Framework for Factor Mining via Empirical Validation Hyeonjin Kim et.al. 2608.30192 null
2026-08-31 GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space Boqi Chen et.al. 2608.30188 null
2026-08-31 Understanding Stage-Wise Utility-Risk Trade-offs in LLM Agent Memory Chuanchao Zang et.al. 2608.30177 null
2026-08-31 Science sandboxes measure the scientific capability of AI agents Arya S. Rao et.al. 2608.30165 null
2026-08-28 Offline-Verifiable Accountability for Cross-Organization Agent Messaging: A Preserved Evidence-Bundle Approach Adil Alshammari et.al. 2608.28542 null
2026-08-28 Recognition Without Enforcement: Configuration-Dependent Failures in LLM Agent Instruction Arbitration and External Control Jun Wen Leong et.al. 2608.28502 null
2026-08-28 On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code Plugin Marketplaces Ahmed Hereiz et.al. 2608.28497 null
2026-08-28 ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL Zhuoshi Pan et.al. 2608.28476 null
2026-08-28 Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning Minghui Xu et.al. 2608.28447 null
2026-08-28 Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction Qing Ye et.al. 2608.28439 null
2026-08-28 LUCID: An Agentic AI Framework on Digital-Twin in the Loop for QoS-Guaranteeing Robotic Control Hyeonsu Lyu et.al. 2608.28437 null
2026-08-28 Prove2Me: An Open Collaborative Platform for Scaling Math Formalization Shuze Chen et.al. 2608.28433 null
2026-08-28 When Verified Source Becomes Attack Input: Defending Smart Contracts Against LLM-Based Vulnerability Scanning Mingyuan Huang et.al. 2608.28400 null
2026-08-28 RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents Yupeng Zhang et.al. 2608.28399 null
2026-08-28 PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems Hanglong Lv et.al. 2608.28378 null
2026-08-28 EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses Tanmay Sah et.al. 2608.28363 null
2026-08-28 AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents Pengze Li et.al. 2608.28345 null
2026-08-28 LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Yi Wang et.al. 2608.28281 null
2026-08-28 STEGNav: Spatio-Temporal Event Graph Reasoning for Multimodal Lifelong Object Navigation Yang Chen et.al. 2608.28279 null
2026-08-28 Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation Linze Wu et.al. 2608.28276 null
2026-08-28 CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning Yang Chen et.al. 2608.28266 null
2026-08-28 Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration Xiaoqing Wang et.al. 2608.28264 null
2026-08-28 Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation Tianle Wang et.al. 2608.28241 null
2026-08-28 Adaptive Strategy Generation for Boundary Value Exploration Beyond Numeric Inputs Sabinakhon Akbarova et.al. 2608.28230 null
2026-08-27 WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution Liyan Tang et.al. 2608.27454 null
2026-08-27 Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach? Ting Yan et.al. 2608.27443 null
2026-08-27 RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution Junjie Zhang et.al. 2608.27439 null
2026-08-27 Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit Yisen Xi et.al. 2608.27427 null
2026-08-27 Embodied Scene Rearrangement Planning Canzhi Chen et.al. 2608.27371 null
2026-08-27 INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment Yutong Zhang et.al. 2608.27348 null
2026-08-27 When Context Gets Root: Privilege Escalation in LLM Harnesses Xingbang He et.al. 2608.27299 null
2026-08-27 Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search Yuan Chang et.al. 2608.27266 null
2026-08-27 What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents Xingshan Zeng et.al. 2608.27260 null
2026-08-27 Assessing Company Contributions to Societal Resilience: Extending the Societal Capacity Assessment Framework to Agentic AI Catherine Simons et.al. 2608.27238 null
2026-08-27 SPA: Securing Persistent LLM Agents Across Queries with Plan-First Information-Flow Control Dylan Girrens et.al. 2608.27234 null
2026-08-27 TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution Tommaso Bendinelli et.al. 2608.27182 null
2026-08-27 Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable Pranav Aggarwal et.al. 2608.27167 null
2026-08-27 When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents Xiaokun Guo et.al. 2608.27146 null
2026-08-27 GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL Zike Yuan et.al. 2608.27142 null
2026-08-27 Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents Chenhao Wu et.al. 2608.27141 null
2026-08-27 TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation Jingyi Zheng et.al. 2608.27127 null
2026-08-27 The Framing Gap: Indirect Prompt-Injection Exfiltration Defeats Surface-Level Defenses in Tool-Using Agents Md Habibur Rahman et.al. 2608.27092 null
2026-08-27 FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets Kuan-Hao Tseng et.al. 2608.27021 null
2026-08-27 ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions Rui Xie et.al. 2608.26991 null
2026-08-26 Agentic Autoresearch for Cell-Edge Power Control: Radically Redefining the Researcher’s Role Ahmad Khan et.al. 2608.26093 null
2026-08-26 SwarmWorld: Stigmergic technological evolution in societies of language-model agents Subhadeep Pal et.al. 2608.26081 null
2026-08-26 VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following Min Zeng et.al. 2608.26013 null
2026-08-26 A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks Tongyan Hu et.al. 2608.26008 null
2026-08-26 AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs Sheng Liang et.al. 2608.26004 null
2026-08-26 ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs Songyuan Li et.al. 2608.25992 null
2026-08-26 Spatial-Knowledge-Graph-Grounded LLM Agents for Neighborhood Livability Evaluation Haiyan Hao et.al. 2608.25952 null
2026-08-26 Candidate supply and answer selection shape the value of LLM judging in multi-agent systems Jia-Hao Ji et.al. 2608.25937 null
2026-08-26 AI Agentic Selective Laser Sintering Process Optimization Peter Pak et.al. 2608.25928 null
2026-08-26 Code World Model: Coding Agent as World Brain Yiwen Chen et.al. 2608.25927 null
2026-08-26 Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence Shengyi Pan et.al. 2608.25905 null
2026-08-26 SkillShield: Prompt-Space Security Skills for LLM Coding Agents Xiaodong Wu et.al. 2608.25817 null
2026-08-26 LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI Agents Weiming Li et.al. 2608.25777 null
2026-08-26 EVOMAL: Self-Poisoning in Self-Evolving Coding Agents Xiaodong Wu et.al. 2608.25776 null
2026-08-26 Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences Natasha Ureyang et.al. 2608.25771 null
2026-08-26 HypoForge: A Self-Improving Multi-Agent Framework for Automated Hypothesis Generation and Testing via Scientific Skill Learning Ziqing Qian et.al. 2608.25770 null
2026-08-26 Reassembling Distributed Risk: Trajectory-Conditioned Action Generation for Multi-Turn Agent Safety Yanbo Dai et.al. 2608.25711 null
2026-08-26 AI Slop and Hallucinations in Vulnerability Assessment: A Survey on Reasoning Failures and Trustworthy Mitigation Junchen Ding et.al. 2608.25667 null
2026-08-26 Beyond Scaling: Self-Evolving LLM Agents for Hardware Kernel Optimization via an Experience-Driven Workflow and Experience Graph Memory Siyuan Chen et.al. 2608.25570 null
2026-08-26 AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research Xintong Zhang et.al. 2608.25559 null
2026-08-25 SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL Kai Ruan et.al. 2608.24870 null
2026-08-25 BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes Fei Tang et.al. 2608.24848 null
2026-08-25 SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents Shidong Yang et.al. 2608.24747 null
2026-08-25 Meta $^n$ : Recursive Self-Improvement through Emergent Depth Zae Myung Kim et.al. 2608.24735 null
2026-08-25 Simthesizer: An Agent-Driven Simulation Framework for LLM Serving Systems Wonung Kim et.al. 2608.24650 null
2026-08-25 A Literate Programming Environment for Human and Machine Agents Adam T. Burke et.al. 2608.24644 null
2026-08-25 IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents Bo Ren et.al. 2608.24588 null
2026-08-25 Joint Optimization of Tool Creation and Use for Large Language Model Agents Zhi Rui Tam et.al. 2608.24571 null
2026-08-25 EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents Lihang Zeng et.al. 2608.24570 null
2026-08-25 When “Must” Becomes “Maybe”: Constraint Weakening in LLM Agent Workflows Yiheng Sun et.al. 2608.24569 null
2026-08-25 StrokeGuard: A Multi-Agent Guided System for Prehospital Stroke Assessment Wentao Yang et.al. 2608.24555 null
2026-08-25 PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents Zhi-Kai Chen et.al. 2608.24509 null
2026-08-25 From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use Rongfeng Guo et.al. 2608.24368 null
2026-08-25 Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems Yarden Bakish et.al. 2608.24361 null
2026-08-25 The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents Roy Ganz et.al. 2608.24358 null
2026-08-25 ReproAgent: Contract-Guided Paper-to-Code Reproduction Xue Hu et.al. 2608.24291 null
2026-08-25 Observability and Fault Injection for LLM-Based Multi-Agent Systems in Software Engineering Zahra Seyedghorban et.al. 2608.24271 null
2026-08-25 SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction Xue Hu et.al. 2608.24252 null
2026-08-25 DeepRepoQA: Code Repository Question Answering with Deep Agent Exploration Weihan Peng et.al. 2608.24221 null
2026-08-25 Paritok-4B: Intent-Conditioned Context Compression for Coding Agents Jiayu Shi et.al. 2608.24188 null
2026-08-24 SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? Deyao Hong et.al. 2608.23564 null
2026-08-24 Prime Agent: A Self-Improving RLM Harness Seth Karten et.al. 2608.23552 null
2026-08-24 The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams Summer Eunhyung Ann et.al. 2608.23541 null
2026-08-24 EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards Zhiqing Cui et.al. 2608.23525 null
2026-08-24 An Interactive Agent for Requirement-Driven Candidate Sourcing Yuanpeng He et.al. 2608.23501 null
2026-08-24 MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters ChengAo Shen et.al. 2608.23473 null
2026-08-24 InjecMEM: Memory Injection Attack on LLM Agent Memory Systems Hanling Tian et.al. 2608.23471 null
2026-08-24 Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep Pedro Santos et.al. 2608.23395 null
2026-08-24 DPIAgent: Divide, Protocol, Isolate for Agentic Reproduction Test Generation Hao Liu et.al. 2608.23341 null
2026-08-24 Can Coding Agents Build Robust Baselines? A Skill-Based Approach for Automating the Medical Imaging Model-Development Pipeline Eugenia Moris et.al. 2608.23336 null
2026-08-24 Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents Wenqi Liu et.al. 2608.23329 null
2026-08-24 From Natural Language Policies to Executable Obligations: A Verification Harness for Dependable In-Car LLM Agents Radouane Bouchekir et.al. 2608.23282 null
2026-08-24 NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration Chang Liu et.al. 2608.23179 null
2026-08-24 Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech Generation Sujoy Nath et.al. 2608.23152 null
2026-08-24 Molecular LLM Agents: From Architectural Design to Scientific Autonomy Jiatong Li et.al. 2608.23104 null
2026-08-24 ARGUS: MCP-Grounded Root Cause Analysis for Kubernetes Incidents Ergi Senja et.al. 2608.23084 null
2026-08-24 AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models Saurav Singla et.al. 2608.23078 null
2026-08-24 Signal or Noise? A Benchmark Study of Agent Skills in Web Development Ziyue Yang et.al. 2608.23067 null
2026-08-24 From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation Xiangxin Zhang et.al. 2608.23045 null
2026-08-24 AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Sungho Park et.al. 2608.23041 null
2026-08-21 AI with Authority, from Application to Silicon Jason Hickey et.al. 2608.21356 null
2026-08-21 Asymmetric Capacity Allocation in Self-Refinement Pipelines Zhuoyi Yang et.al. 2608.21345 null
2026-08-21 Invisible Agents, Uninformed Patients: Towards Responsible Deployment Of Autonomous AI Diagnostic Agents In Sub-Saharan Africa Percy Brown et.al. 2608.21326 null
2026-08-21 AI-to-AI Code Reviews of GitHub Pull Requests Niruthiha Selvanayagam et.al. 2608.21311 null
2026-08-21 Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis Qisheng Lu et.al. 2608.21310 null
2026-08-21 Benchmarking Patent Drafting from Inventor-Style Disclosures Lekang Jiang et.al. 2608.21249 null
2026-08-21 Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models Zhuoyuan Li et.al. 2608.21247 null
2026-08-21 AID-Guard: Stateful Authorization for Delegated Agent Effects Yingzhe Tong et.al. 2608.21159 null
2026-08-21 Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Yuyuan Feng et.al. 2608.21156 null
2026-08-21 TraceGrant: A Contract-Governed Security Framework for the Task-Effect Lifecycle of Networked LLM Agents Bohao Liao et.al. 2608.21126 null
2026-08-21 ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents Kai Wang et.al. 2608.21101 null
2026-08-21 Don’t Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents Yanze Jiang et.al. 2608.21027 null
2026-08-21 Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models Tonglin Yan et.al. 2608.20975 null
2026-08-21 ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction Linhao Zhong et.al. 2608.20920 null
2026-08-21 Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique Sahil Mishra et.al. 2608.20777 null
2026-08-21 Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol Guodong Xu et.al. 2608.20729 null
2026-08-21 ArtiMo: Agent-Driven Articulated Mesh Animation Chunyu Zou et.al. 2608.20699 null
2026-08-21 VortexChat: An agentic framework for autonomous multi-objective integrated photonic design Faqian Chong et.al. 2608.20688 null
2026-08-21 The Claws in Plain Sight: Unauthorized Context Disclosure through LLM Agent Tool Calls Ben Dong et.al. 2608.20658 null
2026-08-21 ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection Chunyi Wang et.al. 2608.20637 null
2026-08-20 An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction Narges Ahmadi et.al. 2608.20320 null
2026-08-20 AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement Yizhe Chi et.al. 2608.20318 null
2026-08-20 MidTool: Mid-training Data Synthesis for Agentic Tool Use Fengqing Jiang et.al. 2608.20314 null
2026-08-20 Complete Symbols of Equivariant Pseudodifferential Operators on Noncompact Symmetric Spaces Satwata Hans et.al. 2608.20313 null
2026-08-20 Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents Yiyang Feng et.al. 2608.20274 null
2026-08-20 From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation Zhijun Gao et.al. 2608.20195 null
2026-08-20 Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Atsuyuki Miyai et.al. 2608.20169 null
2026-08-20 Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design Poomphob Suwannapichat et.al. 2608.20099 null
2026-08-20 Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees Yu Chen et.al. 2608.19993 null
2026-08-20 ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance Yiyang Luo et.al. 2608.19974 null
2026-08-20 G-MARK: Grounded Multi-Agent Reasoning for Cooperative Driving via Knowledge Graphs Bhavya Gupta et.al. 2608.19964 null
2026-08-20 Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis Zijiao Chen et.al. 2608.19902 null
2026-08-20 MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection Yue Wang et.al. 2608.19901 null
2026-08-20 EnvHarness: Awakening Static Worlds for Agent Learning Chengsong Huang et.al. 2608.19880 null
2026-08-20 A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries Mahyar Abbasian et.al. 2608.19875 null
2026-08-20 PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents Seongjae Kang et.al. 2608.19861 null
2026-08-20 Inadvertent Context Leakage in Language Models Jaiden Fairoze et.al. 2608.19857 null
2026-08-20 Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions Stephen Barrett et.al. 2608.19816 null
2026-08-20 MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents Bo Qian et.al. 2608.19803 null
2026-08-20 SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Zhipeng Xu et.al. 2608.19799 null
2026-08-19 SPADE: Self-Play in Adaptive Synthetic Executable Environments Bo Liu et.al. 2608.19197 null
2026-08-19 Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication Ramneet Kaur et.al. 2608.19161 null
2026-08-19 Quantum circuit optimization using deep reinforcement learning: Applications across multiple gate sets Khoa Dang Tao et.al. 2608.19103 null
2026-08-19 What is Missing from AI Post-Training AI: An Empirical Analysis Joy Jia Yin Lim et.al. 2608.19072 null
2026-08-19 Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering Pradeep Murugesan et.al. 2608.19029 null
2026-08-19 DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning Zijie Meng et.al. 2608.18878 null
2026-08-19 CauSec: Unboxing the Causal Drivers of Static Vulnerability Analysis Performance Md Akram Khan et.al. 2608.18876 null
2026-08-19 SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents Qingyao Li et.al. 2608.18852 null
2026-08-19 Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction Zijian Xiao et.al. 2608.18840 null
2026-08-19 A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation Manoj N M et.al. 2608.18740 null
2026-08-19 CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes Kumal Hewagamage et.al. 2608.18734 null
2026-08-19 DocClaw: A Unified Agentic System for Intelligent Document Processing Siqi Xiang et.al. 2608.18685 null
2026-08-19 RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training Yugu Li et.al. 2608.18682 null
2026-08-19 Code Health in LLM-Based Test Generation: Effectiveness and Token Efficiency Freya Wirdemann et.al. 2608.18645 null
2026-08-19 PILOT Technical Report Jiuning Lin et.al. 2608.18637 null
2026-08-19 CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence Yutong Cheng et.al. 2608.18613 null
2026-08-19 AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin Bang Xie et.al. 2608.18588 null
2026-08-19 PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMs Yaqi Li et.al. 2608.18573 null
2026-08-19 CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks Pattaraphon Kenny Wongchamcharoen et.al. 2608.18554 null
2026-08-19 Measuring Proof Burden in Public Bounty Listings: A RentAHuman Case Study Iman YeckehZaare et.al. 2608.18547 null
2026-08-18 Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating Daria Leshchikova et.al. 2608.18058 null
2026-08-18 StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents Yining Hua et.al. 2608.18050 null
2026-08-18 Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation Zhikai Xu et.al. 2608.18034 null
2026-08-18 aDSL: Agentic 3D Creation via Joint Agent-Program Design Rui-Huan Wang et.al. 2608.17975 null
2026-08-18 EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection Lei Jiang et.al. 2608.17933 null
2026-08-18 CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion Zheling Tan et.al. 2608.17911 null
2026-08-18 AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis Yangtian Liu et.al. 2608.17834 null
2026-08-18 Edge-Native Embodied Intelligence for Action-Aware Wireless Edge Networks Yiru Wang et.al. 2608.17774 null
2026-08-18 D $^2$ ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory Xule Liu et.al. 2608.17756 null
2026-08-18 Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment Zhen Zhang et.al. 2608.17713 null
2026-08-18 GADR: Gathering Architecture Decision Records from Meeting Transcriptions Lucas Daniel Costa da Silva et.al. 2608.17694 null
2026-08-18 Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch Jialong Li et.al. 2608.17684 null
2026-08-18 Benchmarking Automated Security Patch Backporting: How Far Are We? Jincheng Yang et.al. 2608.17671 null
2026-08-18 GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities Haoran Bu et.al. 2608.17665 null
2026-08-18 Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback Kang Peng et.al. 2608.17587 null
2026-08-18 HODAgent: Towards On-Demand, Responsive Humanoids for Physical World Human Interaction Wang Warren Chen et.al. 2608.17584 null
2026-08-18 GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting Qijian Tian et.al. 2608.17535 null
2026-08-18 Agent Lightning v1.0: Towards Harnessed Agentic RL Zhiyuan He et.al. 2608.17528 null
2026-08-18 Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents Ram Rachum et.al. 2608.17524 null
2026-08-18 Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context Yiwen Zhao et.al. 2608.17499 null
2026-08-17 Don’t Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory Bingxin Xu et.al. 2608.16889 null
2026-08-17 Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation Jiawei Liu et.al. 2608.16843 null
2026-08-17 Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning Minh-Ha Nguyen et.al. 2608.16831 null
2026-08-17 When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents Jiawei Liu et.al. 2608.16806 null
2026-08-17 When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding Giuseppe Destefanis et.al. 2608.16801 null
2026-08-17 Neurosymbolic Embodied Agents Mohammad Albinhassan et.al. 2608.16794 null
2026-08-17 Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis Reza Fayyazi et.al. 2608.16775 null
2026-08-17 TDD-Agent: Test-Driven Reasoning for Code Generation Hongyue Yu et.al. 2608.16742 null
2026-08-17 Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors David Eric Austin et.al. 2608.16707 null
2026-08-17 PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning Veit Laule et.al. 2608.16637 null
2026-08-17 The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks Bardia Mohammadi et.al. 2608.16630 null
2026-08-17 Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning Peng Du et.al. 2608.16620 null
2026-08-17 Characterizing Agentic Flooding of Government Services Chris Schmitz et.al. 2608.16603 null
2026-08-17 Zetta $ζ$ : An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Xin Ding et.al. 2608.16590 null
2026-08-17 Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents Batu El et.al. 2608.16578 null
2026-08-17 VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience Jianming Chen et.al. 2608.16544 null
2026-08-17 HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents Shen Liu et.al. 2608.16447 null
2026-08-17 D2-ScaleAgent: Dual-Dimensional Scaling for Long Document Understanding Hao Zhang et.al. 2608.16417 null
2026-08-17 Towards Risk-free AI Agent Deployment Yintong Huo et.al. 2608.16411 null
2026-08-17 A Policy Algebra for Trust-Preserving Agentic AI Execution Bhaskar Tripathi et.al. 2608.16402 null
2026-08-14 Validating LLM-Modernized Scientific Software Through Differential Fault Injection Evan Coleman et.al. 2608.14527 null
2026-08-14 Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers Taenyun Kim et.al. 2608.14522 null
2026-08-14 Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training Hanfeng Lu et.al. 2608.14498 null
2026-08-14 Twin: Playing an Unknown Game with a Test-Time Digital Twin Alexy Skoutnev et.al. 2608.14490 null
2026-08-14 SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning Panjing He et.al. 2608.14452 null
2026-08-14 Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports Beatrice Alessandra Motetti et.al. 2608.14446 null
2026-08-14 The Past and Future of AI Scientists Ross D. King et.al. 2608.14407 null
2026-08-14 AgentRewind: Recoverable Execution for Long-Horizon LLM Agents Yu Zhuang et.al. 2608.14380 null
2026-08-14 Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages Chih-Hsuan Yang et.al. 2608.14375 null
2026-08-14 ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond Mingming Zhao et.al. 2608.14354 null
2026-08-14 ATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata Learning Ignacio D. Lopez-Miguel et.al. 2608.14352 null
2026-08-14 Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in LLM Agents Zhizhao Guan et.al. 2608.14339 null
2026-08-14 Multi-Agent Reinforcement Learning for Joint Handover Management and Power Allocation in Multi-Orbit Satellite Networks Yassine Afif et.al. 2608.14335 null
2026-08-14 Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL Xiaojun Wu et.al. 2608.14312 null
2026-08-14 TimeSage-EV: A Live Benchmark for Agentic Time Series Analysis in Evolving Environments Qingren Yao et.al. 2608.14270 null
2026-08-14 Polaris : Multi Agentic System for Conversational Enterprise Analytics Varuni H K et.al. 2608.14246 null
2026-08-14 Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions Xiaokai Yan et.al. 2608.14132 null
2026-08-14 LegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy Workflows Thilo Reintjes et.al. 2608.14131 null
2026-08-14 A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents Ismail El Hamraoui et.al. 2608.14109 null
2026-08-14 AppLooper: An Agentic Application Engineering Loop for Accountable Release with Virtual-User Feedback Zihong He et.al. 2608.14093 null
2026-08-13 AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Yaxin Luo et.al. 2608.13560 null
2026-08-13 OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Bobo Li et.al. 2608.13558 null
2026-08-13 QuoteBench: How Matched Scores Can Hide Command-Path Failures Shangao Li et.al. 2608.13547 null
2026-08-13 Vero: Can AI Agents Build Formally Verified Software Repositories? Zhe Ye et.al. 2608.13522 null
2026-08-13 Intern-S2-Preview: Scientific Agentic Foundation Model Lei Bai et.al. 2608.13505 null
2026-08-13 MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination Saisha Shetty et.al. 2608.13476 null
2026-08-13 UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models Yukun Dai et.al. 2608.13453 null
2026-08-13 Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes Aimilios Hadjiliasi et.al. 2608.13420 null
2026-08-13 Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Yiwei Li et.al. 2608.13417 null
2026-08-13 Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks Muhammad Hannan Akram et.al. 2608.13394 null
2026-08-13 Thermal transport in crystals: from the quantum Dyson equation to mesoscopic phonon hydrodynamics Enrico Di Lucente et.al. 2608.13339 null
2026-08-13 Training AI Scientists to Replicate Research Damon Falck et.al. 2608.13331 null
2026-08-13 TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems Shunwen Bai et.al. 2608.13221 null
2026-08-13 Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Zechuan Wang et.al. 2608.13179 null
2026-08-13 SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents Chang Liu et.al. 2608.13173 null
2026-08-13 Validation-Centric AI-Assisted GPU Porting of a 250,000+ Line Legacy Weather Simulation Code Tetsuya Hoshino et.al. 2608.13122 null
2026-08-13 Formal Verification of Quantum Ancilla Safety Jiqi Li et.al. 2608.13099 null
2026-08-13 Semantic Radiance Fields as Simulators for Spatial Reasoning in Real-World Scenes Nico Heider et.al. 2608.13095 null
2026-08-13 VALG: An Agentic System for ML Theory Research Dechen Zhang et.al. 2608.13060 null
2026-08-13 Latent On-Policy Self-Distillation Guibin Zhang et.al. 2608.13040 null
2026-08-12 AVA-Encoder: Towards Agent-Native Video Representation Learning Chuyue Li et.al. 2608.12313 null
2026-08-12 The Role Specialization Model (RSM): Coordinating LLM-Based Tools in Agentic Software Development - An Exploratory Case Study Carlos Alberto Fernández-y-Fernández et.al. 2608.12311 null
2026-08-12 DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation Yan Deng et.al. 2608.12308 null
2026-08-12 Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior Yusuf Pisan et.al. 2608.12292 null
2026-08-12 VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies Ankita Rajaram Naik et.al. 2608.12282 null
2026-08-12 PACE-SIMS: Checkpoint-Gated Autonomous SIMS Characterization with AI-Agent Quality Control Anton V Ievlev et.al. 2608.12277 null
2026-08-12 Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents Junliang Liu et.al. 2608.12273 null
2026-08-12 One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Simon Yu et.al. 2608.12253 null
2026-08-12 An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS Yuzhong Shen et.al. 2608.12249 null
2026-08-12 VICBench: A Multi-Language Benchmark for Code Vulnerability Detection Jin Lu et.al. 2608.12246 null
2026-08-12 Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs Yung-Hsu Yang et.al. 2608.12179 null
2026-08-12 Rethinking Agent Security as a Networking Problem Van Tran et.al. 2608.12172 null
2026-08-12 GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings Shivali Dalmia et.al. 2608.12133 null
2026-08-12 SAG: SQL-Retrieval Augmented Generation with Query-Time Dynamic Hyperedges Yuchao Wu et.al. 2608.12129 null
2026-08-12 SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks Tao Yu et.al. 2608.12127 null
2026-08-12 Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control Josef Liyanjun Chen et.al. 2608.12123 null
2026-08-12 No One to Blame: A Framework of Constitutive AI Unaccountability Long Hoang Nguyen et.al. 2608.12104 null
2026-08-12 CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations Xingyu Yan et.al. 2608.12002 null
2026-08-12 Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection Chaoran Chen et.al. 2608.11977 null
2026-08-12 LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation Zhixin Zhang et.al. 2608.11967 null
2026-08-11 Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration Alan Li et.al. 2608.11195 null
2026-08-11 Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents Sourabrata Mukherjee et.al. 2608.11110 null
2026-08-11 HUI360: A 360° Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation Raphael Lorenzo-Louis et.al. 2608.11051 null
2026-08-11 Who Are You Explaining To? A Multi-Agent System for Audience-Aware XAI Narratives Francesco Musicco et.al. 2608.11033 null
2026-08-11 Understanding the Architecture of Coding Agents: An Exploratory Study Using a Research Prototype Marco Tulio Valente et.al. 2608.10934 null
2026-08-11 ComBodied Agents: a New Paradigm of Human-Centric Agentic AI Qianggang Ding et.al. 2608.10915 null
2026-08-11 FormaTheoria: Constructing Large-Scale Lean Theories from Mathematical Literature $-$ Toward the Formalization of the Classification of Finite Simple Groups Tianjiao Nie et.al. 2608.10894 null
2026-08-11 VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? Xiaohongshu Inc et.al. 2608.10875 null
2026-08-11 MIRA: Medical Image Reflection for Agentic Diagnosis Shengzhi Wang et.al. 2608.10827 null
2026-08-11 A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem Suraj Kumar et.al. 2608.10760 null
2026-08-11 Mitigating Context Interference for Reliable and Efficient Search Agents Boyang Xue et.al. 2608.10743 null
2026-08-11 REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems Zixing Chen et.al. 2608.10669 null
2026-08-11 Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome Fabrizio Russo et.al. 2608.10664 null
2026-08-11 A Study of Cursorrules Files in GitHub Open Source Projects Shuang Sun et.al. 2608.10622 null
2026-08-11 Toward the Cognitive–Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent Zitong Shan et.al. 2608.10618 null
2026-08-11 On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models Md Jafrin Hossain et.al. 2608.10530 null
2026-08-11 MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows Yiqi Wang et.al. 2608.10509 null
2026-08-11 MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph Jung Hwan Lee et.al. 2608.10504 null
2026-08-11 Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases Davood Wadi et.al. 2608.10503 null
2026-08-11 From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents Caili Yu et.al. 2608.10502 null
2026-08-07 Interaction Creates Dynamical AI Behavior Absent in Isolation Bella Xinrui Li et.al. 2608.07457 null
2026-08-07 Strategy-first synthesis planning for complex natural products Daniel Armstrong et.al. 2608.07454 null
2026-08-07 SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent Mingxuan Zheng et.al. 2608.07449 null
2026-08-07 PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents Mohammad Amanlou et.al. 2608.07438 null
2026-08-07 Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Jiacheng Miao et.al. 2608.07437 null
2026-08-07 ResidencyRL: Reinforcement Learning in Simulated Clinical Environments Valentin Liévin et.al. 2608.07418 null
2026-08-07 An End-to-End Agent Auditing Engine Haoning Wang et.al. 2608.07346 null
2026-08-07 Learning Long-Term Educational Investment Policies under Residential Sorting Honglei Guo et.al. 2608.07295 null
2026-08-07 DRL-Based Secure Transmission for Rotatable Antenna-Enabled Low-Altitude ISAC Systems Chuan Liu et.al. 2608.07170 null
2026-08-07 Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory Taeil Kim et.al. 2608.07169 null
2026-08-07 NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs Aditya Katkar et.al. 2608.07167 null
2026-08-07 DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training Xucong Wang et.al. 2608.07147 null
2026-08-07 PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery Sumaiya Islam et.al. 2608.07126 null
2026-08-07 Social Facilitation of Creative Reflection: AI-agents and Humans Olga Sutskova et.al. 2608.06980 null
2026-08-07 CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows Zhu Wang et.al. 2608.06961 null
2026-08-07 Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints Paul-Peter Arslan et.al. 2608.06949 null
2026-08-07 Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery Taolin Han et.al. 2608.06931 null
2026-08-07 TRIBE: Predicting Team Performance via Communication Behavior Ensembles Ali Jalal-Kamali et.al. 2608.06926 null
2026-08-07 Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiation Massimiliano Luca et.al. 2608.06922 null
2026-08-07 Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework Jing Chen et.al. 2608.06909 null
2026-08-06 The Bitter Lesson of Tool Calling Ishan Patel et.al. 2608.06370 null
2026-08-06 AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games Boning Li et.al. 2608.06362 null
2026-08-06 Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents Praphul Chandra et.al. 2608.06353 null
2026-08-06 TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories Yunjia Qi et.al. 2608.06346 null
2026-08-06 Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents Tao Wang et.al. 2608.06312 null
2026-08-06 QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction Mutasim Fuad Sarker et.al. 2608.06294 null
2026-08-06 The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images Zhiheng Wang et.al. 2608.06270 null
2026-08-06 Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints Omid Bazgir et.al. 2608.06265 null
2026-08-06 EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Zishan Xu et.al. 2608.06197 null
2026-08-06 Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents Jiaming Wei et.al. 2608.06171 null
2026-08-06 iARCS: Iterative Agentic RL for Controllable 3D Scene Generation Saugat Adhikari et.al. 2608.06161 null
2026-08-06 Learning Globally Reusable Skills for Coding Agents Chen Yang et.al. 2608.06153 null
2026-08-06 Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture Leo Sambrook et.al. 2608.06130 null
2026-08-06 From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems Manideep Dhar et.al. 2608.06112 null
2026-08-06 When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories Xiaoqing Wu et.al. 2608.06057 null
2026-08-06 ASGE-RR: Agentic Service Graph Embedding with Revisable Reservations for Dynamic AI-Agent Calls Trond Vatten et.al. 2608.06033 null
2026-08-06 From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models Jiale Han et.al. 2608.06020 null
2026-08-06 Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation He Kong et.al. 2608.05999 null
2026-08-06 OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents Ning Xu et.al. 2608.05990 null
2026-08-06 AgentExecutor: Partial Code Execution via Agentic Context Generation Junkai Chen et.al. 2608.05959 null
2026-08-05 OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling Indraneil Paul et.al. 2608.05141 null
2026-08-05 Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models Yuezhang Peng et.al. 2608.05126 null
2026-08-05 Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming Yanting Wang et.al. 2608.05108 null
2026-08-05 CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs Hung Truong Thanh Nguyen et.al. 2608.05107 null
2026-08-05 Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite Xiawei Yue et.al. 2608.05095 null
2026-08-05 ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation Xiaoyan Gu et.al. 2608.05026 null
2026-08-05 RAC: Reference-Aware Activation Compression for Communication-Efficient Split LLM Inference Guotao Yang et.al. 2608.04991 null
2026-08-05 EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement Jun Nie et.al. 2608.04968 null
2026-08-05 CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications Brendan Smith et.al. 2608.04942 null
2026-08-05 State2State: Environment-Derived Mid-Training for LLM Agents Xuanyu Lei et.al. 2608.04934 null
2026-08-05 Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments Haoming Xu et.al. 2608.04933 null
2026-08-05 A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination Wenxiao Zhao et.al. 2608.04872 null
2026-08-05 MemoryCPT: An End-to-End Agent Memory Framework for Cost-Performance Trade-off Songxin Lei et.al. 2608.04843 null
2026-08-05 Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? Jinyi Han et.al. 2608.04828 null
2026-08-05 Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First Ishaan Bhola et.al. 2608.04804 null
2026-08-05 Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation Sarthak Harne et.al. 2608.04794 null
2026-08-05 Embedding Large Language Models into Flow Controls: An Agentic Framework for Adaptive and Trustworthy Automated Cooking Zihan Song et.al. 2608.04768 null
2026-08-05 Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems Kartikey Singh Bhandari et.al. 2608.04746 null
2026-08-05 LoginTrap: Uncovering Task-Agnostic Phishing-Style Indirect Prompt Injection Attacks against LLM-based Web Agents Longtao Guo et.al. 2608.04741 null
2026-08-05 Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools Atul Anand et.al. 2608.04719 null
2026-08-04 SocietyBench: Forecasting Counterfactual Social-World Evolution Zhenran Wang et.al. 2608.04009 null
2026-08-04 PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Shuhan Xue et.al. 2608.04003 null
2026-08-04 Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations Zizhao Hu et.al. 2608.03970 null
2026-08-04 A game theory for foundation models shows new paths to rational cooperation through similarity inference Alexander Meulemans et.al. 2608.03958 null
2026-08-04 Implementing Causal Perception: Competing SCMs and Situated Fairness Jose M. Álvarez et.al. 2608.03917 null
2026-08-04 MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning Martin Böckling et.al. 2608.03882 null
2026-08-04 ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? Tianyi Guan et.al. 2608.03874 null
2026-08-04 MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents Jiaming Chen et.al. 2608.03844 null
2026-08-04 Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers Sajjad Khan et.al. 2608.03836 null
2026-08-04 History Matters: Meta-policy Delegation with Heterogeneous Multi-agent Reinforcement Learning Ziqing Lu et.al. 2608.03833 null
2026-08-04 Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure Holly Lewis et.al. 2608.03800 null
2026-08-04 AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding Yuxiang Duan et.al. 2608.03779 null
2026-08-04 AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits Shuo Ren et.al. 2608.03738 null
2026-08-04 Accountability Asymmetry and Structural Trust in Autonomous AI Systems Nathan DeBardeleben et.al. 2608.03670 null
2026-08-04 Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Maksymilian Wolski et.al. 2608.03644 null
2026-08-04 Formal Verification of Agentic Systems over Operational Data Alejandro J. Mercado et.al. 2608.03609 null
2026-08-04 Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents William Bolton et.al. 2608.03606 null
2026-08-04 DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction Xuyang Liu et.al. 2608.03591 null
2026-08-04 GenOS: Compositional Certificates for Semantic Robustness in AI Code Generation Corrado Priami et.al. 2608.03588 null
2026-08-04 From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities Mengying Zhou et.al. 2608.03585 null
2026-08-04 A Challenge-Nonce Freshness Gap in Project Veraison’s TPM Reference Schemes, Found by Appraising Application-Layer Action Evidence End-to-End Anton Sokolov et.al. 2608.03534 null
2026-08-04 Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS Fengjunjie Pan et.al. 2608.03524 null
2026-08-04 Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks Christophe D. Hounwanou et.al. 2608.03502 null
2026-08-04 Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design Zejun Liu et.al. 2608.03501 null
2026-08-04 WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks Prince Zizhuang Wang et.al. 2608.03499 null
2026-08-04 MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification Qiming Li et.al. 2608.03474 null
2026-08-04 ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning Xiuhui You et.al. 2608.03468 null
2026-08-04 LeanMem: Simple and Efficient Long-Term Memory for LLM Agents Yuxin Liao et.al. 2608.03463 null
2026-08-04 Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory Jakub Rada et.al. 2608.03420 null
2026-08-04 Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance Can Wang et.al. 2608.03403 null
2026-08-04 Self-Evolving Coding Agents Hao Zhou et.al. 2608.03392 null
2026-08-04 Traceable Multi-Agent System for Knowledge-Based Forecasting Junhyeok Kang et.al. 2608.03339 null
2026-08-03 ACEM: A Cost Estimation Model for Agentic Software Engineering Mohammad El-Ramly et.al. 2608.02582 null
2026-08-03 RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States Yi Yang et.al. 2608.02508 null
2026-08-03 SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Yuqiao Tan et.al. 2608.02499 null
2026-08-03 Real-Time Detection and Repair of LLM Agent Failures Sunny Dubey et.al. 2608.02464 null
2026-08-03 ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Wei-Jung Huang et.al. 2608.02444 null
2026-08-03 Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce Shicheng Fan et.al. 2608.02441 null
2026-08-03 Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning Yiran Gao et.al. 2608.02422 null
2026-08-03 Analyzing GPU Performance in Virtualized Environments: A~Case Study Adel Belkhiri et.al. 2608.02414 null
2026-08-03 Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training Zhiyuan Wang et.al. 2608.02391 null
2026-08-03 PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise Abdulrahman AlRabah et.al. 2608.02372 null
2026-08-03 ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step Vernon Toh et.al. 2608.02358 null
2026-08-03 SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents Yue Yao et.al. 2608.02356 null
2026-08-03 Global Optimization and Inference-Time Region Grafting for Agentic Workflows Donghyeok Koh et.al. 2608.02353 null
2026-08-03 Qwen-CUA: Native Computer Use for (almost) Everything Dunjie Lu et.al. 2608.02352 null
2026-08-03 Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation Stefan Hut et.al. 2608.02345 null
2026-08-03 Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit Jingxi Wei et.al. 2608.02302 null
2026-08-03 MechGeo: Autoformalizing and Proving Euclidean Geometry in Lean 4 Hao Shen et.al. 2608.02295 null
2026-08-03 Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning Yiqing Liu et.al. 2608.02291 null
2026-08-03 Homebot: A Personal AI Agent for Conversational Home Assistance and Automation Shengyuan Ye et.al. 2608.02254 null
2026-08-03 PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs Haojie Hu et.al. 2608.02218 null
2026-08-02 Control Under Compression: Reliability Frontiers for Tool-Using Agents Yinghan Hou et.al. 2608.01056 null
2026-08-02 Don’t Offer What Can’t Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale Ortal Ashkenazi et.al. 2608.01050 null
2026-08-02 What Could the Agent See at 19:05? Generating Temporal Enterprise Scenarios from Real Research and Replaying Them to Evaluate Agents Tezan Sahu et.al. 2608.01042 null
2026-08-02 From AI Technical Debt to Agentic Technical Debt: A Systematic Mapping of Root Causes and Manifestations in Agentic AI Systems Muhammad Tukur et.al. 2608.01001 null
2026-08-02 PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Sudipta Paul et.al. 2608.00969 null
2026-08-02 AgenTag: Attribution of AI Coding Agents from Behavioral Fingerprints Taher A. Ghaleb et.al. 2608.00966 null
2026-08-02 Claim Plane: Reliability Gains and the Limits of Selective Concurrency for Parallel Coding Agents: A 30-Pair, Three-Seed Confirmatory Study of Deterministic Pre-Write Admission Maxim Nikolaev et.al. 2608.00947 null
2026-08-02 Neuro-Symbolic Participation Governance for Verifiable AI Agents in Open Digital Twin Ecosystems Juan Li et.al. 2608.00937 null
2026-08-02 Unified remnant models for aligned-spin, precessing, and eccentric binary black hole mergers Tousif Islam et.al. 2608.00934 null
2026-08-02 Augmented Backpressure for Decentralized Management of Agentic Networks Zuyuan Zhang et.al. 2608.00914 null
2026-08-02 Practical Online KV Cache Compaction for LLM Agents: An Empirical Study Yujian Liu et.al. 2608.00902 null
2026-08-01 Rethinking Agentic Kernel Generation for Emerging Accelerators Ruijie Gao et.al. 2608.00894 null
2026-08-01 Goal-Oriented Logic-based Semantic Communication for Neuro-Symbolic Reasoning with Applications onto Autonomous Driving Ahmet Faruk Saz et.al. 2608.00878 null
2026-08-01 Turning Interaction History into Execution State: A Runtime Layer for Long-Horizon Coding Agents Zehao Wang et.al. 2608.00808 null
2026-08-01 AgentSLABench: Evaluating and Benchmarking Agentic Systems Under Resource Constraints Meher Bhaskar Madiraju et.al. 2608.00805 null
2026-08-01 Hardware-rooted attestation for AI-agent evidence: composing IETF RATS with action evidence packages Anton Sokolov et.al. 2608.00801 null
2026-08-01 Safety Invariants for Agents Orchestrating Irreversible State Transitions: A Four-Dimensional Formalism Evaluated on Public Ledgers Zhaoming Yin et.al. 2608.00783 null
2026-08-01 Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations Xinshun Feng et.al. 2608.00711 null
2026-08-01 OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Yunhao Chen et.al. 2608.00677 null
2026-08-01 ParticleGen: A Multi-Agent System for Particle Effects Generation Junhao Zhuge et.al. 2608.00629 null
2026-07-31 TokTier: Exact Stateful Tokenization for Agentic LLM Serving Zhenyu Zhang et.al. 2607.29678 null
2026-07-31 ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Boyang Zhang et.al. 2607.29677 null
2026-07-31 Reusing Past Repairs Through Hierarchical Trajectory Abstraction for Coding Agents Yisen Xu et.al. 2607.29658 null
2026-07-31 AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers Tianyu Huai et.al. 2607.29626 null
2026-07-31 Educating the Agentic Engineer: Curricula, Collaboration, and Continuous Learning in the AI Era Mamdouh Alenezi et.al. 2607.29610 null
2026-07-31 From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale Chandra Maddila et.al. 2607.29516 null
2026-07-31 Know It, Act on It: Investigating Memory Utilization in LLM Personalization Zhaoxin Feng et.al. 2607.29433 null
2026-07-31 Beyond Component Testing: Validating Agentic AI Systems Fabio Orazio Mirto et.al. 2607.29405 null
2026-07-31 Zero-Mem: Zero-Token Memory Operations for LLM Agents Yilin Xiao et.al. 2607.29377 null
2026-07-31 SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery Jiamin Wu et.al. 2607.29347 null
2026-07-31 Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents Minghui Pan et.al. 2607.29254 null
2026-07-31 Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation Goutham Ramakrishnan et.al. 2607.29250 null
2026-07-31 Don’t Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Ruiming Liang et.al. 2607.29246 null
2026-07-31 CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents Blaise Delattre et.al. 2607.29190 null
2026-07-31 Execution-First Synthetic Tool-Use Trace Generation for LLM Agents Hafsa Ouajdi et.al. 2607.29175 null
2026-07-31 Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory Jinghan Xu et.al. 2607.29167 null
2026-07-31 Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework Leonid Kondrashov et.al. 2607.29069 null
2026-07-31 TransMem: Transforming Hidden States into Memory for Large Language Models Haodong Lei et.al. 2607.29032 null
2026-07-31 EasyBCI Agent: Towards Universal Neural Data Preprocessing for Brain-Computer Interfaces Yu Zhu et.al. 2607.29007 null
2026-07-31 Scaling Scientific Discovery Environments for Turn-Level Agentic RL Yucheng Xu et.al. 2607.28990 null
2026-07-30 AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis Bing Yan et.al. 2607.28618 null
2026-07-30 Beacon: Knowing When and How to Perform Agentic Visual Reasoning Qixun Wang et.al. 2607.28595 null
2026-07-30 Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments Haomin Qi et.al. 2607.28591 null
2026-07-30 Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs Woongkyu Lee et.al. 2607.28573 null
2026-07-30 ORCA-bench: How Ready Are Language Model Agents for Oncall? Albert Gong et.al. 2607.28545 null
2026-07-30 MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems Mao-xun Huang et.al. 2607.28527 null
2026-07-30 AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration Xinxing Ren et.al. 2607.28430 null
2026-07-30 LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger Enjun Du et.al. 2607.28374 null
2026-07-30 Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents Mingdai Yang et.al. 2607.28330 null
2026-07-30 One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence Cesare Zavattari et.al. 2607.28317 null
2026-07-30 Tycho: Active Abstraction with Programmatic World Models for ARC-AGI-3 Jens Lehmann et.al. 2607.28287 null
2026-07-30 MemHarness: Memory Is Reconstructed, Not Replayed Rong Wu et.al. 2607.28272 null
2026-07-30 EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents Luigi Sigillo et.al. 2607.28229 null
2026-07-30 FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification Haoqing Wang et.al. 2607.28225 null
2026-07-30 Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews Brian Jabarian et.al. 2607.28222 null
2026-07-30 Bridging Probabilistic LLMs and Deterministic Statistical Validation: The PROVE Multi-Agent Framework for Clinical Trial Reporting Zhaohua Lu et.al. 2607.28218 null
2026-07-30 Vibe-FDTR: An agent-oriented framework for reproducible frequency-domain thermoreflectance data analysis Fuwei Yang et.al. 2607.28200 null
2026-07-30 Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents Mingxiao Liu et.al. 2607.28165 null
2026-07-30 MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck Dongyi Liu et.al. 2607.28103 null
2026-07-30 VIG-RL: Learning to Search and Insert for Verified Image Grounding Qinhan Yu et.al. 2607.28055 null
2026-07-29 Can AI agents conduct open-ended AI research? Early evidence from two case studies Peter Kirgis et.al. 2607.27191 null
2026-07-29 Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork Peter Tisnikar et.al. 2607.27177 null
2026-07-29 OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding Jingbo Zhou et.al. 2607.27155 null
2026-07-29 MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis Yihao Chen et.al. 2607.27146 null
2026-07-29 Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents Yicheng Feng et.al. 2607.27083 null
2026-07-29 Constraining the shape of dark matter haloes using only starlight II. Tests of the technique with objects of known gravitational potential Jorge Sanchez Almeida et.al. 2607.27001 null
2026-07-29 AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents Ruoyu Wang et.al. 2607.26998 null
2026-07-29 TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning Jinhu Qi et.al. 2607.26977 null
2026-07-29 Assurance-Scoped Reliability for Agentic Networks: Capturing the State That Matters Bilgehan Erman et.al. 2607.26953 null
2026-07-29 VITAL-RAG: Invariance Race for Context Allocation in Coding Agents Zijian Lu et.al. 2607.26937 null
2026-07-29 What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation Vishisht Choudhary et.al. 2607.26935 null
2026-07-29 CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents Qianru Li et.al. 2607.26910 null
2026-07-29 Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents Amirmohammad Farzaneh et.al. 2607.26865 null
2026-07-29 Before Agents Speak: Pre-hoc Failure Risk Inference in Multi-Agent Systems Shi Lin et.al. 2607.26836 null
2026-07-29 Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions Shi Lin et.al. 2607.26820 null
2026-07-29 A First Look at Coding Agents’ Compliance with AI Contribution Rules in Open-Source Communities Wenhao Yang et.al. 2607.26819 null
2026-07-29 Practice Makes Policies: Bootstrapping and Consolidating Robotic Capabilities from Zero Human Demonstrations Jialiang Li et.al. 2607.26809 null
2026-07-29 SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response Lehan Wang et.al. 2607.26791 null
2026-07-29 SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution Zhiyuan Yao et.al. 2607.26784 null
2026-07-29 CodeSpec: Dual Executable Specifications for Agentic Long-Horizon Feature Development Peiding Wang et.al. 2607.26777 null
2026-07-29 Metis: Memory Foundation Model Zeyu Zhang et.al. 2607.26760 null
2026-07-29 UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks Zhilun Zhou et.al. 2607.26724 null
2026-07-29 PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems Kaiwen Jiang et.al. 2607.26710 null
2026-07-29 Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection Yikun Li et.al. 2607.26656 null
2026-07-28 UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams Siyu Xia et.al. 2607.26017 null
2026-07-28 Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Fengxiang Wang et.al. 2607.25993 null
2026-07-28 \textsc{IH-Benchmark}: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications Conor McCauley et.al. 2607.25987 null
2026-07-28 Who is scientific code for? Maintaining human-readable landmarks in agent-written code Elle O’Brien et.al. 2607.25975 null
2026-07-28 Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA Carlos Celemin et.al. 2607.25921 null
2026-07-28 Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks Ravi Kant Sharma et.al. 2607.25914 null
2026-07-28 Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Stefan Krsteski et.al. 2607.25891 null
2026-07-28 Distributing Security Controls Through Harness Engineering William Robert Gore et.al. 2607.25890 null
2026-07-28 RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement Fanqing Meng et.al. 2607.25886 null
2026-07-28 Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks Bart Custers et.al. 2607.25877 null
2026-07-28 HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs Yu Hao et.al. 2607.25853 null
2026-07-28 Lowering the implementation barrier of neutral-atom quantum computing with agentic workflows Constantin Dalyac et.al. 2607.25834 null
2026-07-28 C-RE-ACT: Causal RE-ACTing Agent for O-RAN Forensic Triage Pau Baguer et.al. 2607.25828 null
2026-07-28 Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL Jiabao Ji et.al. 2607.25816 null
2026-07-28 Towards Faithful Sentimental Image Captioning via Evidence-Aware Multi-Agent Reasoning Tiecheng Cai et.al. 2607.25789 null
2026-07-28 WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing Hao Liang et.al. 2607.25765 null
2026-07-28 Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction Xinyi Hong et.al. 2607.25718 null
2026-07-28 F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill Florian Krebs et.al. 2607.25637 null
2026-07-28 Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines Federico Cabitza et.al. 2607.25620 null
2026-07-28 SkillGate: Cost Efficient Runtime Malicious Skill File Detection in Coding Agents Rui Yang et.al. 2607.25619 null
2026-07-27 Data Pyramid for Embodied Manipulation Yifan Ye et.al. 2607.24744 null
2026-07-27 ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding Hangjie Yuan et.al. 2607.24743 null
2026-07-27 Kimi K3: Open Frontier Intelligence Kimi Team et.al. 2607.24653 null
2026-07-27 Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents Arseny Kravchenko et.al. 2607.24625 null
2026-07-27 Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair Xueping Gao et.al. 2607.24604 null
2026-07-27 SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents Hang Ni et.al. 2607.24588 null
2026-07-27 CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding Jinlong Yang et.al. 2607.24582 null
2026-07-27 Designing Within the Lines: Practitioners’ Perspectives and Visualisation Tool Evaluation in the Arabic Context Muna Alebri et.al. 2607.24571 null
2026-07-27 Distributed Coordination for Resilient Multi-UAV Remote Sensing: A Photovoltaic Inspection Case Study Guillermo GP-Lenza et.al. 2607.24482 null
2026-07-27 Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families Dushyant Sharma et.al. 2607.24339 null
2026-07-27 Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG Mohamed Amine Janati et.al. 2607.24313 null
2026-07-27 INS-ActBench: A Comprehensive Benchmark for Assessing Professional Actuarial Capability of Large Language Models Changyu Chen et.al. 2607.24273 null
2026-07-27 6G: From Connectivity Infrastructure to Guaranteed Digital Services David Soldani et.al. 2607.24185 null
2026-07-27 Falsifiable Commitment Planning for Self-Correcting Web Agents Guangyi Liu et.al. 2607.24167 null
2026-07-27 Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness Yang Li et.al. 2607.24162 null
2026-07-27 MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents Yiwen Ma et.al. 2607.24097 null
2026-07-27 TCellAlign: Cross-study T-cell Populations Alignment with Nomenclature-Guided Multi-Agent Workflow Pengyu Xie et.al. 2607.24093 null
2026-07-27 Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation Mohan Manivannan et.al. 2607.24006 null
2026-07-27 ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents Wenhao Lan et.al. 2607.23999 null
2026-07-27 HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows Qingyi Yang et.al. 2607.23983 null
2026-07-26 E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Weihuang Zheng et.al. 2607.23722 null
2026-07-26 LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratories Haobo Wang et.al. 2607.23704 null
2026-07-26 Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems Mingzhou Fan et.al. 2607.23678 null
2026-07-26 Rethinking Logic Optimization Operators: Theory-Derived Operator Compression via Agentic Source Analysis Keren Zhu et.al. 2607.23672 null
2026-07-26 Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents Aayush Kumar et.al. 2607.23670 null
2026-07-26 Where Is the Cost of Third-Party API Routers in Agentic Software Development? Donghao Fu et.al. 2607.23624 null
2026-07-26 JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Yunlong Lin et.al. 2607.23588 null
2026-07-26 Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents Zhaoxi Zhang et.al. 2607.23586 null
2026-07-26 Private Again: AI Agents Restore Anonymity—Foreclosing Discrimination and Its Proof Anirban Mukherjee et.al. 2607.23539 null
2026-07-26 Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis Xinhao Yao et.al. 2607.23524 null
2026-07-26 Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents Xinyu Gao et.al. 2607.23444 null
2026-07-26 Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels Haining Zheng et.al. 2607.23438 null
2026-07-25 Constitutional governance for societies of AI agents in the built environment: a research agenda Ali Ghoroghi et.al. 2607.23336 null
2026-07-25 AlloBench: Measuring Online Tool Allocation Capability in LLM Agents Daniel Wang et.al. 2607.23332 null
2026-07-25 Accountable yet Anonymous AI Agents - Split-Knowledge Binding in National Agent-Identity Layer in China Yifan He et.al. 2607.23207 null
2026-07-25 False Prophets: On the Security of World Models in Agentic Systems Erik Imgrund et.al. 2607.23147 null
2026-07-25 SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows Summer Sun et.al. 2607.23123 null
2026-07-25 VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy Xinyan Zhong et.al. 2607.23006 null
2026-07-25 WCM: World-Cognition Model for Generalizable Human-Robot Interaction Yuzhen Chen et.al. 2607.22999 null
2026-07-25 Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline Qing Yang et.al. 2607.22997 null
2026-07-24 Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Siyuan Huang et.al. 2607.22529 null
2026-07-24 The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents Darshan Tank et.al. 2607.22520 null
2026-07-24 CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference Jiyuan Tan et.al. 2607.22511 null
2026-07-24 Where FactsGo Missing: A LayerwiseTaxonomy and Per-Layer Attribution of Information Omissionin Air-Gapped LLM Agent Pipelines Santhiya Rajan et.al. 2607.22448 null
2026-07-24 Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture Halil Burak Noyan et.al. 2607.22445 null
2026-07-24 A Human-Augmenting Agentic Workflow for Observational Causal Inference Winston Chou et.al. 2607.22443 null
2026-07-24 Vibe Coding: An Experiment with Test-Driven Development Moritz Mock et.al. 2607.22406 null
2026-07-24 A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation Fin Gentzen et.al. 2607.22400 null
2026-07-24 Agentic Root Cause Analysis through Evidence-Grounded Reasoning Amaury Wei et.al. 2607.22385 null
2026-07-24 IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation Varun Gumma et.al. 2607.22375 null
2026-07-24 SMEFT-Pheno-Agent: a natural-language-driven AI agent for machine-learning-assisted Standard Model Effective Field Theory phenomenology Yu-Chen Guo et.al. 2607.22331 null
2026-07-24 Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG Chuangtao Ma et.al. 2607.22319 null
2026-07-24 Agentic CPU-GPU Scheduling for Heterogeneous AI Workloads Tianxi Lu et.al. 2607.22242 null
2026-07-24 AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment Ziyao Huang et.al. 2607.22241 null
2026-07-24 DeFiScreener: Efficient DeFi Attack Pre-screening in Smart Contracts via Historical Case Matching Rui Cao et.al. 2607.22184 null
2026-07-24 Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents Valentin Tablan et.al. 2607.22157 null
2026-07-24 Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode Nanbeige Lab et.al. 2607.22083 null
2026-07-24 IDSTune: A Multi-Agent Collaborative Framework for Integrated Database System Tuning Yiyan Li et.al. 2607.22031 null
2026-07-24 Are Production Cloud Skills Adequately Tested? Measuring and Governing Skill Test Coverage in Practice Haotian Si et.al. 2607.22015 null
2026-07-24 Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents Suman Navaratnarajah et.al. 2607.22014 null
2026-07-23 OpenForgeRL: Train Harness-native Agents in Any Environment Xiao Yu et.al. 2607.21557 null
2026-07-23 Benchmarking Agents for Proving Theorems in Quantum Algorithms and Quantum Information Lei Zhang et.al. 2607.21533 null
2026-07-23 GS-Agent: Creating 4D Physical Worlds With Generative Simulation Hongxin Zhang et.al. 2607.21522 null
2026-07-23 Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems Gaurav Dadhich et.al. 2607.21503 null
2026-07-23 Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry Natan Levy et.al. 2607.21495 null
2026-07-23 Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks Mack Nixon et.al. 2607.21482 null
2026-07-23 AREX: Towards a Recursively Self-Improving Agent for Deep Research Shuqi Lu et.al. 2607.21461 null
2026-07-23 PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning Yipeng Shi et.al. 2607.21419 null
2026-07-23 VoLN: Vision-Only Long-Horizon Navigation—Paradigm, Benchmark, and Method Jiabin Lou et.al. 2607.21400 null
2026-07-23 FedAgentKE: Federated Semantic Knowledge Evolution for Heterogeneous Agents Weihao Li et.al. 2607.21361 null
2026-07-23 Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation M. Llambí-Morillas et.al. 2607.21325 null
2026-07-23 GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG Paolo Pedinotti et.al. 2607.21324 null
2026-07-23 The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents – and What Actually Works Yu Wang et.al. 2607.21273 null
2026-07-23 ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders Zhongyuan Peng et.al. 2607.21217 null
2026-07-23 Explainability Framework for Policy-Aware Autonomous Agents Heather Merhout et.al. 2607.21209 null
2026-07-23 Causal-AgentIR: Self-Evolving Causal Memory for Adaptive Image Restoration Agents Hu Gao et.al. 2607.21125 null
2026-07-23 AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning Qinfeng Li et.al. 2607.21106 null
2026-07-23 HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices Wei Liu et.al. 2607.21019 null
2026-07-23 EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization Lihuang Fang et.al. 2607.21013 null
2026-07-23 Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents Swapnanil Saha et.al. 2607.20972 null
2026-07-21 OmniReasoner: Thinking with Long Audio-Video via Native Tool Use Yu Chen et.al. 2607.19339 null
2026-07-21 CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents Qijia He et.al. 2607.19338 null
2026-07-21 ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D Lena Libon et.al. 2607.19321 null
2026-07-21 Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes Daniel Pearson et.al. 2607.19297 null
2026-07-21 BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance Harmon Bhasin et.al. 2607.19262 null
2026-07-21 FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents Xianfu Cheng et.al. 2607.19238 null
2026-07-21 HACO: Hedged Agent Computing for Reliable LLM Systems Enhan Li et.al. 2607.19215 null
2026-07-21 Teleportation Game: Quantum Teleportation in Multi-Agent Systems for Interactive Music Eduardo Reck Miranda et.al. 2607.19212 null
2026-07-21 Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning Ubayd Ali Bapoo et.al. 2607.19117 null
2026-07-21 FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling Jialong Zuo et.al. 2607.19038 null
2026-07-21 CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement Zhihao Yang et.al. 2607.19036 null
2026-07-21 Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio Jialian Li et.al. 2607.18985 null
2026-07-21 Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts Haodi Fan et.al. 2607.18970 null
2026-07-21 TraceDev: A Traceability-Driven Multi-agent Framework for Requirement-to-Code Development Mingyu Chen et.al. 2607.18886 null
2026-07-21 PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents Tianyue Jiang et.al. 2607.18859 null
2026-07-21 Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents SangJin Park et.al. 2607.18826 null
2026-07-21 AgentTrails: Towards Trust and Reuse for Agentic Tasks Eden Wu et.al. 2607.18816 null
2026-07-21 AI Tour Meeting: Group Travel Planning by LLM Agents Daisuke Kikuta et.al. 2607.18806 null
2026-07-21 AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents Kunlun Zhu et.al. 2607.18754 null
2026-07-21 Beyond Text Editing: Algebraic Manipulation of Source Code Kevin Pulo et.al. 2607.18742 null
2026-07-20 SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Yuhang Wang et.al. 2607.18213 null
2026-07-20 FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Krish Agarwal et.al. 2607.18171 null
2026-07-20 TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization Alex Mathai et.al. 2607.18161 null
2026-07-20 O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning Mei Yuan et.al. 2607.18142 null
2026-07-20 AI Agent Communications in AI-Native 6G Network: Status, Challenges and Opportunities Qiang Duan et.al. 2607.18138 null
2026-07-20 FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering Jijun Chi et.al. 2607.18102 null
2026-07-20 Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data Nursultan Askarbekuly et.al. 2607.18064 null
2026-07-20 Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security Devina Jain et.al. 2607.18063 null
2026-07-20 Test Coverage Analysis of Agentic Pull Requests Atish Kumar Dipongkor et.al. 2607.18057 null
2026-07-20 Evidence-in-the-Loop: Trace-Driven Optimization for Customer-Service LLM Agents Chunming Wu et.al. 2607.18039 null
2026-07-20 Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go? Yimeng Chen et.al. 2607.17986 null
2026-07-20 Harness Engineering for LLM-Driven GPU Kernel Generation Yue Shui et.al. 2607.17979 null
2026-07-20 RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control Di Lu et.al. 2607.17951 null
2026-07-20 How Agent Skills Fail under Long Contexts: A White-Box Study in Code Auditing Yue Xue et.al. 2607.17937 null
2026-07-20 (Over)Reliance on Test Agents in AI-Assisted Software Testing Eduard Paul Enoiu et.al. 2607.17927 null
2026-07-20 Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss Kemal Devrim Kafadar et.al. 2607.17914 null
2026-07-20 Stress Testing Concept Erasure with Large Language Model Agents Yuyang Xue et.al. 2607.17890 null
2026-07-20 Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI Bogdan Raduta et.al. 2607.17883 null
2026-07-20 Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory Ganesh Senrayan et.al. 2607.17879 null
2026-07-20 PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model Li Xian et.al. 2607.17806 null
2026-07-17 Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs Like Liu et.al. 2607.16193 null
2026-07-17 When Does Muon Help Agentic Reinforcement Learning? Kai Ruan et.al. 2607.16169 null
2026-07-17 When Do Multi-Agent Systems Help? An Information Bottleneck Perspective Wendi Yu et.al. 2607.16133 null
2026-07-17 ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning Binglin Zhou et.al. 2607.16131 null
2026-07-17 Student Evaluation of Repeated AI Feedback Across a Semester of Writing Andres Karjus et.al. 2607.16115 null
2026-07-17 LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization Mazene Ameur et.al. 2607.16066 null
2026-07-17 Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning Ajay Patel et.al. 2607.16057 null
2026-07-17 SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery SciForge Team et.al. 2607.16038 null
2026-07-17 DebrisTracer: Reliable Tracking in Hypervelocity Impact Fast Imaging Théophane Loloum et.al. 2607.15986 null
2026-07-17 Code-Poisoning Property Inference Attacks Xukun Luan et.al. 2607.15970 null
2026-07-17 DSWorld: A Data Science World Model for Efficient Autonomous Agents Zherui Yang et.al. 2607.15901 null
2026-07-17 Yarrow: Reconciling Effects Handlers and Region-Based Memory Management Anders Alnor Mathiasen et.al. 2607.15876 null
2026-07-17 Agentic Synthesis against Counterexample-Supplemented Sketches Muness Castle et.al. 2607.15854 null
2026-07-17 Knowledge-Centric Agents for Workflow Generation Zhendong Li et.al. 2607.15845 null
2026-07-17 AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets Ming Chen et.al. 2607.15781 null
2026-07-17 Making Agent-Mediated Contributions Governable: A Project-Level Governance Manifest for Open-Source AI Collaboration Jinjin Gao et.al. 2607.15769 null
2026-07-17 Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents Lujia Zhang et.al. 2607.15715 null
2026-07-17 Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports Jingyi Chen et.al. 2607.15684 null
2026-07-17 Beyond Detection: Agentic Attack Synthesis and Simulation for Smart Contracts Xianhao Zhang et.al. 2607.15673 null
2026-07-17 ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning Shuaiyu Zhou et.al. 2607.15660 null
2026-07-16 Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Paul Kassianik et.al. 2607.15263 null
2026-07-16 SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Yuyao Zhang et.al. 2607.15257 null
2026-07-16 ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors Christos Korgialas et.al. 2607.15246 null
2026-07-16 When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space Weimeng Wang et.al. 2607.15218 null
2026-07-16 Plover: Steering GUI Agents through Plan-Centric Interaction Madhumitha Venkatesan et.al. 2607.15193 null
2026-07-16 Can We Trust Item Response Theory for AI Evaluation? Han Jiang et.al. 2607.15190 null
2026-07-16 Scaling Behavior Foundation Model for Humanoid Robots Weishuai Zeng et.al. 2607.15163 null
2026-07-16 Setup Complete, Now You Are Compromised: Weaponizing Setup Instructions Against AI Coding Agents Aadesh Bagmar et.al. 2607.15143 null
2026-07-16 Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents Dylan Van Mulders et.al. 2607.15095 null
2026-07-16 BrainPilot: Automating Brain Discovery with Agentic Research Haoxuan Li et.al. 2607.15079 null
2026-07-16 ANet Patu-1: The Value of Connection in the Agent Network Mu Yuan et.al. 2607.15053 null
2026-07-16 OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Chengyu Shen et.al. 2607.14989 null
2026-07-16 LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Changhai Zhou et.al. 2607.14952 null
2026-07-16 Human-Robot Interaction in GenAI Architectures via the Agent-Client Protocol Jesus Moncada-Ramirez et.al. 2607.14919 null
2026-07-16 FirmPilot: Evidence-Guided Multi-Agent Environment Recovery for IoT Firmware Rehosting Yanbing Shen et.al. 2607.14903 null
2026-07-16 StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows Sizhong Qin et.al. 2607.14896 null
2026-07-16 Proof-or-Stop: Don’t Trust the Agent, Trust the Evidence – Loop Engineering for Verifiable Evidence-Gated Lifecycle Control Jek Huang et.al. 2607.14890 null
2026-07-16 CAMB v2: cosmological power spectra for high-precision surveys Antony Lewis et.al. 2607.14854 null
2026-07-16 Ground-Side Mission Plan Compilation with Policy-as-Code Guardrails for Cloud-Native Satellite Platforms Hsiu-Chi Tsai et.al. 2607.14798 null
2026-07-16 SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Jinyang Wu et.al. 2607.14777 null
2026-07-15 Early Adoption of Agentic Coding Tools by GitHub Projects Maliha Noushin Raida et.al. 2607.14037 null
2026-07-15 TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Leitian Tao et.al. 2607.13988 null
2026-07-15 Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation Sanket Badhe et.al. 2607.13987 null
2026-07-15 NNStar: An end-to-end AI agent for nuclear matter and neutron star physics Yao Ma et.al. 2607.13930 null
2026-07-15 Experience Memory Graph: One-Shot Error Correction for Agents Wenjun Wang et.al. 2607.13884 null
2026-07-15 Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild Ting Lei et.al. 2607.13881 null
2026-07-15 SPyCE: Skill-Policy Co-evolution for Multimodal Agents Ru Zhang et.al. 2607.13854 null
2026-07-15 EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent Junlong Li et.al. 2607.13792 null
2026-07-15 How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement Alexandra E. Michael et.al. 2607.13718 null
2026-07-15 AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Zichen Ding et.al. 2607.13705 null
2026-07-15 Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity Xiaotian Luo et.al. 2607.13683 null
2026-07-15 When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects Yongren Shi et.al. 2607.13679 null
2026-07-15 Explaining Reinforcement Learning Agents via Inductive Logic Programming Celeste Veronese et.al. 2607.13655 null
2026-07-15 Exploratory, Communicative, and Deployable: Vision-Driven Embodied Agents for Open-World Mobile Manipulation Boyu Mi et.al. 2607.13653 null
2026-07-15 UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following Kun Yu et.al. 2607.13621 null
2026-07-15 STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle Sagar Deb et.al. 2607.13618 null
2026-07-15 Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System David Krongauz et.al. 2607.13608 null
2026-07-15 Analogical Deep Research: Retrieving and Integrating Historical Analogies for Foresight Analysis Yongqiang Chen et.al. 2607.13602 null
2026-07-15 Active Trust Management for Successful Human-Robot Teaming: Moving from a Trust Repair to a Trust Satisficing Perspective Nicola Webb et.al. 2607.13595 null
2026-07-15 SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing Tianyu Chen et.al. 2607.13594 null
2026-07-14 Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution Junjie Yin et.al. 2607.13034 null
2026-07-14 PalmClaw: A Native On-Device Agent Framework for Mobile Phones Hongru Cai et.al. 2607.13027 null
2026-07-14 Software Supply Chains are Dead: Use-Case-Oriented Regeneration Tanmay Singla et.al. 2607.13021 null
2026-07-14 Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition Bruce Coburn et.al. 2607.12911 null
2026-07-14 Hy-Embodied-VLM-1.0: Efficient Physical-World Agents Ziyi Wang et.al. 2607.12894 null
2026-07-14 MetaInfer: A Knowledge Only LLM Inference Engine Generator SKILL Toolbox Zhenwen Miao et.al. 2607.12875 null
2026-07-14 Human-AI Agent Interaction as a Neuroplastic Training Environment Eranga Bandara et.al. 2607.12823 null
2026-07-14 Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents Xing Zhang et.al. 2607.12790 null
2026-07-14 Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration Quanyan Zhu et.al. 2607.12662 null
2026-07-14 Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs Junyu Ren et.al. 2607.12650 null
2026-07-14 A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism Chengguang Gan et.al. 2607.12640 null
2026-07-14 Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making? Minh Khoi Ho et.al. 2607.12631 null
2026-07-14 Instance-Enriched Semantic Maps for Visual Language Navigation Jiho Hong et.al. 2607.12630 null
2026-07-14 KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill Yunxin Li et.al. 2607.12625 null
2026-07-14 PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis Junhui Wang et.al. 2607.12624 null
2026-07-14 Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing Amin Beheshti et.al. 2607.12619 null
2026-07-14 How Agentic Is Agentic Commerce? A Population-Scale Measurement of x402 Adoption and Authenticity Shengchen Ling et.al. 2607.12575 null
2026-07-14 Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Yubo Wang et.al. 2607.12463 null
2026-07-14 Trust but Verify? Uncovering the Security Debt of Autonomous Coding Agents A H M Nazmus Sakib et.al. 2607.12428 null
2026-07-14 Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions Huihao Jing et.al. 2607.12406 null
2026-07-14 Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents Yaopei Zeng et.al. 2607.12397 null
2026-07-14 PM-Bench: Evaluating Prospective Memory in LLM Agents Genglin Liu et.al. 2607.12385 null
2026-07-14 Skills That Don’t Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents Weifeng Yuan et.al. 2607.12340 null
2026-07-13 A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation Yunhai Feng et.al. 2607.11874 null
2026-07-13 Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding Kerui Chen et.al. 2607.11844 null
2026-07-13 When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems Yibo Hu et.al. 2607.11751 null
2026-07-13 Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming Xutao Mao et.al. 2607.11698 null
2026-07-13 From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence Yuanzhi Liang et.al. 2607.11689 null
2026-07-13 Heuristic Learning for Active Flow Control Using Coding Agents Paul Garnier et.al. 2607.11565 null
2026-07-13 PaperRouter-Agent: A Content-Grounded LLM Agent for Personalized Hierarchical Paper Routing Keshen Zhou et.al. 2607.11564 null
2026-07-13 Linux disk encryption and self-encrypting drives – A case study on Opal2 drives security Milan Brož et.al. 2607.11563 null
2026-07-13 Towards Human-level Dexterous Teleoperation Puhao Li et.al. 2607.11481 null
2026-07-13 UMoE:Unlocking Every Expert in Domain-Specific Training Xuefeng Li et.al. 2607.11444 null
2026-07-13 ToFu: A White-Box, Token-Efficient Agent Harness for Researchers Junhao Ruan et.al. 2607.11423 null
2026-07-13 Agentic Routing: The Harness-Native Data Flywheel Xinchen Liu et.al. 2607.11399 null
2026-07-13 TerraRepair: A Tool-Grounded LLM Agent for Infrastructure-as-Code Repair Minase Mekete Mengistu et.al. 2607.11390 null
2026-07-13 A Glimpse into Long-term Physical Coexistence with Intelligent Robots Weiqi Jin et.al. 2607.11377 null
2026-07-13 OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis Yongqian Sun et.al. 2607.11357 null
2026-07-13 Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents Chenglin Yu et.al. 2607.11346 null
2026-07-13 Toward AI-Agent-Driven Particle Transport Simulations: Implementation of AI-Assisted Workflows for PHITS Tatsuhiko Sato et.al. 2607.11309 null
2026-07-13 FlowArk: Boosting Agentic Data-flow Analysis for Android Apps via Context-Aware Knowledge Reuse Yiming Zhang et.al. 2607.11308 null
2026-07-13 Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization Tian-Shuo Liu et.al. 2607.11307 null
2026-07-13 Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation Praneeth Narisetty et.al. 2607.11288 null
2026-07-10 VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents Katherine Swinea et.al. 2607.09653 null
2026-07-10 Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory Chengkai Zhu et.al. 2607.09632 null
2026-07-10 LLM for EDA in Front-End Design: Challenges and Opportunities Kangwei Xu et.al. 2607.09616 null
2026-07-10 Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation Kaiji Zhou et.al. 2607.09600 null
2026-07-10 Writing Bug Reports for Software Repair Agents: What Information Matters Most? Vincenzo Luigi Bruno et.al. 2607.09553 null
2026-07-10 Failure as a Process: An Anatomy of CLI Coding Agent Trajectories Xiangxin Zhao et.al. 2607.09510 null
2026-07-10 All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models Pan Li et.al. 2607.09502 null
2026-07-10 Shared Selective Persistent Memory for Agentic LLM Systems Sanjana Pedada et.al. 2607.09493 null
2026-07-10 ProofCouncil: An LLM Agent for Solving Open Mathematical Problems Johannes Schmitt et.al. 2607.09474 null
2026-07-10 Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review Jingbo Chen et.al. 2607.09403 null
2026-07-10 Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks Nuocheng Yang et.al. 2607.09330 null
2026-07-10 LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Yanzhen Chen et.al. 2607.09322 null
2026-07-10 Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents Izumi Takahara et.al. 2607.09195 null
2026-07-10 Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning Xingzhi Qian et.al. 2607.09179 null
2026-07-10 Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift Dan C. Hsu et.al. 2607.09175 null
2026-07-10 Present but Rescaled: Chat-to-Agent Transfer of Additive Activation Steering Lucas Pinto et.al. 2607.09156 null
2026-07-10 ReProAgent: Tool-Augmented Multi-Stage Agentic Generation of Bug Reproduction Tests from Issue Reports Quanjun Zhang et.al. 2607.09123 null
2026-07-10 Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows Quanjun Zhang et.al. 2607.09101 null
2026-07-10 Inside the Skill Market: From Software Engineering Activities to Reusable Agent Skills Jialun Cao et.al. 2607.09065 null
2026-07-10 ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning Kunbo Zhang et.al. 2607.09059 null
2026-07-09 UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks Zhekai Chen et.al. 2607.08768 null
2026-07-09 DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation Yunchao Yao et.al. 2607.08751 null
2026-07-09 Workflow as Knowledge: Semantic Persistence for LLM-Mediated Workflows Emanuele Quinto et.al. 2607.08740 null
2026-07-09 ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation QiHong Chen et.al. 2607.08691 null
2026-07-09 SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets Shilin Ou et.al. 2607.08681 null
2026-07-09 WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search Xiaoshuai Song et.al. 2607.08662 null
2026-07-09 Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study Eugene Ng Yi Sheng et.al. 2607.08652 null
2026-07-09 Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Ali Larian et.al. 2607.08647 null
2026-07-09 UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing Xinlong Zhao et.al. 2607.08646 null
2026-07-09 Early to Share, Late to Save: Synchronisation-Driven Communication Gating in Bandwidth-Constrained Cooperative VLN Arav Gupta et.al. 2607.08504 null
2026-07-09 The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality Masahiro Fujita et.al. 2607.08495 null
2026-07-09 Harness VLA: Steering Frozen VLAs into Reliable Manipulation Primitives via Memory-Guided Agents Yixian Zhang et.al. 2607.08448 null
2026-07-09 OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice Qian Jiang et.al. 2607.08423 null
2026-07-09 Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination Runzhe Liu et.al. 2607.08403 null
2026-07-09 TRACE: A Two-Channel Robust Attribution Watermark via Complementary Embeddings for LLM-Agent Trajectories Zheng Gao et.al. 2607.08400 null
2026-07-09 Token-Flow Firewall: Semantic Runtime Auditing for Persistent AI Agents Puji Wang et.al. 2607.08395 null
2026-07-09 Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction Sophia Koehler et.al. 2607.08233 null
2026-07-09 Out of Sight: Compression-Aware Content Protection against Agentic Crawlers Xuefei Wang et.al. 2607.08180 null
2026-07-09 ASMR: Agentic Schema Generation for Ship Maintenance Report Writing Sohrab Namazi Nia et.al. 2607.08177 null
2026-07-09 Prismata: Confining Cross-Site Prompt Injection in Web Agents Corban Villa et.al. 2607.08147 null
2026-07-08 Agent Delivery Engineering Predictive Reliability Framework Dexing Liu et.al. 2607.07689 null
2026-07-08 SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents Tianming Sha et.al. 2607.07676 null
2026-07-08 A hierarchical memory architecture overcomes context limits in long-horizon multi-agent computational modeling Shivendra G. Tewari et.al. 2607.07666 null
2026-07-08 ATLAS: Automated HLS for DL-Optimized FPGAs Ruthwik Reddy Sunketa et.al. 2607.07643 null
2026-07-08 Future Confidence Distillation in Large Language Models Sahil Kale et.al. 2607.07626 null
2026-07-08 Rethinking Code Performance Benchmarks for LLMs Nhat Minh Le et.al. 2607.07619 null
2026-07-08 What Makes a Good Bug Report for an AI Agent? Lara Khatib et.al. 2607.07593 null
2026-07-08 Creativity from Friction: Human-AI Interaction for Exploratory Structural Design Ricardo Maia Avelino et.al. 2607.07521 null
2026-07-08 Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Zhenyu Hou et.al. 2607.07508 null
2026-07-08 Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents Harry Owiredu-Ashley et.al. 2607.07474 null
2026-07-08 SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis Songhan Wang et.al. 2607.07467 null
2026-07-08 Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions Yang Shi et.al. 2607.07461 null
2026-07-08 Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents Vikas Reddy et.al. 2607.07405 null
2026-07-08 Agentic Data Environments Elaine Ang et.al. 2607.07397 null
2026-07-08 Physics-Audited Agentic Discovery in Scientific Machine Learning Diab W. Abueidda et.al. 2607.07379 null
2026-07-08 From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents Haipeng Ding et.al. 2607.07321 null
2026-07-08 Predicting LLM Safety Before Release by Simulating Deployment Marcus Williams et.al. 2607.07184 null
2026-07-08 Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning Zetian Hu et.al. 2607.07178 null
2026-07-08 Certifying Ghosts: How Cybersecurity AI Agents Break the EU Cyber Resilience Act Víctor Mayoral-Vilches et.al. 2607.07109 null
2026-07-08 Seeing and Reflecting: Multimodal Memory-Enhanced Agent Collaboration for Recommendation Hao Cong et.al. 2607.07108 null
2026-07-07 Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade Kai Ruan et.al. 2607.06503 null
2026-07-07 Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities So Hasegawa et.al. 2607.06482 null
2026-07-07 From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b Taeyun Roh et.al. 2607.06452 null
2026-07-07 An Experimental Design Approach to Evaluating Agentic AI’s Autonomous Model Discovery Hao He et.al. 2607.06413 null
2026-07-07 RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Evgeny Shilov et.al. 2607.06411 null
2026-07-07 VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery Jiazi Wang et.al. 2607.06374 null
2026-07-07 Harnessing Code Agents for Automatic Software Verification Shuangxiang Kan et.al. 2607.06341 null
2026-07-07 AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operation Chenyu Zhao et.al. 2607.06273 null
2026-07-07 Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale Ziting Wang et.al. 2607.06233 null
2026-07-07 Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows Tianyang Liu et.al. 2607.06229 null
2026-07-07 Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Yijun Zhang et.al. 2607.06223 null
2026-07-07 LogicHunter: Testing LLM Agent Frameworks with an Agentic Oracle Minghui Long et.al. 2607.06195 null
2026-07-07 What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents Rui Shu et.al. 2607.06184 null
2026-07-07 EAGOR: Embodied Reasoning in Omni-direction Shriram Damodaran et.al. 2607.06165 null
2026-07-07 LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability Chenxu Wang et.al. 2607.06157 null
2026-07-07 When Does Tool Use Increase the Expressive Power of Finite-Precision Recurrent Models? Nikola Zubić et.al. 2607.06155 null
2026-07-07 CurateEvo: Data-Curation Evolving for Agentic Post-Training Dingzirui Wang et.al. 2607.06140 null
2026-07-07 Causal Inference with Video Features as Treatments Kentaro Nakamura et.al. 2607.06126 null
2026-07-07 WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation Wei Dong et.al. 2607.06118 null
2026-07-07 Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development Rohit Mehra et.al. 2607.06101 null
2026-07-06 CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Yujiang Li et.al. 2607.05378 null
2026-07-06 Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation Jiaqi Peng et.al. 2607.05377 null
2026-07-06 SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints Dylan Zongmin Liu et.al. 2607.05363 null
2026-07-06 OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement Adriana Laurindo Monteiro et.al. 2607.05346 null
2026-07-06 PiSAs: Benchmarking Contextual Integrity in Multi-User Agentic Systems Shubham Gupta et.al. 2607.05318 null
2026-07-06 MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution Zefeng Wang et.al. 2607.05297 null
2026-07-06 Untrusted Content Masking for Web Agents with Security Guarantees Kristina Nikolić et.al. 2607.05277 null
2026-07-06 Latent Programming Horizons in Coding Agents André Silva et.al. 2607.05188 null
2026-07-06 ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization Mahnoor Shahid et.al. 2607.05185 null
2026-07-06 AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments Zhiheng Xi et.al. 2607.05174 null
2026-07-06 On the risk of coding before testing: An empirical study on LLM-based test generation workflow Michael Konstantinou et.al. 2607.05139 null
2026-07-06 PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference Akshat Jani et.al. 2607.05134 null
2026-07-06 When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games Jerick Shi et.al. 2607.05132 null
2026-07-06 Agent Data Injection Attacks are Realistic Threats to AI Agents Woohyuk Choi et.al. 2607.05120 null
2026-07-06 Commensal image plane transient search methods with the SKAO Alex Andersson et.al. 2607.05118 null
2026-07-06 Smooth Reduced Rank Regression with P-splines Mark de Rooij et.al. 2607.05096 null
2026-07-06 Toward Trustworthy Large Language Model Agents in Healthcare Hadi Hasan et.al. 2607.05055 null
2026-07-06 Your Agent’s Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses Neeraj Karamchandani et.al. 2607.05029 null
2026-07-06 TACTIC-KG: Toward Small Agent Teams for Cyber Threat Intelligence Knowledge Graph Construction Mouhamed Amine Bouchiha et.al. 2607.05001 null
2026-07-06 STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Qiuyi Qi et.al. 2607.04963 null
2026-07-02 Distributed Attacks in Persistent-State AI Control Josh Hills et.al. 2607.02514 null
2026-07-02 What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates Arman Ghaffarizadeh et.al. 2607.02507 null
2026-07-02 Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas Yuxuan Li et.al. 2607.02504 null
2026-07-02 Controllable Sim Agents with Behavior Latents Juanwu Lu et.al. 2607.02496 null
2026-07-02 Language Models as Measurement Apparatus for Culture Kent K. Chang et.al. 2607.02459 null
2026-07-02 Adoption and Ecosystem Health: A Longitudinal Analysis of Open-Source Multi-Agent Frameworks Xi Zhang et.al. 2607.02453 null
2026-07-02 EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments Zhilin Wang et.al. 2607.02440 null
2026-07-02 Steerability via constraints: a substrate for scalable oversight of coding agents Thomas Winninger et.al. 2607.02389 null
2026-07-02 HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification for Spanish Easy-to-Read Generation Lourdes Moreno et.al. 2607.02381 null
2026-07-02 Understanding Agent-Based Patching of Compiler Missed Optimizations Batu Guan et.al. 2607.02370 null
2026-07-02 Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Skill Malware Zimo Ji et.al. 2607.02357 null
2026-07-02 Coding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps Instructions Zimo Ji et.al. 2607.02294 null
2026-07-02 AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Xiangchen Cheng et.al. 2607.02255 null
2026-07-02 Copewell: A Multi-Agent Swarm Architecture for Equitable Mental Wellness Support Seren Yenikent et.al. 2607.02245 null
2026-07-02 Criticality-Based Guard Rail Validation for AI Agent Decisions in Autonomous Telecom Networks Ravi Kant Sharma et.al. 2607.02210 null
2026-07-02 UA-ChatDev: Uncertainty-Aware Multi-Agent Collaboration for Reliable Software Development Temitayo Olamilekan Ogunsusi et.al. 2607.02186 null
2026-07-02 Coding-agents can replicate scientific machine learning papers Atharva Hans et.al. 2607.02134 null
2026-07-02 ContextNest: Verifiable Context Governance for Autonomous AI Agent Misha Sulpovar et.al. 2607.02116 null
2026-07-02 Prompt Coverage Adequacy Florian Tambon et.al. 2607.02057 null
2026-07-02 PACE: A Proxy for Agentic Capability Evaluation Yueqi Song et.al. 2607.02032 null
2026-07-01 RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue Zhihao Lin et.al. 2607.01213 null
2026-07-01 Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Zhi Chen et.al. 2607.01211 null
2026-07-01 Optimal Resource Utilization for Autonomous Laboratory Orchestrators Austin McDannald et.al. 2607.01188 null
2026-07-01 Emergence of Preferential Attachment and Glass-Ceiling Effects in Autonomous Networks of LLMs Yiming Zhang et.al. 2607.01148 null
2026-07-01 Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains Changguo Jia et.al. 2607.01136 null
2026-07-01 Autonomous Scientific Discovery via Iterative Meta-Reflection Bingchen Zhao et.al. 2607.01131 null
2026-07-01 Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Ran Yan et.al. 2607.01120 null
2026-07-01 Technical Report: Asynchronous Distributed Trajectory Estimation of Multi-Robot Systems Adam Pooley et.al. 2607.01106 null
2026-07-01 Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering James C. Davis et.al. 2607.01087 null
2026-07-01 Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use Song-Lin Lv et.al. 2607.01084 null
2026-07-01 Agentic generation of verifiable rules for deterministic, self-expanding reaction classification Daniel Armstrong et.al. 2607.01061 null
2026-07-01 Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates Elias Najarro et.al. 2607.01047 null
2026-07-01 From Registry to Repository: How AI Agent Skills Are Written, Adapted, and Maintained Haoyu Gao et.al. 2607.00911 null
2026-07-01 Calibrating the Instrument: Controllability of an LLM-Driven Synthetic Population Mirko Degli Esposti et.al. 2607.00910 null
2026-07-01 Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents Ádám Kovács et.al. 2607.00895 null
2026-07-01 SessionBound: Turning Enterprise Task Approval into Budgeted Database Sessions Minmin Wu et.al. 2607.00751 null
2026-07-01 Self-GC: Self-Governing Context for Long-Horizon LLM Agents Xubin Hao et.al. 2607.00692 null
2026-07-01 AGI Maze as a Benchmark Framework for World-Modeling Agents Alexey Potapov et.al. 2607.00627 null
2026-07-01 Ai2-Kit: Streamlining AI-Accelerated Ab Initio Workflows for Complex Chemical Systems Sheng Bi et.al. 2607.00613 null
2026-07-01 Vehicle Routing Problem Meets Large Language Models: An Overview and Perspectives Xianchao Xiu et.al. 2607.00604 null
2026-06-30 QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents Sergio Hernández-Gutiérrez et.al. 2606.32034 null
2026-06-30 Generative Skill Composition for LLM Agents Xinyu Zhao et.al. 2606.32025 null
2026-06-30 TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models Shiyi Chen et.al. 2606.31976 null
2026-06-30 MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments Qingyun Liu et.al. 2606.31966 null
2026-06-30 Delegation Rights: Property, Agency, and Investment Incentives in the Age of AI Agents Yukun Zhang et.al. 2606.31935 null
2026-06-30 MVP-Nav: Multi-layer Value Map Planner Navigator Wenyuan Xie et.al. 2606.31919 null
2026-06-30 Non-classical Topological Evidence Logic Igor Sedlár et.al. 2606.31888 null
2026-06-30 An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping Renan Souza et.al. 2606.31831 null
2026-06-30 JETO-Bench: A Reproducible Benchmark for Execution Time Improvement Patches in Java Khashayar Etemadi et.al. 2606.31767 null
2026-06-30 A Conversational Agentic Interface to Physics-Based Household Digital Twins for Residential Energy Decision Support Costas Mylonas et.al. 2606.31744 null
2026-06-30 ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping Jiacheng Chen et.al. 2606.31693 null
2026-06-30 ECHO: Prune to act, trace to learn with selective turn memory in agentic RL Zijun Xie et.al. 2606.31650 null
2026-06-30 Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents Utsav Garg et.al. 2606.31648 null
2026-06-30 A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems Seyed Bagher Hashemi Natanzi et.al. 2606.31639 null
2026-06-30 A Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agents Javal Vyas et.al. 2606.31635 null
2026-06-30 FormIDEAble: Safe and Socially-aware Autonomous Systems Livia Lestingi et.al. 2606.31572 null
2026-06-30 ACE: Pluggable Adaptive Context Elasticizer across Agents Ning Liao et.al. 2606.31564 null
2026-06-30 DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation Siyu Yan et.al. 2606.31537 null
2026-06-30 Design and Implementation of Agentic Orchestrations and Orchestration of Agents Stefanie Rinderle-Ma et.al. 2606.31518 null
2026-06-30 Governance Gaps in Agent Interoperability Protocols: What MCP, A2A, and ACP Cannot Express Richard Kang et.al. 2606.31498 null
2026-06-29 Self-Evolving World Models for LLM Agent Planning Xuan Zhang et.al. 2606.30639 null
2026-06-29 GROW $^2$ : Grounding Which and Where for Robot Tool Use Yuhong Deng et.al. 2606.30632 null
2026-06-29 UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image Mohamed el amine boudjoghra et.al. 2606.30608 null
2026-06-29 MESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems Kunyang Li et.al. 2606.30602 null
2026-06-29 SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions Mohit Raghavendra et.al. 2606.30573 null
2026-06-29 Attractor States Emerge in Multi-Turn LLM Conversations Ting-Wen Ko et.al. 2606.30571 null
2026-06-29 Forensic Trajectory Signatures for Agent Memory Poisoning Detection Jun Wen Leong et.al. 2606.30566 null
2026-06-29 TraceLab: Characterizing Coding Agent Workloads for LLM Serving Kan Zhu et.al. 2606.30560 null
2026-06-29 Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing Dvir Alsheich et.al. 2606.30555 null
2026-06-29 To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks Jessica Hutchison et.al. 2606.30549 null
2026-06-29 MAS-Lab: A Specification-Driven Validation Framework for Reliable Multi-Agent Systems Jordan Augé et.al. 2606.30546 null
2026-06-29 TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech Sathvik Manikantan Napa Ugandhar et.al. 2606.30543 null
2026-06-29 Entity Binding Failures in Tool-Augmented Agents Rahul Suresh Babu et.al. 2606.30531 null
2026-06-29 Collective cooperation without individual fidelity in LLM agents Henrique Ferraz de Arruda et.al. 2606.30454 null
2026-06-29 Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents Bojie Li et.al. 2606.30383 null
2026-06-29 Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation Bertram Taetz et.al. 2606.30266 null
2026-06-29 TACO: Tool-Augmented Credit Optimization for Agentic Tool Use Mingkuan Feng et.al. 2606.30251 null
2026-06-29 DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning Xinxin Chen et.al. 2606.30189 null
2026-06-29 MirrorCode: AI can rebuild entire programs from behavior alone Tom Adamczewski et.al. 2606.30182 null
2026-06-29 On the Internet, Nobody Knows You’re an LLM Bot: Unmasking Web Agents with Multi-Layer Fingerprinting Iliana Fayolle et.al. 2606.30119 null
2026-06-29 Automating the Design of Embodied AgentArchitectures Jian Zhou et.al. 2606.30111 null
2026-06-29 LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard Binyan Xu et.al. 2606.30005 null
2026-06-29 SWE-Together: Evaluating Coding Agents in Interactive User Sessions Yifan Wu et.al. 2606.29957 null
2026-06-29 SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning Tianyu Jin et.al. 2606.29932 null
2026-06-29 AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes Anjali Rao et.al. 2606.29871 null
2026-06-29 Neural Procedural Memory: Empowering LLM Agents with Implicit Activation Steering Chengfeng Zhao et.al. 2606.29824 null
2026-06-29 Experience Graphs: The Data Foundation for Self-Improving Agents Gang Liao et.al. 2606.29823 null
2026-06-29 MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory Kuan Wang et.al. 2606.29788 null
2026-06-29 CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents Bo Qu et.al. 2606.29771 null
2026-06-29 TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging Guangyu Meng et.al. 2606.29763 null
2026-06-29 Do Recommendation Algorithms Work When Users Are LLM Agents? A Case Study on Moltbook Daming Li et.al. 2606.29762 null
2026-06-29 MicroAgent: Context-Augmented Multi-Agent Framework for Automatic Microservice Decomposition Zishan Su et.al. 2606.29742 null
2026-06-29 Attraction, Not Adaptation: How AI Agent Communities Develop Distinct Linguistic Identities Daming Li et.al. 2606.29722 null
2026-06-29 A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents Liu Zewen et.al. 2606.29719 null
2026-06-29 Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop Chenmu Zhang et.al. 2606.29717 null
2026-06-26 Agentic Hardware Design as Repository-Level Code Evolution Cunxi Yu et.al. 2606.28279 null
2026-06-26 Agent-Native Immune System: Architecture, Taxonomy, and Engineering Bo Shen et.al. 2606.28270 null
2026-06-26 Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software Daniel Russo et.al. 2606.28235 null
2026-06-26 HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration Jiaxin Li et.al. 2606.28215 null
2026-06-26 LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior Qinhong Zhou et.al. 2606.28182 null
2026-06-26 How Humans, Bots, and Agents Communicate About Vulnerabilities in Pull Requests Pien Rooijendijk et.al. 2606.28125 null
2026-06-26 ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents Shijing Hu et.al. 2606.28061 null
2026-06-26 From Detection to Action: Using LLM Agents for Fault-Tolerant Control Javal Vyas et.al. 2606.28011 null
2026-06-26 AdvancedShelLM: A Stateful Multi-Agent LLM Honeypot for SSH Deception Muris Sladić et.al. 2606.27990 null
2026-06-26 ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering ZhengXian Wu et.al. 2606.27974 null
2026-06-26 AI Persuasive Framing in Collective Dilemmas Anders Giovanni Møller et.al. 2606.27951 null
2026-06-26 It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents Yiming Sun et.al. 2606.27944 null
2026-06-26 When Multi-Robot Systems Meet Agentic AI:Towards Embodied Collective Intelligence Yuxuan Yan et.al. 2606.27929 null
2026-06-26 Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs Avni Mittal et.al. 2606.27909 null
2026-06-26 LLM Agents as Static Level-k Players in Behavioural Games Po Han Teo et.al. 2606.27845 null
2026-06-26 NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning Shiyun Zhao et.al. 2606.27826 null
2026-06-26 ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents Qitai Tan et.al. 2606.27814 null
2026-06-26 Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents Xinyuan Song et.al. 2606.27806 null
2026-06-26 GenWorld: Empirically Grounded Urban Simulation Infrastructure for Scalable LLM-Agent Studies Gen Li et.al. 2606.27650 null
2026-06-26 Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety Ting Ma et.al. 2606.27632 null
2026-06-25 CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Angela Cui et.al. 2606.27350 null
2026-06-25 Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning Tianyi Men et.al. 2606.27330 null
2026-06-25 Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search Ping Liu et.al. 2606.27291 null
2026-06-25 Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy Junhao Shi et.al. 2606.27251 null
2026-06-25 NOVA: A Verification-Aware Agent Harness for Architecture Evolution in Industrial Recommender Systems Shaohua Liu et.al. 2606.27243 null
2026-06-25 Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts Zhengyuan Liu et.al. 2606.27233 null
2026-06-25 A hardware-safety-gated system for LLM-written native ARTIQ control code on a trapped-ion platform Duanyang Wang et.al. 2606.27231 null
2026-06-25 A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO Fabiana Fournier et.al. 2606.27188 null
2026-06-25 OpenRCA 2.0: From Outcome Labels to Causal Process Supervision Aoyang Fang et.al. 2606.27154 null
2026-06-25 Joint Learning of Experiential Rules and Policies for Large Language Model Agents Shicheng Ye et.al. 2606.27136 null
2026-06-25 Mostly Automatic Translation of Language Interpreters from C to Safe Rust Bo Wang et.al. 2606.27122 null
2026-06-25 The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development Hartwig Grabowski et.al. 2606.27045 null
2026-06-25 Semantic Early-Stopping for Iterative LLM Agent Loops Sahil Shrivastava et.al. 2606.27009 null
2026-06-25 How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring Zhihao Lin et.al. 2606.26979 null
2026-06-25 Toward Agentic SysAdmin: Rethinking System Administration with AI Agents Gianmaria Frigo et.al. 2606.26960 null
2026-06-25 A Deterministic Control Plane for LLM Coding Agents Padmaraj Madatha et.al. 2606.26924 null
2026-06-25 Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Zekai Zhang et.al. 2606.26907 null
2026-06-25 EconSimulacra: A Digital Twin Platform of Socio-Economic Systems Powered by LLM Agents Ryuji Hashimoto et.al. 2606.26883 null
2026-06-25 EGG: An Expert-Guided Agent Framework for Kernel Generation Yaochen Han et.al. 2606.26758 null
2026-06-25 Knowledge-Based Pull Requests: A Trusted Workflow for Agent-Mediated Knowledge Collaboration Xinyu Zhang et.al. 2606.26721 null
2026-06-24 Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Changdae Oh et.al. 2606.26080 null
2026-06-24 The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems Seth Dobrin et.al. 2606.26057 null
2026-06-24 Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem Xihan Xiong et.al. 2606.26028 null
2026-06-24 Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It Yupu Hao et.al. 2606.26027 null
2026-06-24 Autodata: An agentic data scientist to create high quality synthetic data Ilia Kulikov et.al. 2606.25996 null
2026-06-24 Explainable Control Framework (XCF) based on Fuzzy Model-Agnostic Explanation and LLM Agent-Supported Interface Faliang Yin et.al. 2606.25941 null
2026-06-24 Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs Adeeb Zaman et.al. 2606.25899 null
2026-06-24 The Web4 Agent Economy: A Large-Scale Empirical Study of the Landscape, Challenges, and Opportunities Yuhan Jin et.al. 2606.25876 null
2026-06-24 Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents Peng Xu et.al. 2606.25852 null
2026-06-24 AI Snitches Get Glitches: Towards Evading Agentic Surveillance Hyejun Jeong et.al. 2606.25836 null
2026-06-24 Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability Yang Tian et.al. 2606.25819 null
2026-06-24 GUI agent: Guided Exploration of User-Sensitive Screens Aradhana Nayak et.al. 2606.25705 null
2026-06-24 Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints Fangzheng Li et.al. 2606.25605 null
2026-06-24 IntentTester: Intent-Driven Multi-agent Framework for Cross-Library Test Migration Yi Gao et.al. 2606.25588 null
2026-06-24 BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Hanyang Wang et.al. 2606.25556 null
2026-06-24 Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models Xinyu Lian et.al. 2606.25519 null
2026-06-24 SAGE-Nav: Leveraging LLM Planning and Alignment Fusion for Hierarchical Scene Graph-Guided Navigation Hao Su et.al. 2606.25497 null
2026-06-24 The Interplay of Harness Design and Post-Training in LLM Agents Kyungmin Kim et.al. 2606.25447 null
2026-06-24 Beyond Next-Observation Prediction: Agent-Authored World Modeling for Sequential Decision Making Guangfeng Cai et.al. 2606.25421 null
2026-06-24 BrainAgent: A Large Language Model-Driven Multi-Agent Framework for Autonomous Brain Signal Understanding Yangxuan Zhou et.al. 2606.25400 null
2026-06-23 SHERLOC: Structured Diagnostic Localization for Code Repair Agents Hovhannes Tamoyan et.al. 2606.24820 null
2026-06-23 MANGO: Automated Multi-Agent Test Oracle Generation for Vision-Language-Action Models Pablo Valle et.al. 2606.24815 null
2026-06-23 Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce Filippos Ventirozos et.al. 2606.24783 null
2026-06-23 DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects Shiyu Li et.al. 2606.24779 null
2026-06-23 Are We Ready For An Agent-Native Memory System? Wei Zhou et.al. 2606.24775 null
2026-06-23 SupplyNet: Supporting Visual Exploratory Learning in Supply Chain via Contextual Multi-Agent Simulation Yanjia Li et.al. 2606.24694 null
2026-06-23 Automated Summarization of Software Documents: An LLM-based Multi-Agent Approach Duc S. H. Nguyen et.al. 2606.24689 null
2026-06-23 Agentic Collaborative Cognition for Zero-Shot 3D Understanding Wenxin Wang et.al. 2606.24649 null
2026-06-23 SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation Chenyang Zhu et.al. 2606.24626 null
2026-06-23 Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity Yuanhe Zhao et.al. 2606.24623 null
2026-06-23 Degeneracy-Aware Resilient Resource Allocation in Cell-Free Cache-Aided MU-MIMO Networks Sayanti Ghosh et.al. 2606.24611 null
2026-06-23 CONDUCTOR: An LLM-Orchestrated Digital Twin for Uncertainty-Aware Distribution Grid Operations Antonio Alcántara et.al. 2606.24609 null
2026-06-23 Qwen-AgentWorld: Language World Models for General Agents Yuxin Zuo et.al. 2606.24597 null
2026-06-23 MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery Enze Ma et.al. 2606.24595 null
2026-06-23 AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability Khanak Khandelwal et.al. 2606.24589 null
2026-06-23 NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Yuru Wang et.al. 2606.24530 null
2026-06-23 Bayesian control for coding agents Theodore Papamarkou et.al. 2606.24453 null
2026-06-23 Agentic Generation of AST Transformation Rules for Fixing Breaking Updates Frank Reyes et.al. 2606.24446 null
2026-06-23 ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling Heng Ping et.al. 2606.24437 null
2026-06-23 Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories Arsham Khosravani et.al. 2606.24429 null
2026-06-22 Semantic Browsing: Controllable Diversity for Image Generation Sara Dorfman et.al. 2606.23679 null
2026-06-22 AIR: Adaptive Interleaved Reasoning with Code in MLLMs Cong Han et.al. 2606.23678 null
2026-06-22 HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory Xiaolin Zhou et.al. 2606.23565 null
2026-06-22 Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference Yuhang Gan et.al. 2606.23521 null
2026-06-22 Continuity equations in the Generalised Lagrangian Mean theory Vladimir A. Vladimirov et.al. 2606.23481 null
2026-06-22 AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction Shanhui Zhao et.al. 2606.23449 null
2026-06-22 Detecting Malicious Agent Skills in the Wild using Attention Bacem Etteib et.al. 2606.23416 null
2026-06-22 Superhuman AI for Generals.io Using Self-Play Reinforcement Learning Matej Straka et.al. 2606.23348 null
2026-06-22 Group Selection Promotes Prosocial Prompts in Populations of LLM Agents Luis Celiktemel et.al. 2606.23343 null
2026-06-22 VideoAgent: All-in-One Framework for Video Understanding and Editing Hengji Zhou et.al. 2606.23327 null
2026-06-22 Test-Driven, AI-Assisted Learning: Replacing Lectures with Weekly Closed-Book Tests Jin-Guo Liu et.al. 2606.23315 null
2026-06-22 IOI: Decoupling Kinematics and Physics for Interactive World Models Chengyu Bai et.al. 2606.23296 null
2026-06-22 GIF: Locally Sound Geometric Information Flow Control for LLMs Adam Storek et.al. 2606.23277 null
2026-06-22 Wireless Personal Agent: Extending Wireless Intelligence from Networks to Terminals Jiedan Tan et.al. 2606.23255 null
2026-06-22 RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation Feifei Bian et.al. 2606.23221 null
2026-06-22 MuPPET: A Benchmark for Contextual Privacy of LLM Assistants in Multi-Party Conversations Elena Sofia Ruzzetti et.al. 2606.23217 null
2026-06-22 Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory Zewen Liu et.al. 2606.23195 null
2026-06-22 Position: Correct Answer, Wrong Mechanism – When AI Scientists Defend General Claims Their Own Data Contradicts Steven Young Eulig et.al. 2606.23175 null
2026-06-22 Understanding the (In)Security of Vibe-Coded Applications Junquan Deng et.al. 2606.23130 null
2026-06-22 Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation Julia Belikova et.al. 2606.23127 null
2026-06-21 MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop Yikun Fu et.al. 2606.22557 null
2026-06-21 Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents Shiyang Chen et.al. 2606.22528 null
2026-06-21 Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents Igor Santos-Grueiro et.al. 2606.22504 null
2026-06-21 Governed AI-Assisted Engineering: Graduated Human Oversight for Agentic Code Generation in Regulated Domains Richard Kang et.al. 2606.22484 null
2026-06-21 A Differentiable Atari VCS:A Complex, Fully Known Ground Truth for Explainable AI Andreas Maier et.al. 2606.22447 null
2026-06-21 SVGym (SciVerseGym): An Environment for Reinforcement Learning and Bayesian Optimization in Crystal Discovery Bin Cao et.al. 2606.22425 null
2026-06-21 Code Isn’t Memory: A Structural Codebase Index Inside a Coding Agent Ishaan Bhola et.al. 2606.22417 null
2026-06-21 PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Jiayu Liu et.al. 2606.22388 null
2026-06-21 MetaPS: Adaptive Programmatic Strategy Selection for Market Agents Jiaxiang Chen et.al. 2606.22385 null
2026-06-21 Hypothesis-Driven Skill Optimization for LLM Agents Fangxin Shang et.al. 2606.22330 null
2026-06-20 Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases Yiwei Hou et.al. 2606.22263 null
2026-06-20 Quantifying Theoretical AI Alignment Guarantees: Receiver-Utility Bounds in Bayesian Persuasion Eric Yachbes et.al. 2606.22226 null
2026-06-20 When Is Emergent Consensus Real? A Measured Coupling Gain and a Validity Diagnostic for LLM Agent Societies Dongxu Yang et.al. 2606.22203 null
2026-06-20 TraceView: Interactive Visualization of Agentic Program Repair Trajectories Amirali Sajadi et.al. 2606.22110 null
2026-06-20 CodeTeam: An LLM-Powered Multi-Agent Framework for Repository-Level Code Generation Yifei Wang et.al. 2606.22082 null
2026-06-20 Skills for the future software profession: beyond agentic AI! Sungmin Kang et.al. 2606.21894 null
2026-06-20 Learning the ARTS of Search for Automated Discovery Gurusha Juneja et.al. 2606.21891 null
2026-06-20 AgentRiskBOM: A Risk-Scoping Security Bill of Materials for Agentic AI Systems Srimonti Dutta et.al. 2606.21877 null
2026-06-20 Harness-MU: A Safe, Governed, and Effective Harness for Multi-User LLM Agents Wangxuan Fan et.al. 2606.21856 null
2026-06-20 Measuring What Persists: Conditioning Mechanisms and a Geometric Framework for AI Agent Identity Andrew Tanner et.al. 2606.21843 null
2026-06-18 Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving Liang Su et.al. 2606.20537 null
2026-06-18 Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes Jun He et.al. 2606.20520 null
2026-06-18 S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Yalun Dai et.al. 2606.20515 null
2026-06-18 Probe-and-Refine Tuning of Repository Guidance for Coding Agents Asa Shepard et.al. 2606.20512 null
2026-06-18 Efficient and Sound Probabilistic Verification for AI Agents Alaia Solko-Breslin et.al. 2606.20510 null
2026-06-18 Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems Zewen Liu et.al. 2606.20493 null
2026-06-18 LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems Hanwool Lee et.al. 2606.20408 null
2026-06-18 AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning Zepeng Li et.al. 2606.20373 null
2026-06-18 AgenticDB: Agentic Performance Reconfiguration for Database Workloads Xinyue Yang et.al. 2606.20318 null
2026-06-18 Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference Huang Peng et.al. 2606.20245 null
2026-06-18 A Multi-Agent system for Multi-Objective constrained optimization Federica Filippini et.al. 2606.20236 null
2026-06-18 N-Version Programming with Coding Agents Javier Ron et.al. 2606.20158 null
2026-06-18 Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform Hyeonna Choi et.al. 2606.20120 null
2026-06-18 When Does Streaming Tool Use Help? Characterizing Tool-Intent Stabilization in Streaming Retrieval-Augmented Generation Elroy Galbraith et.al. 2606.20113 null
2026-06-18 PACMS: Submodular Context Selection as a Pluggable Engine for LLM Agents Manu Ghulyani et.al. 2606.20047 null
2026-06-18 See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View Fanfu Xue et.al. 2606.20045 null
2026-06-18 AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models Masahiro Kato et.al. 2606.20041 null
2026-06-18 When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents Kaiyue Yang et.al. 2606.20023 null
2026-06-18 Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning Yanxi Chen et.al. 2606.20002 null
2026-06-18 ENPIRE: Agentic Robot Policy Self-Improvement in the Real World Wenli Xiao et.al. 2606.19980 null
2026-06-17 Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning Jisoo Kim et.al. 2606.19340 null
2026-06-17 Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents Anoushka Vyas et.al. 2606.19319 null
2026-06-17 Accelerating Network-Agent Dispersion: Territorial Behavior and Directionally Biased Lazy Random Walks Li Zeng et.al. 2606.19294 null
2026-06-17 TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology Hannah Le et.al. 2606.19245 null
2026-06-17 Runtime Compliance Verification for AI Agents Nafiseh Kahani et.al. 2606.19242 null
2026-06-17 STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability Haipeng Luo et.al. 2606.19236 null
2026-06-17 CodeSentinel: A Three-Layer Defense Against Indirect Prompt Injection in Code Contexts Po-Han Cheng et.al. 2606.19235 null
2026-06-17 PhantomSkill: Malicious Code Injection in Agent Skill Ecosystems Yu-Ting Lin et.al. 2606.19191 null
2026-06-17 AdsMind: A Physics-Grounded Multi-Agent System for Self-Correcting Discovery of Adsorption Configurations on Heterogeneous Catalyst Surfaces Zongmin Zhang et.al. 2606.19152 null
2026-06-17 A Technical Taxonomy of LLM Agent Communication Protocols Linus Sander et.al. 2606.19135 null
2026-06-17 Towards an Agent-First Web: Redesigning the Web for AI Agents Eranga Bandara et.al. 2606.19116 null
2026-06-17 PYPILINE: Malicious PyPI Package Detection via Suspicious API Knowledge and Agent Workflow Siyuan Pang et.al. 2606.19063 null
2026-06-17 RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents Ruishan Fang et.al. 2606.19047 null
2026-06-17 TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction Moon Ye-Bin et.al. 2606.18996 null
2026-06-17 Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents Emmanuel Aboah Boateng et.al. 2606.18947 null
2026-06-17 Generative-Model Predictive Planning for Navigation in Partially Observable Environments Thomas Quilter et.al. 2606.18888 null
2026-06-17 WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents Yehang Zhang et.al. 2606.18847 null
2026-06-17 Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning Xiaoyue Xu et.al. 2606.18831 null
2026-06-17 GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents Zhe Ren et.al. 2606.18829 null
2026-06-17 ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch Tengfei Lyu et.al. 2606.18803 null
2026-06-16 Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion Nils Morbitzer et.al. 2606.18250 null
2026-06-16 ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues Shanda Li et.al. 2606.18237 null
2026-06-16 EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation Qi Chai et.al. 2606.18235 null
2026-06-16 All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code Dipayan Banik et.al. 2606.18168 null
2026-06-16 Learning Cardiac Electrophysiology Digital Twins Through Agentic Discovery of Hybrid Structure Ziqi Zhou et.al. 2606.18154 null
2026-06-16 WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning Yuwei Zhang et.al. 2606.18147 null
2026-06-16 Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So Josef Liyanjun Chen et.al. 2606.18144 null
2026-06-16 Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models Jasmine Brazilek et.al. 2606.18142 null
2026-06-16 On the Reliability of Networks of AI Agents: Density Evolution, Stopping Sets, and Architecture Optimization Ehsan Aghazadeh et.al. 2606.18121 null
2026-06-16 Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications Divyansh Srivastava et.al. 2606.18068 null
2026-06-16 Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose Xueping Gao et.al. 2606.18051 null
2026-06-16 ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents Ander Alvarez et.al. 2606.18037 null
2026-06-16 LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling Jian Yang et.al. 2606.18023 null
2026-06-16 LLM Consumer Behavior Theory: Foundations of a Novel Research Field Manon Reusens et.al. 2606.18005 null
2026-06-16 How Inference Compute Shapes Frontier LLM Evaluation Jessica McFadyen et.al. 2606.17930 null
2026-06-16 Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization Aueaphum Aueawatthanaphisut et.al. 2606.17915 null
2026-06-16 GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine? Tongxu Luo et.al. 2606.17861 null
2026-06-16 From Ad Hoc Pilots to Repeatable Patterns: Structuring Drone Collaboration in Emergency Services with DroneLets Dzmitry Katsiuba et.al. 2606.17839 null
2026-06-16 Environment-Grounded Automated Prompt Optimization for LLM Game Agents Rean Clive Fernandes et.al. 2606.17838 null
2026-06-16 A Framework for Evaluating Agentic Skills at Scale Maksim Shaposhnikov et.al. 2606.17819 null
2026-06-15 Context-Aware RL for Agentic and Multimodal LLMs Peiyang Xu et.al. 2606.17053 null
2026-06-15 Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio Anzhe Xie et.al. 2606.17041 null
2026-06-15 TokenPilot: Cache-Efficient Context Management for LLM Agents Buqiang Xu et.al. 2606.17016 null
2026-06-15 Agent trajectories as programs: fingerprinting and programming coding-agent behavior Hamidah Oderinwale et.al. 2606.16988 null
2026-06-15 Directory-Aware Query and Maintenance in Vector Databases Mengzhao Wang et.al. 2606.16903 null
2026-06-15 Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization Dongbin Na et.al. 2606.16898 null
2026-06-15 Human-on-the-Bridge: Scalable Evaluation for AI Agents Fouad Bousetouane et.al. 2606.16871 null
2026-06-15 Evolution & Foundation: AI Shares Creative Control Dylan Banarse et.al. 2606.16849 null
2026-06-15 Towards LLM Accelerated Rapid Reviews for Software Tool Discovery – Case for Log Anomaly Detection Jesse Nyyssölä et.al. 2606.16839 null
2026-06-15 CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agents Shubham Tiwari et.al. 2606.16824 null
2026-06-15 GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents Rahul Suresh Babu et.al. 2606.16813 null
2026-06-15 LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control Anqi Zou et.al. 2606.16802 null
2026-06-15 OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models Tianyi Lin et.al. 2606.16774 null
2026-06-15 Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents Tianyi Zhang et.al. 2606.16769 null
2026-06-15 Witnesses and Counterexamples for Timed Bisimulation Alexander Lieb et.al. 2606.16736 null
2026-06-15 A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions Jianghan Shen et.al. 2606.16733 null
2026-06-15 AgentFairBench: Do LLM Agents Discriminate When They Act? Triveni Morla et.al. 2606.16723 null
2026-06-15 Misinformation Propagation in Benign Multi-Agent Systems Jonas Becker et.al. 2606.16710 null
2026-06-15 User as Code: Executable Memory for Personalized Agents Bojie Li et.al. 2606.16707 null
2026-06-15 Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents Zewen Liu et.al. 2606.16682 null
2026-06-12 AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition Jixuan Chen et.al. 2606.14674 null
2026-06-12 Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows Shikun Liu et.al. 2606.14672 null
2026-06-12 Towards In Silico Cancer Therapy Design: An Agent-Based Approach for GPU-Accelerated Molecular Pathway Simulation Stefano Maestri et.al. 2606.14603 null
2026-06-12 When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime Wei Wu et.al. 2606.14589 null
2026-06-12 SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model Xiaoxin Lu et.al. 2606.14574 null
2026-06-12 From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails Yuguang Zhou et.al. 2606.14517 null
2026-06-12 From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Yongheng Zhang et.al. 2606.14502 null
2026-06-12 When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More Zhongyuan Wang et.al. 2606.14476 null
2026-06-12 GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge Pavan C Shekar et.al. 2606.14470 null
2026-06-12 tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration Minseo Kim et.al. 2606.14445 null
2026-06-12 Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments Mykola Vysotskyi et.al. 2606.14397 null
2026-06-12 Communication Policy Evolution for Proactive LLM Agents Xinbei Ma et.al. 2606.14314 null
2026-06-12 Retrospective Progress-Aware Self-Refinement for LLM Agent Training Xinbei Ma et.al. 2606.14302 null
2026-06-12 Large Language Model Based Agent for Automated Discovery in Computational Physics Hang Lin et.al. 2606.14266 null
2026-06-12 Security in a Workflow: Exploring Role-Based Agentic Architectures for Vulnerability Handling Srijita Basu et.al. 2606.14261 null
2026-06-12 HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry Tingyang Chen et.al. 2606.14249 null
2026-06-12 SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing Haowen Gao et.al. 2606.14239 null
2026-06-12 Selective Agentic Recovery for UAV Autonomy with a Persistent Mission Runtime Taewoo Park et.al. 2606.14219 null
2026-06-12 Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL Yinglun Zhu et.al. 2606.14211 null
2026-06-12 When Should Agent Trust Be Conditional? Characterizing and Attacking Skill-Conditional Reputation in Agent Swarms Yihan Xia et.al. 2606.14200 null
2026-06-11 EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments Jundong Xu et.al. 2606.13681 null
2026-06-11 Mana: Dexterous Manipulation of Articulated Tools Zhao-Heng Yin et.al. 2606.13677 null
2026-06-11 HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents Yaxin Du et.al. 2606.13663 null
2026-06-11 EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Amy Xin et.al. 2606.13662 null
2026-06-11 Recursive Agent Harnesses Elias Lumer et.al. 2606.13643 null
2026-06-11 AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility Xiaoyuan Liu et.al. 2606.13608 null
2026-06-11 EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis Harihara Muralidharan et.al. 2606.13602 null
2026-06-11 ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages Tanmoy Kanti Halder et.al. 2606.13572 null
2026-06-11 A Reactive Redistribution Mechanism for STL Tasks in Multi-Agent Systems Under Time-Varying Communication Gregorio Marchesini et.al. 2606.13479 null
2026-06-11 Exploring Systems-Thinking Approaches to Loss of Control Risk Aurelio Carlucci et.al. 2606.13474 null
2026-06-11 Understanding the Rejection of Fixes Generated by Agentic Pull Requests – Insights from the AIDev Dataset Mahmoud Abujadallah et.al. 2606.13468 null
2026-06-11 From Traditional Automation to Embodied Wireless Intelligence: Vision-Language-Action Empowered Physics-Aware Communication Networks Genze Jiang et.al. 2606.13458 null
2026-06-11 Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull Requests Ali Arabat et.al. 2606.13449 null
2026-06-11 MiniMax Sparse Attention Xunhao Lai et.al. 2606.13392 null
2026-06-11 Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents Zihao Wang et.al. 2606.13385 null
2026-06-11 An LLM System for Autonomous Variational Quantum Circuit Design Kenya Sakka et.al. 2606.13380 null
2026-06-11 IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing Tao Hu et.al. 2606.13368 null
2026-06-11 Can I Buy Your KV Cache? Luoyuan Zhang et.al. 2606.13361 null
2026-06-11 SkillCAT: Contrastive Assessment and Topology-Aware Skill Self-Evolution for LLM Agents Kunfeng Chen et.al. 2606.13317 null
2026-06-11 Under What Conditions Can a Machine Become Genuinely Creative? Yong Zeng et.al. 2606.13196 null
2026-06-10 DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners? Jadelynn Dao et.al. 2606.12402 null
2026-06-10 APPO: Agentic Procedural Policy Optimization Xucong Wang et.al. 2606.12384 null
2026-06-10 Should LLM Agents Decide in Social Simulations? Comparing Finite-State and LLM-Based Decision Policies Alejandro Buitrago López et.al. 2606.12369 null
2026-06-10 Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks Mengyu Zheng et.al. 2606.12344 null
2026-06-10 OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents Jin Xie et.al. 2606.12341 null
2026-06-10 PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents Ripon Chandra Malo et.al. 2606.12329 null
2026-06-10 A Five-Plane Reference Architecture for Runtime Governance of Production AI Agents Krti Tallam et.al. 2606.12320 null
2026-06-10 The Impossibility of Eliciting Latent Knowledge Korbinian Friedl et.al. 2606.12268 null
2026-06-10 DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems Zhongyu Xia et.al. 2606.12236 null
2026-06-10 InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Ziang Yan et.al. 2606.12195 null
2026-06-10 From Agent Identity to Agent Economy: Measuring the Operational Readiness of ERC-8004 AI Agents Rischan Mafrur et.al. 2606.12128 null
2026-06-10 ChargeBD: Character-Aware Heterogeneous Agent Reasoning for Guided Engineering in Battery Development Rui Huang et.al. 2606.12057 null
2026-06-10 A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design Wanting Wang et.al. 2606.12040 null
2026-06-10 MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning Shang Ma et.al. 2606.12018 null
2026-06-10 Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents Frank Xiao et.al. 2606.11998 null
2026-06-10 ParseFixer: An Agentic Framework for Document Parsing via Selective Multimodal Correction LeKai Yu et.al. 2606.11977 null
2026-06-10 Exploration Structure in LLM Agents for Multi-File Change Localization Akeela Darryl Fattha et.al. 2606.11976 null
2026-06-10 Toward Generalist Autonomous Research via Hypothesis-Tree Refinement Jiajie Jin et.al. 2606.11926 null
2026-06-10 Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction Baoyang Jiang et.al. 2606.11909 null
2026-06-10 When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models Xuan Dong et.al. 2606.11906 null
2026-06-09 EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents Weixian Xu et.al. 2606.11182 null
2026-06-09 Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories Kevin Qinghong Lin et.al. 2606.11176 null
2026-06-09 ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity Andrew Bo Liu et.al. 2606.11150 null
2026-06-09 OpenPCC: Open and Confidential LLM Serving on Commodity TEEs Haoling Zhou et.al. 2606.11145 null
2026-06-09 TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Heming Zou et.al. 2606.11119 null
2026-06-09 Revealing information – or not – in a social network of traders Patrick Allmis et.al. 2606.11053 null
2026-06-09 LLM-Mediated Demand Response Coordination in Smart Microgrids J. de Curtò et.al. 2606.11050 null
2026-06-09 Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Liya Zhu et.al. 2606.11042 null
2026-06-09 Understanding and mitigating the risks of OpenClaw for non-technical users: A practical guide with Skill Junchang Zheng et.al. 2606.11007 null
2026-06-09 Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam? Tengchao Lv et.al. 2606.10956 null
2026-06-09 Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages Aman Sharma et.al. 2606.10933 null
2026-06-09 Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution Xucong Wang et.al. 2606.10917 null
2026-06-09 Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation Yupu Hao et.al. 2606.10875 null
2026-06-09 Moonshine: An Autonomous Mathematical Research Agent Centered on Conjecture Generation Xiaoyang Chen et.al. 2606.10806 null
2026-06-09 Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use Zhixin Ma et.al. 2606.10803 null
2026-06-09 Evaluating Research-Level Math Proofs via Strict Step-Level Verification Yifeng Sun et.al. 2606.10799 null
2026-06-09 AutoPDE: Reliable Agentic PDE Solving via Explicitly Represented Solver Strategies Huanshuo Dong et.al. 2606.10752 null
2026-06-09 Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation Yuchen Ling et.al. 2606.10749 null
2026-06-09 MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents Yv Zhang et.al. 2606.10742 null
2026-06-09 DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch Jiale Zhao et.al. 2606.10728 null
2026-06-08 OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics Mingxian Lin et.al. 2606.09826 null
2026-06-08 Limit Theory for $N$-Player $α$ -Potential Games Xin Guo et.al. 2606.09815 null
2026-06-08 iMaC: Translating Actions into Motion and Contact Images for Embodied World Models Zhenyu Wu et.al. 2606.09813 null
2026-06-08 FASE: Fast Adaptive Semantic Entropy for Code Quality Shizhe Lin et.al. 2606.09800 null
2026-06-08 SIGA: Self-Evolving Coding-Agent Adapters for Scientific Simulation Matthew Ho et.al. 2606.09774 null
2026-06-08 Collaborative Human-Agent Protocol (CHAP) Arsalan Shahid et.al. 2606.09751 null
2026-06-08 HDSL: A Hierarchical Domain-Specific Language for Structured 3D Indoor Scene Generation and Localized Editing with LLM Agents Letian Li et.al. 2606.09738 null
2026-06-08 SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research Pu Ning et.al. 2606.09730 null
2026-06-08 IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking Zechen Sun et.al. 2606.09709 null
2026-06-08 (Auto)formalization is supposed to be easy: Trellis process semantics for spelling out rigorous proofs Wesley Pegden et.al. 2606.09674 null
2026-06-08 MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding Jie Zhang et.al. 2606.09641 null
2026-06-08 Agentic Persona Generation with Critique-Refinement: An Industrial Evaluation Mohammad Hossein Amini et.al. 2606.09637 null
2026-06-08 AGENTSERVESIM: A Hardware-aware Simulator for Multi-Turn LLM Agent Serving Rakibul Hasan Rajib et.al. 2606.09613 null
2026-06-08 InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees Shaoyang Cui et.al. 2606.09550 null
2026-06-08 SecureClaw: Clawing Back Control of LLM Agents Yuhan Ma et.al. 2606.09549 null
2026-06-08 Memory Beyond Recall: A Dual-Process Cognitive Memory System for Self-Evolving LLM Agents Tianxiang Fei et.al. 2606.09483 null
2026-06-08 H2HMem: A Multimodal Memory Benchmark for Agents in Human-Human Interactions Shiping Zhu et.al. 2606.09461 null
2026-06-08 AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning Bojie Rong et.al. 2606.09447 null
2026-06-08 A Robust Agentic Framework for Expert-Level Automation of Atomistic Simulations Yutack Park et.al. 2606.09422 null
2026-06-08 What Should a Skill Remember? Quality-Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents Qinghua Xing et.al. 2606.09421 null
2026-06-08 RunAgent SuperBrowser: A Theory of Autonomous Web Navigation Grounded in Human Browsing Behaviour Radeen Mostafa et.al. 2606.09399 null
2026-06-08 Capability-Aligned Hierarchical Learning for Tool-Augmented LLMs Haotong Yang et.al. 2606.09371 null
2026-06-08 Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory Haoran Sun et.al. 2606.09365 null
2026-06-08 Bespoke-Card: Why Tune When You Can Generate? Synthesizing Workload-Specific Cardinality Estimators Johannes Wehrstein et.al. 2606.09361 null
2026-06-08 Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents Jianwei Tai et.al. 2606.09315 null
2026-06-08 Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning Haoran Xu et.al. 2606.09290 null
2026-06-08 MAGIS: Evidence-Based Multi-Agent Reasoning for Interpretable Strabismus Clinical Decision-Making Xikai Tang et.al. 2606.09249 null
2026-06-08 Self-Paced Curriculum Reinforcement Learning for Autonomous Superbike Racing in Simulation Luca Ghisi et.al. 2606.09236 null
2026-06-08 Trustworthy Smart Fabs via Professional Proxies: Scaling Safe and Sustainable by Design (SSbD) through Industrial Data Spaces Han-Teng Liao et.al. 2606.09227 null
2026-06-05 Accelerated Decentralized Stochastic Gradient Descent for Strongly Convex Optimization Ming Sun et.al. 2606.07496 null
2026-06-05 How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope Jeremy Yang et.al. 2606.07489 null
2026-06-05 OPENPATH: A Supervisor–Specialist Agent System for Personalized, Accessible, and Multi-stop Urban Trip Planning Ziyang Xiong et.al. 2606.07486 null
2026-06-05 Agentic Very Much! Adoption of Coding Agent in New GitHub Projects Romain Robbes et.al. 2606.07448 null
2026-06-05 Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning Haoyuan Li et.al. 2606.07436 null
2026-06-05 Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills Chuan Xiao et.al. 2606.07412 null
2026-06-05 Audio-Oscar: A Multi-Agent System for Complex Audio Scene Generation, Orchestration, and Refinement Yifan Duan et.al. 2606.07397 null
2026-06-05 Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests Thanawat Lodkaew et.al. 2606.07379 null
2026-06-05 Self-evolving LLM agents with in-distribution Optimization Yudi Zhang et.al. 2606.07367 null
2026-06-05 Hierarchical Certified Semantic Commitment for Byzantine-Resilient LLM-Agent Collaboration Haoran Xu et.al. 2606.07316 null
2026-06-05 QBugLM: An Agentic Benchmarking Framework for LLM-based Quantum Software Debugging An B. B. Pham et.al. 2606.07314 null
2026-06-05 CAPE: Contrastive Action-conditioned Parallel Encoding for Embodied Planning Cong Chen et.al. 2606.07304 null
2026-06-05 SWE-Explore: Benchmarking How Coding Agents Explore Repositories Shaoqiu Zhang et.al. 2606.07297 null
2026-06-05 MMAE: A Massive Multitask Audio Editing Benchmark Ziyang Ma et.al. 2606.07229 null
2026-06-05 Learning Multi-Agent Communication Protocol: Study on Information Entropy Efficiency in MARL Xinren Zhang et.al. 2606.07200 null
2026-06-05 From Privacy to Workflow Integrity: Communication-Graph Metadata in Autonomous Agent Interoperability Bijaya Dangol et.al. 2606.07150 null
2026-06-05 MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills Wenbo Guo et.al. 2606.07131 null
2026-06-05 SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating Zequn Xie et.al. 2606.07074 null
2026-06-05 TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents Vijitha Mittapalli et.al. 2606.07054 null
2026-06-05 Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning Yijin Zhou et.al. 2606.06976 null
2026-06-04 Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators Chenming Zhu et.al. 2606.06476 null
2026-06-04 MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery Shangheng Du et.al. 2606.06473 null
2026-06-04 Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement Jui-Hui Chung et.al. 2606.06468 null
2026-06-04 Benchmark Everything Everywhere All at Once Shiyun Xiong et.al. 2606.06462 null
2026-06-04 Will the Agent Recuse Itself? Measuring LLM-Agent Compliance with In-Band Access-Deny Signals Thamilvendhan Munirathinam et.al. 2606.06460 null
2026-06-04 Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Zhuoming Chen et.al. 2606.06453 null
2026-06-04 Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads Yasmine Omri et.al. 2606.06448 null
2026-06-04 CollabSim: A CSCW-Grounded Methodology for Investigating Collaborative Competence of LLM Agents through Controlled Multi-Agent Experiments Jiaju Chen et.al. 2606.06399 null
2026-06-04 Humans’ ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration Jiaju Chen et.al. 2606.06388 null
2026-06-04 WebMCP Tool Surface Poisoning: Runtime Manipulation Attacks on LLM Agents Lin-Fa Lee et.al. 2606.06387 null
2026-06-04 StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset Zhengqian Wu et.al. 2606.06338 null
2026-06-04 From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws Mengzhuo Chen et.al. 2606.06324 null
2026-06-04 ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents Rahul Suresh Babu et.al. 2606.06284 null
2026-06-04 DAST: A VLM-LLM Framework for Cross-Interface Anomaly Detection in O-RAN Francesco Spinelli et.al. 2606.06261 null
2026-06-04 TOKI: A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory Ziming Wang et.al. 2606.06240 null
2026-06-04 From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents Patrick Wilhelm et.al. 2606.06223 null
2026-06-04 Learning to Contest: Decentralized Robust Fairness in Cooperative MARL via Cross-Attention Can Savcı et.al. 2606.06162 null
2026-06-04 A Finite Certificate for the Positive $n=9$ Vasc Inequality Dakai Guo et.al. 2606.06136 null
2026-06-04 LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents Aofan Yu et.al. 2606.06087 null
2026-06-04 SkillComposer: Learning to Evolve Agent Skills for Specification and Generalization Qi Zhang et.al. 2606.06079 null
2026-05-29 EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision Rosario Forte et.al. 2605.31557 null
2026-05-29 Accessing Exotic Hadronic States via Charmed-Meson Femtoscopy in Relativistic Heavy-Ion Collisions Jiaxing Zhao et.al. 2605.31527 null
2026-05-29 If LLMs Have Human-Like Attributes, Then So Does Age of Empires II Adrian de Wynter et.al. 2605.31514 null
2026-05-29 Skill Reuse as Compression in Agentic RL Zhikun Xu et.al. 2605.31509 null
2026-05-29 GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization Zaid Khan et.al. 2605.31464 null
2026-05-29 PithTrain: A Compact and Agent-Native MoE Training System Ruihang Lai et.al. 2605.31463 null
2026-05-29 Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information Antonio Valerio Miceli-Barone et.al. 2605.31445 null
2026-05-29 Answer-Set-Programming-based Abstractions for Reinforcement Learning Rafael Bankosegger et.al. 2605.31444 null
2026-05-29 DynaTree: Dynamic Agentic Retrieval Tree for Time-Sensitive News Retrieval Siyuan Qi et.al. 2605.31377 null
2026-05-29 HypoAgent: An Agentic Framework for Interactive Abductive Hypothesis Generation over Knowledge Graphs Yisen Gao et.al. 2605.31370 null
2026-05-29 Learning to Adapt: Self-Improving Web Agent via Cognitive-Aware Exploration Weile Chen et.al. 2605.31365 null
2026-05-29 Personalized to Persuade: The Effects of Contextualization and Warmth on Trust and Reliance in Conversational AI Mert Yazan et.al. 2605.31275 null
2026-05-29 Mellum2 Technical Report Marko Kojic et.al. 2605.31268 null
2026-05-29 COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Tianyi Zhou et.al. 2605.31264 null
2026-05-29 ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models Kaiwen Xue et.al. 2605.31251 null
2026-05-29 Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning Wenlun Zhang et.al. 2605.31174 null
2026-05-29 From Evidence to Design: Developing an AI-Augmented UX Research Point of View for Digital Wellbeing in Emergency and Public Safety Contexts Olumuyiwa Ayorinde et.al. 2605.31146 null
2026-05-29 Don’t Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning Navin Sriram Ravie et.al. 2605.31119 null
2026-05-29 Extending the UXR Point of View Playbook: Triangulating Insights in Complex Developer Domains Sarah Kianfar et.al. 2605.31104 null
2026-05-29 Task-Focused Memorization for Multimodal Agents Tao Zou et.al. 2605.31075 null
2026-05-28 Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software Nhat-Minh Nguyen et.al. 2605.30353 null
2026-05-28 SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations Qinpei Luo et.al. 2605.30345 null
2026-05-28 Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents Anany Kotawala et.al. 2605.30335 null
2026-05-28 RoboWits: Unexpected Challenges for Robotic Creative Problem Solving Chunru Lin et.al. 2605.30326 null
2026-05-28 Gram: Assessing sabotage propensities via automated alignment auditing David Lindner et.al. 2605.30322 null
2026-05-28 SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents Grant Hamblin et.al. 2605.30314 null
2026-05-28 Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Ziyan Liu et.al. 2605.30159 null
2026-05-28 AnomalyAgent: Training-Free Agentic Models for Zero-/Few-Shot Anomaly Detection Yi Zhang et.al. 2605.30140 null
2026-05-28 Enhancing Multi-Agent Communication through Attention Steering with Context Relevance Hongxiang Zhang et.al. 2605.30136 null
2026-05-28 EvoRepair: Enhancing Vulnerability Repair Agents Through Experience-Based Self-Evolution Haichuan Hu et.al. 2605.30105 null
2026-05-28 SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge? Jiamin Chen et.al. 2605.30104 null
2026-05-28 Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison Tiancheng Yang et.al. 2605.30087 null
2026-05-28 HEART-Bench: Do LLM Agents Exhibit Human-like Psychology? Weihan Peng et.al. 2605.30058 null
2026-05-28 Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection Geremy Loachamín-Suntaxi et.al. 2605.30042 null
2026-05-28 Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas Víctor Gallego et.al. 2605.30003 null
2026-05-28 KairosAgent: Agentic Time Series Forecasting with Fused Semantic Reasoning Kun Feng et.al. 2605.30002 null
2026-05-28 Compass: Navigating Global Marine Lead Data Integration through Expert-Guided LLM Agent Yiming Liu et.al. 2605.29966 null
2026-05-28 Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction Hongtao Wang et.al. 2605.29960 null
2026-05-28 Formalizing Mathematics at Scale Ahmad Rammal et.al. 2605.29955 null
2026-05-28 When Should AI Read the Room? Public Perceptions of Social Intelligence in AI Agents Leena Mathur et.al. 2605.29938 null
2026-05-25 From Model Scaling to System Scaling: Scaling the Harness in Agentic AI Shangding Gu et.al. 2605.26112 null
2026-05-25 DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking Matt L. Wiemann et.al. 2605.26087 null
2026-05-25 Automated Benchmark Auditing for AI Agents and Large Language Models Junlin Wang et.al. 2605.26079 null
2026-05-25 Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use Tianda Sun et.al. 2605.26037 null
2026-05-25 CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists Junlin Yang et.al. 2605.26029 null
2026-05-25 STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Yiming Liang et.al. 2605.26014 null
2026-05-25 Causal methods for LLM development and evaluation Dennis Frauen et.al. 2605.25998 null
2026-05-25 When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation Liyun Zhang et.al. 2605.25981 null
2026-05-25 Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents Haoyi Hu et.al. 2605.25971 null
2026-05-25 Mitigating Provenance-Role Collapse in Long-Term Agents via Typed Memory Representation Zhengda Jin et.al. 2605.25869 null
2026-05-25 When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills Yunfei Wang et.al. 2605.25832 null
2026-05-25 Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network Qiming Ye et.al. 2605.25815 null
2026-05-25 Collaborative Threat-Aware Autonomy (CTAA) Rajnikant Sharma et.al. 2605.25741 null
2026-05-25 Meta-Engineering Harnesses for AI-Native Software Production: A Contract-Driven Adversarial Verification Architecture with Early Deployment Report Satadru Sengupta et.al. 2605.25665 null
2026-05-25 Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents Hao-Hsuan Chen et.al. 2605.25632 null
2026-05-25 CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents Bowen Wang et.al. 2605.25624 null
2026-05-25 DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning Guochao Jiang et.al. 2605.25604 null
2026-05-25 Specular gradient methods for nonsmooth convex optimization in Euclidean spaces: a subgradient selection strategy Kiyuob Jung et.al. 2605.25490 null
2026-05-25 ATWL: A Formal Language for Representing, Comparing, and Reusing Visual Analytics Workflows Natalia Andrienko et.al. 2605.25489 null
2026-05-25 Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki Haoliang Ming et.al. 2605.25480 null
2026-05-22 Routing Equilibrium in Mixed-Autonomy Traffic Networks with Altruistic Autonomous Agents Lihui Yi et.al. 2605.23782 null
2026-05-22 MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection Zhewen Tan et.al. 2605.23723 null
2026-05-22 OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents Jiahao Ying et.al. 2605.23657 null
2026-05-22 RF Instrument Agent (RFIA): Empowering RF Instruments with Natural Language Understanding, Scheduling and Execution of Complex Tasks Chunhui Li et.al. 2605.23636 null
2026-05-22 Investigation of the Two-Dimensional Velocity Field of the Large-Scale Coronal Wave from September 6, 2011 using the SOLERwave Tool Markus Baumgartner-Steinleitner et.al. 2605.23599 null
2026-05-22 Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents Yuandao Cai et.al. 2605.23574 null
2026-05-22 LiveFigure: Generating Editable Scientific Illustration with VLM Agents Chenyang Shao et.al. 2605.23527 null
2026-05-22 B-GRTO: Bootstrapped Group Relative Tool Optimization for Referring Segmentation Mario Markov et.al. 2605.23500 null
2026-05-22 AI Assurance: A Comprehensive Testing Strategy for Enterprise AI Systems Chitra Badagi et.al. 2605.23459 null
2026-05-22 Socially fluent AI decouples conversational signals from source identity in online interaction Lixiang Yan et.al. 2605.23426 null
2026-05-22 When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems Zehao Wang et.al. 2605.23414 null
2026-05-22 From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning Ranxu zhang et.al. 2605.23382 null
2026-05-22 Security, Privacy, and Ethical Risks in OpenClaw Yutong Jin et.al. 2605.23330 null
2026-05-22 Human-in-the-Loop Multi-Agent Ventilator Decision Support with Contextual Bandit Preference Learning Sijia Li et.al. 2605.23320 null
2026-05-22 Parallel Context Compaction for Long-Horizon LLM Agent Serving Musa Cim et.al. 2605.23296 null
2026-05-22 When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming Francesco Corielli et.al. 2605.23278 null
2026-05-22 Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework Hyunjee Park et.al. 2605.23273 null
2026-05-22 EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation Songlin Yang et.al. 2605.23271 null
2026-05-22 6G Communication Networks Enabling Embodied Agents: Architecture and Prototype Lipeng Dai et.al. 2605.23263 null
2026-05-22 Design and Report Benchmarks for Knowledge Work Yining Hua et.al. 2605.23262 null
2026-05-21 MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems Qianshu Cai et.al. 2605.22794 null
2026-05-21 DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback Yunpeng Dong et.al. 2605.22781 null
2026-05-21 Towards a General Intelligence and Interface for Wearable Health Data Girish Narayanswamy et.al. 2605.22759 null
2026-05-21 Self-Evolving Multi-Agent Systems via Decentralized Memory Guangya Hao et.al. 2605.22721 null
2026-05-21 WorkstreamBench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance Thomson Yen et.al. 2605.22664 null
2026-05-21 Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety Piercosma Bisconti et.al. 2605.22643 null
2026-05-21 Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning Banghao Chi et.al. 2605.22642 null
2026-05-21 Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents Ting Liu et.al. 2605.22634 null
2026-05-21 Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents Asaf Yehudai et.al. 2605.22608 null
2026-05-21 Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard Sahar Abdelnabi et.al. 2605.22568 null
2026-05-21 GraphFlow: A Graph-Based Workflow Management for Efficient LLM-Agent Serving Ao Li et.al. 2605.22566 null
2026-05-21 Why Are Agentic Pull Requests Merged or Rejected? An Empirical Study Sien Reeve O. Peralta et.al. 2605.22534 null
2026-05-21 “Refactoring Runaway”: Understanding and Mitigating Tangled Refactorings in Coding Agents for Issue Resolution Zhao Tian et.al. 2605.22526 null
2026-05-21 Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost Simon Dennis et.al. 2605.22502 null
2026-05-21 DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QA Jianing Yin et.al. 2605.22411 null
2026-05-21 AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture Zi Ye et.al. 2605.22366 null
2026-05-21 Benchmarking Autonomous Agents against Temporal, Spatial, and Semantic Evasions Jianan Ma et.al. 2605.22321 null
2026-05-21 Cross-domain benchmarks reveal when coordinated AI agents improve scientific inference from partial evidence Fiona Y. Wong et.al. 2605.22300 null
2026-05-21 SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval Ningyuan Li et.al. 2605.22219 null
2026-05-21 Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Jinyang Wu et.al. 2605.22177 null
2026-05-20 Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling Caleb Winston et.al. 2605.21470 null
2026-05-20 Mem- $π$ : Adaptive Memory through Learning When and What to Generate Xiaoqiang Wang et.al. 2605.21463 null
2026-05-20 Quality and Security Signals in AI-Generated Python Refactoring Pull Requests Mohamed Almukhtar et.al. 2605.21453 null
2026-05-20 Agentic Model Checking Youcheng Sun et.al. 2605.21434 null
2026-05-20 What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema Mahdi Naser Moghadasi et.al. 2605.21404 null
2026-05-20 Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment Roland Pihlakas et.al. 2605.21401 null
2026-05-20 VIPER-MCP: Detecting and Exploiting Taint-Style Vulnerabilities in Model Context Protocol Servers Pengyu Sun et.al. 2605.21392 null
2026-05-20 SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents Bingchen Zhao et.al. 2605.21384 null
2026-05-20 Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents Akshay Manglik et.al. 2605.21347 null
2026-05-20 Frontier: Towards Comprehensive and Accurate LLM Inference Simulation Yicheng Feng et.al. 2605.21312 null
2026-05-20 APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents Yibo Li et.al. 2605.21240 null
2026-05-20 Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints Alexi Canesse et.al. 2605.21085 null
2026-05-20 Causal Past Logic for Runtime Verification of Distributed LLM Agent Workflows Benedikt Bollig et.al. 2605.20923 null
2026-05-20 GenAI-Driven Threat Detection with Microsoft Security Copilot Scott Freitas et.al. 2605.20896 null
2026-05-20 Governance by Construction for Generalist Agents Segev Shlomov et.al. 2605.20874 null
2026-05-20 ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection Yingjia Xu et.al. 2605.20867 null
2026-05-20 MemGym: a Long-Horizon Memory Environment for LLM Agents Wujiang Xu et.al. 2605.20833 null
2026-05-20 DynaMate2: Democratization of Agentic AI for Expert-Designed Custom Workflows Orlando A. Mendible-Barreto et.al. 2605.20819 null
2026-05-20 DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Haoyang Zhang et.al. 2605.20755 null
2026-05-20 Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Amit Roth et.al. 2605.20744 null
2026-05-19 ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning Juncheng Wu et.al. 2605.20176 null
2026-05-19 A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents Vasundra Srinivasan et.al. 2605.20173 null
2026-05-19 What Do Evolutionary Coding Agents Evolve? Nico Pelleriti et.al. 2605.20086 null
2026-05-19 CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning Dachuan Shi et.al. 2605.20075 null
2026-05-19 Probing Embodied LLMs: When Higher Observation Fidelity Hurts Problem Solving Oussama Zenkri et.al. 2605.20072 null
2026-05-19 Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Wenjie Tang et.al. 2605.20061 null
2026-05-19 Hunting Vulnerability Variants in AI Infra: Measurement and Reference-Driven Detection Tian Dong et.al. 2605.20051 null
2026-05-19 Does Code Cleanliness Affect Coding Agents? A Controlled Minimal-Pair Study Priyansh Trivedi et.al. 2605.20049 null
2026-05-19 AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration Jiaqi Liu et.al. 2605.20025 null
2026-05-19 When Skills Don’t Help: A Negative Result on Procedural Knowledge for Tool-Grounded Agents in Offensive Cybersecurity Samuel Jacob Chacko et.al. 2605.20023 null
2026-05-19 Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory Jingwei Sun et.al. 2605.19952 null
2026-05-19 PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents Zhuohan Gu et.al. 2605.19932 null
2026-05-19 LLM Agents Make Collective Belief Dynamics Programmable: Challenges and Research Directions Xin He et.al. 2605.19915 null
2026-05-19 A Closed-loop, State-centric, Multi-agent Framework for Passenger Load Estimation from Heterogeneous Data Streams Yiyao Xu et.al. 2605.19834 null
2026-05-19 From Prompts to Pavement Through Time: Temporal Grounding in Agentic Scene-to-Plan Reasoning Ahmed Y. Gado et.al. 2605.19824 null
2026-05-19 Prior Knowledge or Search? A Study of LLM Agents in Hardware-Aware Code Optimization Dmitry Redko et.al. 2605.19782 null
2026-05-19 Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation Yuxuan Gao et.al. 2605.19779 null
2026-05-19 EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design Gioele Molinari et.al. 2605.19743 null
2026-05-19 Physics-in-the-Loop: A Hybrid Agentic Architecture for Validated CAD Engineering Design Elias Berger et.al. 2605.19717 null
2026-05-19 P2DNav: Panorama-to-Downview Reasoning for Zero-shot Vision-and-Language Navigation Kai Sheng et.al. 2605.19634 null
2026-05-18 Aurora: Unified Video Editing with a Tool-Using Agent Yongsheng Yu et.al. 2605.18748 null
2026-05-18 Code as Agent Harness Xuying Ning et.al. 2605.18747 null
2026-05-18 Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Qianhao Yuan et.al. 2605.18740 null
2026-05-18 Robo-Cortex: A Self-Evolving Embodied Agent via Dual-Grain Cognitive Memory and Autonomous Knowledge Induction Nga Teng Chan et.al. 2605.18729 null
2026-05-18 DexHoldem: Playing Texas Hold’em with Dexterous Embodied System Feng Chen et.al. 2605.18727 null
2026-05-18 EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL Minrui Xu et.al. 2605.18703 null
2026-05-18 Contextualized Dynamic Explanations: A Vision Zhicheng Liu et.al. 2605.18698 null
2026-05-18 SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents Yifan Zhou et.al. 2605.18693 null
2026-05-18 Reversa: A Reverse Documentation Engineering Framework for Converting Legacy Software into Operational Specifications for AI Agents Sanderson Oliveira de Macedo et.al. 2605.18684 null
2026-05-18 Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment S. Bensalem et.al. 2605.18672 null
2026-05-18 SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science Nithin Somasekharan et.al. 2605.18630 null
2026-05-18 Latent Action Reparameterization for Efficient Agent Inference Wenhao Huang et.al. 2605.18597 null
2026-05-18 Not What You Asked For: Typographic Attacks in Household Robot Manipulation Ali Iranmanesh et.al. 2605.18593 null
2026-05-18 Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks Yubin Qu et.al. 2605.18583 null
2026-05-18 MA $^{2}$ P: A Meta-Cognitive Autonomous Intelligent Agents Framework for Complex Persuasion Dingyi Zhang et.al. 2605.18572 null
2026-05-18 LongMINT: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems Hyunji Lee et.al. 2605.18565 null
2026-05-18 STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics Tingfeng Hui et.al. 2605.18548 null
2026-05-18 AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment Zhenlin Wei et.al. 2605.18529 null
2026-05-18 AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers Jungang Zou et.al. 2605.18476 null
2026-05-18 One Developer Is All You Need: A Case Study of an AI-Augmented One-Person Squad in a Brownfield Enterprise Marcelo Vilas Boas et.al. 2605.18461 null
2026-05-18 MARS: Technical Report for the CASTLE Challenge at EgoVis 2026 Haoyu Zhang et.al. 2605.18176 null
2026-05-18 Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability Detection Xin Peng et.al. 2605.18153 null
2026-05-18 Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework Mengyu Sun et.al. 2605.18150 null
2026-05-18 LLM-Guided Communication for Cooperative Multi-Agent Reinforcement Learning Sangjun Bae et.al. 2605.18077 null
2026-05-18 A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback Anika Tabassum et.al. 2605.18073 null
2026-05-18 PPAI: Enabling Personalized LLM Agent Interoperability for Collaborative Edge Intelligence Zile Wang et.al. 2605.18067 null
2026-05-18 PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows Kazuki Kawamura et.al. 2605.18032 null
2026-05-18 TeleCom-Bench: How Far Are Large Language Models from Industrial Telecommunication Applications? Jieting Xiao et.al. 2605.18025 null
2026-05-18 Verify-Gated Completion as Admission Control in a Governed Multi-Agent Runtime: A Bounded Architecture Case Study Hai-Duong Nguyen et.al. 2605.17998 null
2026-05-18 LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injectio Lei Zhao et.al. 2605.17986 null
2026-05-18 Generation Navigator: A State-Aware Agentic Framework for Image Generation Jinming Liu et.al. 2605.17969 null
2026-05-18 SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain Lingtao Mao et.al. 2605.17946 null
2026-05-18 Ethical Hyper-Velocity (EHV): A Provably Deterministic Governance-Aware JIT Compiler Architecture for Agentic Systems Riddhi Mohan Sharma et.al. 2605.17909 null
2026-05-18 Curriculum-Guided Heterogeneous Multi-Agent Intelligence for Multi-UAV Cooperative ISAC Kang Yan et.al. 2605.17905 null
2026-05-18 Agentic Chunking and Bayesian De-chunking of AI Generated Fuzzy Cognitive Maps: A Model of the Thucydides Trap Akash Kumar Panda et.al. 2605.17903 null
2026-05-18 DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition Le Zhang et.al. 2605.17900 null
2026-05-18 Evaluating Cognitive Age Alignment in Interactive AI Agents Yifan Shen et.al. 2605.17894 null
2026-05-18 HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents Woongyeng Yeo et.al. 2605.17873 null
2026-05-18 $\boldsymbol{f}$ -OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control Xianwei Chen et.al. 2605.17862 null
2026-05-18 Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents Ahmad Al-Tawaha et.al. 2605.17830 null
2026-05-14 ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both Ziyu Guo et.al. 2605.15198 null
2026-05-14 FutureSim: Replaying World Events to Evaluate Adaptive Agents Shashwat Goel et.al. 2605.15188 null
2026-05-14 Articraft: An Agentic System for Scalable Articulated 3D Asset Generation Matt Zhou et.al. 2605.15187 null
2026-05-14 Is Grep All You Need? How Agent Harnesses Reshape Agentic Search Sahil Sen et.al. 2605.15184 null
2026-05-14 Hand-in-the-Loop: Improving Dexterous VLA via Seamless Interventional Correction Zhuohang Li et.al. 2605.15157 null
2026-05-14 Self-Distilled Agentic Reinforcement Learning Zhengxi Lu et.al. 2605.15155 null
2026-05-14 APWA: A Distributed Architecture for Parallelizable Agentic Workflows Evan Rose et.al. 2605.15132 null
2026-05-14 Understanding How International Students in the U.S. Are Using Conversational AI to Support Cross-Cultural Adaptation Laleh Nourian et.al. 2605.15127 null
2026-05-14 From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents Md Tahmid Rahman Laskar et.al. 2605.15104 null
2026-05-14 Veritas: A Semantically Grounded Agentic Framework for Memory Corruption Vulnerability Detection in Binaries Xinran Zheng et.al. 2605.15097 null
2026-05-14 Concurrency without Model Changes: Future-based Asynchronous Function Calling for LLMs Guangyu Feng et.al. 2605.15077 null
2026-05-14 Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use Renning Pang et.al. 2605.15041 null
2026-05-14 Orchard: An Open-Source Agentic Modeling Framework Baolin Peng et.al. 2605.15040 null
2026-05-14 WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections Tri Cao et.al. 2605.15030 null
2026-05-14 Multi-Agentic Approach for History Matching of Oil Reservoirs Linar Samigullin et.al. 2605.15028 null
2026-05-14 Toward Securing AI Agents Like Operating Systems Lukas Pirch et.al. 2605.14932 null
2026-05-14 Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systems Shihao Qi et.al. 2605.14892 null
2026-05-14 Holistic Evaluation and Failure Diagnosis of AI Agents Netta Madvil et.al. 2605.14865 null
2026-05-14 Do Coding Agents Understand Least-Privilege Authorization? Zheng Yan et.al. 2605.14859 null
2026-05-14 A Deterministic Agentic Workflow for HS Tariff Classification: Multi-Dimensional Rule Reasoning with Interpretable Decisions Yu Zhang et.al. 2605.14857 null
2026-05-13 Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Zhaowei Wang et.al. 2605.13831 null
2026-05-13 Porting the Nonlinear Optimization Library HiOp to Accelerator-Based Hardware Architectures Slaven Peles et.al. 2605.13736 null
2026-05-13 ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles Yitian Yang et.al. 2605.13725 null
2026-05-13 SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems Hongji Pu et.al. 2605.13716 null
2026-05-13 How to Interpret Agent Behavior Jie Gao et.al. 2605.13625 null
2026-05-13 OpenAaaS: An Open Agent-as-a-Service Framework for Distributed Materials-Informatics Research Peng Kang et.al. 2605.13618 null
2026-05-13 Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation Asim Osman et.al. 2605.13554 null
2026-05-13 RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation Chengzhi Shen et.al. 2605.13542 null
2026-05-13 R^2-Mem: Reflective Experience for Memory Search Xinyuan Wang et.al. 2605.13486 null
2026-05-13 PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agents Mikhail Menschikov et.al. 2605.13481 null
2026-05-13 Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents Narek Maloyan et.al. 2605.13471 null
2026-05-13 Cognifold: Always-On Proactive Memory via Cognitive Folding Suli Wang et.al. 2605.13438 null
2026-05-13 Text2Score: Generating Sheet Music From Textual Prompts Keshav Bhandari et.al. 2605.13431 null
2026-05-13 TRIAGE: Evaluating Prospective Metacognitive Control in LLMs under Resource Constraints Zabir Al Nazi et.al. 2605.13414 null
2026-05-13 Building Interactive Real-Time Agents with Asynchronous I/O and Speculative Tool Calling Coleman Hooper et.al. 2605.13360 null
2026-05-13 Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning Qinchuan Cheng et.al. 2605.13335 null
2026-05-13 IdeaForge: A Knowledge Graph-Grounded Multi-Agent Framework for Cross-Methodology Innovation Analysis and Patent Claim Generation Joy Bose et.al. 2605.13311 null
2026-05-13 D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Yucheng Guo et.al. 2605.13276 null
2026-05-13 ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding Xiao Liu et.al. 2605.13228 null
2026-05-13 Hierarchical Attacks for Multi-Modal Multi-Agent Reasoning Hao Zhou et.al. 2605.13213 null
2026-05-12 LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues Di Wu et.al. 2605.12493 null
2026-05-12 ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents Xuhao Hu et.al. 2605.12481 null
2026-05-12 KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference Alireza Nadali et.al. 2605.12471 null
2026-05-12 Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs Guinan Su et.al. 2605.12460 null
2026-05-12 Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling Eilam Shapira et.al. 2605.12411 null
2026-05-12 ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows Wei Liu et.al. 2605.12376 null
2026-05-12 Agent-Based Post-Hoc Correction of Agricultural Yield Forecasts Matthew Beddows et.al. 2605.12375 null
2026-05-12 Classifier Context Rot: Monitor Performance Degrades with Context Length Sam Martin et.al. 2605.12366 null
2026-05-12 BatchBench: Toward a Workload-Aware Benchmark for Autoscaling Policies in Big Data Batch Processing – A Proposed Framework Venkata Krishna Prasanth Budigi et.al. 2605.12272 null
2026-05-12 Social Welfare under Heterogeneous Time Preferences Sarvin Bahmani et.al. 2605.12251 null
2026-05-12 No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents Zixu Yang et.al. 2605.12240 null
2026-05-12 Harness Engineering as Categorical Architecture Bogdan Banu et.al. 2605.12239 null
2026-05-12 TriBand-BEV: Real-Time LiDAR-Only 3D Pedestrian Detection via Height-Aware BEV and High-Resolution Feature Fusion Mohammad Khoshkdahan et.al. 2605.12220 null
2026-05-12 Goal-Oriented Reasoning for RAG-based Memory in Conversational Agentic LLM Systems Jiazhou Liang et.al. 2605.12213 null
2026-05-12 Rollout Cards: A Reproducibility Standard for Agent Research Charlie Masters et.al. 2605.12131 null
2026-05-12 MPEX AI Digital Twins Milestone Report Gary Staebler et.al. 2605.12116 null
2026-05-12 Property-Level Reconstructability of Agent Decisions: An Anchor-Level Pilot Across Vendor SDK Adapter Regimes Oleg Solozobov et.al. 2605.12078 null
2026-05-12 Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction Zhong Guan et.al. 2605.12070 null
2026-05-12 Learning Agentic Policy from Action Guidance Yuxiang Ji et.al. 2605.12004 null
2026-05-12 The SiMPL Method for Multi-Material Topology Optimization Peter Gangl et.al. 2605.11994 null
2026-05-11 AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks Baraa Al Jorf et.al. 2605.10286 null
2026-05-11 Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution Kai Pan et.al. 2605.10223 null
2026-05-11 V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning Zhiwei Ning et.al. 2605.10172 null
2026-05-11 When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews Sandeep Kumar et.al. 2605.10171 null
2026-05-11 RFAmpDesigner: A Self-Evolving Multi-Agent LLM Framework for Automated Radio Frequency Amplifier Design Hang Lu et.al. 2605.10093 null
2026-05-11 Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust Shijun Lei et.al. 2605.10059 null
2026-05-11 Swarm Skills: A Portable, Self-Evolving Multi-Agent System Specification for Coordination Engineering Xinyu Zhang et.al. 2605.10052 null
2026-05-11 Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables Damon McMillan et.al. 2605.10039 null
2026-05-11 TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning Hangchen Liu et.al. 2605.10038 null
2026-05-11 Bridging the Cognitive Gap: A Unified Memory Paradigm for 6G Agentic AI-RAN Xijun Wang et.al. 2605.10036 null
2026-05-11 Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving Aron Distelzweig et.al. 2605.10034 null
2026-05-11 Combining Mechanical and Agentic Specification Inference for Move Wolfgang Grieskamp et.al. 2605.10005 null
2026-05-11 Continual Harness: Online Adaptation for Self-Improving Foundation Agents Seth Karten et.al. 2605.09998 null
2026-05-11 Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference Sietse Schelpe et.al. 2605.09990 null
2026-05-11 TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents Bihui Yu et.al. 2605.09934 null
2026-05-11 FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning Zehua Pei et.al. 2605.09932 null
2026-05-11 Position: Academic Conferences are Potentially Facing Denominator Gaming Caused by Fully Automated Scientific Agents Rong Shan et.al. 2605.09915 null
2026-05-11 RADAR: Redundancy-Aware Diffusion for Multi-Agent Communication Structure Generation Zhen Zhang et.al. 2605.09907 null
2026-05-11 Deterministic vs. LLM-Controlled Orchestration for COBOL-to-Python Modernization Naing Oo Lwin et.al. 2605.09894 null
2026-05-11 Skill Description Deception Attack against Task Routing in Internet of Agents Jiayi He et.al. 2605.09889 null
2026-05-08 LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling Tong Zheng et.al. 2605.08083 null
2026-05-08 The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents Jiayuan Liu et.al. 2605.08060 null
2026-05-08 Collaborator or Assistnat? How AI Coding Agents Partition Work Across Pull Request Lifecycles Young Jo et.al. 2605.08017 null
2026-05-08 Learning CLI Agents with Structured Action Credit under Selective Observation Haoyang Su et.al. 2605.08013 null
2026-05-08 Graph Representation Learning Augmented Model Manipulation on Federated Fine-Tuning of LLMs Hanlin Cai et.al. 2605.07961 null
2026-05-08 Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents? Anmol Gulati et.al. 2605.07937 null
2026-05-08 AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents Zhengkang Guo et.al. 2605.07926 null
2026-05-08 ADKO: Agentic Decentralized Knowledge Optimization Lucas Nerone Rillo et.al. 2605.07863 null
2026-05-08 RelAgent: LLM Agents as Data Scientists for Relational Learning Xingyue Huang et.al. 2605.07840 null
2026-05-08 Unsafe by Flow: Uncovering Bidirectional Data-Flow Risks in MCP Ecosystem Xinyi Hou et.al. 2605.07836 null
2026-05-08 CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios Taein Lim et.al. 2605.07830 null
2026-05-08 Is a team only as strong as its weakest link? Quantifying the short-board effect with AI Agents Xin Xu et.al. 2605.07773 null
2026-05-08 Coding Agents Don’t Know When to Act Thibaud Gloaguen et.al. 2605.07769 null
2026-05-08 Securing the Dark Matter: A Semantic-Enhanced Neuro-Symbolic Framework for Supply Chain Analysis of Opaque Industrial Software Bowei Ning et.al. 2605.07737 null
2026-05-08 SARC: A Governance-by-Architecture Framework for Agentic AI Systems Gaston Besanson et.al. 2605.07728 null
2026-05-08 SOD: Step-wise On-policy Distillation for Small Language Model Agents Qiyong Zhong et.al. 2605.07725 null
2026-05-08 GASim: A Graph-Accelerated Hybrid Framework for Social Simulation Xuan Zhou et.al. 2605.07692 null
2026-05-08 The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting Lauri Lovén et.al. 2605.07671 null
2026-05-08 Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models Berkehan Ünal et.al. 2605.07649 null
2026-05-08 MemCompiler: Compile, Don’t Inject – State-Conditioned Memory for Embodied Agents Xin Ding et.al. 2605.07594 null
2026-05-07 AI Co-Mathematician: Accelerating Mathematicians with Agentic AI Daniel Zheng et.al. 2605.06651 null
2026-05-07 AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents Nithin Somasekharan et.al. 2605.06607 null
2026-05-07 NeuroAgent: LLM Agents for Multimodal Neuroimaging Analysis and Research Lujia Zhong et.al. 2605.06584 null
2026-05-07 ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Wei Gao et.al. 2605.06534 null
2026-05-07 STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? Hanxiang Chao et.al. 2605.06527 null
2026-05-07 Process Matters more than Output for Distinguishing Humans from Machines Milena Rmus et.al. 2605.06524 null
2026-05-07 Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors Jonas Wiedermann-Möller et.al. 2605.06490 null
2026-05-07 ReasonSTL: Bridging Natural Language and Signal Temporal Logic via Tool-Augmented Process-Rewarded Learning Bowen Ye et.al. 2605.06483 null
2026-05-07 Efficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management Haoyu Zheng et.al. 2605.06472 null
2026-05-07 To What Extent Does Agent-generated Code Require Maintenance? An Empirical Study Shota Sawada et.al. 2605.06464 null
2026-05-07 PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitors Xinmiao Huang et.al. 2605.06455 null
2026-05-07 Constraint Decay: The Fragility of LLM Agents in Backend Code Generation Francesco Dente et.al. 2605.06445 null
2026-05-07 AgenticPrecoding: LLM-Empowered Multi-Agent System for Precoding Optimization Zijiu Yang et.al. 2605.06443 null
2026-05-07 Knowledge Graphs, the Missing Link in Agentic AI-based Formal Verification Vaisakh Naduvodi Viswambharan et.al. 2605.06434 null
2026-05-07 Automated alignment is harder than you think Aleksandr Bowkis et.al. 2605.06390 null
2026-05-07 Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level Nan Jia et.al. 2605.06387 null
2026-05-07 From Agent Loops to Deterministic Graphs: Execution Lineage for Reproducible AI-Native Work Josh Rosen et.al. 2605.06365 null
2026-05-07 Prediction and Empowerment: A Theory of Agency through Bridge Interfaces Richard Csaky et.al. 2605.06346 null
2026-05-07 More Than Can Be Said: A Benchmark and Framework for Pre-Question Scientific Ideation Jie Yu et.al. 2605.06345 null
2026-05-07 MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents Ashwani Anand et.al. 2605.06334 null
2026-05-06 LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents Yijun Lu et.al. 2605.05191 null
2026-05-06 Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours The Verkor Team et.al. 2605.05170 null
2026-05-06 PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World Yunhan Yang et.al. 2605.05163 null
2026-05-06 Executable World Models for ARC-AGI-3 in the Era of Coding Agents Sergey Rodionov et.al. 2605.05138 null
2026-05-06 Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime Tianshu Zhu et.al. 2605.05112 null
2026-05-06 Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Xin Yu et.al. 2605.05040 null
2026-05-06 Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation Zhiqing Cui et.al. 2605.05007 null
2026-05-06 Agentic Vulnerability Reasoning on Windows COM Binaries Hwiwon Lee et.al. 2605.05000 null
2026-05-06 Tailoring Scaffolding to Diagnostic Strategies: Theory-Informed LLM-Based Agents Fatma Betul Gures et.al. 2605.04996 null
2026-05-06 Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiers Senkang Hu et.al. 2605.04984 null
2026-05-06 Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games Yidong He et.al. 2605.04906 null
2026-05-06 VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA Haibin He et.al. 2605.04870 null
2026-05-06 Agentic Repository Mining: A Multi-Task Evaluation Johannes Härtel et.al. 2605.04845 null
2026-05-06 DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents Zhaorun Chen et.al. 2605.04808 null
2026-05-06 AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use Chenglin Yang et.al. 2605.04785 null
2026-05-06 SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies Siddhant Saxena et.al. 2605.04637 null
2026-05-06 SensingAgents: A Multi-Agent Collaborative Framework for Robust IMU Activity Recognition Naiyu Zheng et.al. 2605.04608 null
2026-05-06 Accountable Agents in Software Engineering: An Analysis of Terms of Service and a Research Roadmap Christoph Treude et.al. 2605.04532 null
2026-05-06 SADE: Symptom-Aware Diagnostic Escalation for LLM-Based Network Troubleshooting Kuan-Hao Tseng et.al. 2605.04530 null
2026-05-06 KEET: Explaining Performance of GPU Kernels Using LLM Agents Joshua H. Davis et.al. 2605.04467 null
2026-05-05 OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories Yuwen Du et.al. 2605.04036 null
2026-05-05 SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment Joseph Breda et.al. 2605.04012 null
2026-05-05 From Intent to Execution: Composing Agentic Workflows with Agent Recommendation Kishan Athrey et.al. 2605.03986 null
2026-05-05 Generating Proof-of-Vulnerability Tests to Help Enhance the Security of Complex Software Shravya Kanchi et.al. 2605.03956 null
2026-05-05 MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents Jonathan Steinberg et.al. 2605.03952 null
2026-05-05 Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems Jie Zhou et.al. 2605.03900 null
2026-05-05 Evaluating Generative Models as Interactive Emergent Representations of Human-Like Collaborative Behavior Shinas Shaji et.al. 2605.03855 null
2026-05-05 ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting Jiale Chang et.al. 2605.03804 null
2026-05-05 MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents Bronislav Sidik et.al. 2605.03675 null
2026-05-05 Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies Zirui Tang et.al. 2605.03596 null
2026-05-05 A Skill-Based AI Agentic Pipeline for Library of Congress Subject Indexing Eric H. C. Chow et.al. 2605.03537 null
2026-05-05 Multi-Agent Systems for Root Cause Analysis in Microservices Alexander Naakka et.al. 2605.03505 null
2026-05-05 MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents Ishrith Gowda et.al. 2605.03482 null
2026-05-05 CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification Severin Ye et.al. 2605.03476 null
2026-05-05 Robust Agent Compensation (RAC): Teaching AI Agents to Compensate Srinath Perera et.al. 2605.03409 null
2026-05-05 GeoDecider: A Coarse-to-Fine Agentic Workflow for Explainable Lithology Classification Jiahao Wang et.al. 2605.03383 null
2026-05-05 ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection Shihao Weng et.al. 2605.03378 null
2026-05-05 SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents Yipeng Ouyang et.al. 2605.03353 null
2026-05-05 LLM-ADAM: A Generalizable LLM Agent Framework for Pre-Print Anomaly Detection in Additive Manufacturing Ahmadreza Eslaminia et.al. 2605.03328 null
2026-05-05 Revisiting the Travel Planning Capabilities of Large Language Models Bo-Wen Zhang et.al. 2605.03308 null
2026-05-01 Can Coding Agents Reproduce Findings in Computational Materials Science? Ziyang Huang et.al. 2605.00803 null
2026-05-01 Position: agentic AI orchestration should be Bayes-consistent Theodore Papamarkou et.al. 2605.00742 null
2026-05-01 Self-Adaptive Multi-Agent LLM-Based Security Pattern Selection for IoT Systems Saeid Jamshidi et.al. 2605.00741 null
2026-05-01 To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling Qinyuan Wu et.al. 2605.00737 null
2026-05-01 Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory Derong Xu et.al. 2605.00702 null
2026-05-01 DySRec: Dynamic Context-Aware Psychometric Scale Recommendation via Multi-Agent Collaboration Yanzeng Li et.al. 2605.00574 null
2026-05-01 Structure Liberates: How Constrained Sensemaking Produces More Novel Research Output James Mooney et.al. 2605.00557 null
2026-05-01 A11y-Compressor: A Framework for Enhancing the Efficiency of GUI Agent Observations through Visual Context Reconstruction and Redundancy Reduction Michito Takeshita et.al. 2605.00551 null
2026-05-01 SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters Dongxin Guo et.al. 2605.00528 null
2026-05-01 LLM-Oriented Information Retrieval: A Denoising-First Perspective Lu Dai et.al. 2605.00505 null
2026-05-01 Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Kerui Chen et.al. 2605.00444 null
2026-05-01 AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning Haotian Zhao et.al. 2605.00425 null
2026-05-01 Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents Maksym Nechepurenko et.al. 2605.00420 null
2026-05-01 ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Zihan Lin et.al. 2605.00380 null
2026-05-01 Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration Chunlei Meng et.al. 2605.00370 null
2026-05-01 Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning Chengshuai Shi et.al. 2605.00347 null
2026-05-01 AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go? Ranit Karmakar et.al. 2605.00334 null
2026-05-01 On Aubry’s completeness conjecture Tianqi Shi et.al. 2605.00305 null
2026-04-30 High-Probability Convergence in Decentralized Stochastic Optimization with Gradient Tracking Aleksandar Armacki et.al. 2605.00281 null
2026-04-30 Agentic AI for Trip Planning Optimization Application Tiejin Chen et.al. 2605.00276 null
2026-04-30 FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption Yanting Wang et.al. 2604.28157 null
2026-04-30 Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows Chenxin Li et.al. 2604.28139 null
2026-04-30 Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent Sandboxes Tianyuan Wu et.al. 2604.28138 null
2026-04-30 TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering An-Yang Ji et.al. 2604.28076 null
2026-04-30 Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception Neemias B da Silva et.al. 2604.28048 null
2026-04-30 Collaborative Agent Reasoning Engineering (CARE): A Three-Party Design Methodology for Systematically Engineering AI Agents with Subject Matter Experts, Developers, and Helper Agents Rahul Ramachandran et.al. 2604.28043 null
2026-04-30 Exploring Interaction Paradigms for LLM Agents in Scientific Visualization Jackson Vonderhorst et.al. 2604.27996 null
2026-04-30 MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection Weihai Lu et.al. 2604.27934 null
2026-04-30 Building Persona-Based Agents On Demand: Tailoring Multi-Agent Workflows to User Needs Giuseppe Arbore et.al. 2604.27882 null
2026-04-30 Modeling Clinical Concern Trajectories in Language Model Agents Sukesh Subaharan et.al. 2604.27872 null
2026-04-30 Rethinking Agentic Reinforcement Learning In Large Language Models Fangming Cui et.al. 2604.27859 null
2026-04-30 CastFlow: Learning Role-Specialized Agentic Workflows for Time Series Forecasting Bokai Pan et.al. 2604.27840 null
2026-04-30 ObjectGraph: From Document Injection to Knowledge Traversal – A Native File Format for the Agentic Era Mohit Dubey et.al. 2604.27820 null
2026-04-30 WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments Jinchao Li et.al. 2604.27776 null
2026-04-30 Macroscopic photon counting beating the Poisson noise limit Timon Schapeler et.al. 2604.27761 null
2026-04-30 Knowledge Graph Representations for LLM-Based Policy Compliance Reasoning Wilder Baldwin et.al. 2604.27713 null
2026-04-30 Contextual Agentic Memory is a Memo, Not True Memory Binyan Xu et.al. 2604.27707 null
2026-04-30 Bridging Values and Behavior: A Hierarchical Framework for Proactive Embodied Agents Chunhui Zhang et.al. 2604.27699 null
2026-04-30 HAVEN: Hybrid Automated Verification ENgine for UVM Testbench Synthesis with LLMs Chang-Chih Meng et.al. 2604.27643 null
2026-04-30 A stochastic agent-based extension of the GSM2 model for particle therapy: cell-cycle dynamics, dose-rate dependence, and fractionation effects Francesco G. Cordoni et.al. 2604.27630 null
2026-04-29 Artistic Practice Opportunities in CST Evaluations: A Longitudinal Group Deployment of ArtKrit Catherine Liu et.al. 2604.26935 null
2026-04-29 Hot Fixing in the Wild Carol Hanna et.al. 2604.26892 null
2026-04-29 Bian Que: An Agentic Framework with Flexible Skill Arrangement for Online System Operations Bochao Liu et.al. 2604.26805 null
2026-04-29 GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents GLM-V Team et.al. 2604.26752 null
2026-04-29 From Hypotheses to Factors: Constrained LLM Agents in Cryptocurrency Markets Yikuan Huang et.al. 2604.26747 null
2026-04-29 FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow Sina Heidari et.al. 2604.26666 null
2026-04-29 AgentSim: A Platform for Verifiable Agent-Trace Simulation Saber Zerhoudi et.al. 2604.26653 null
2026-04-29 OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory Jinze Li et.al. 2604.26622 null
2026-04-29 AGEL-Comp: A Neuro-Symbolic Framework for Compositional Generalization in Interactive Agents Mahnoor Shahid et.al. 2604.26522 null
2026-04-29 SecMate: Multi-Agent Adaptive Cybersecurity Troubleshooting with Tri-Context Personalization Yair Meidan et.al. 2604.26394 null
2026-04-29 DreamProver: Evolving Transferable Lemma Libraries via a Wake-Sleep Theorem-Proving Agent Youyuan Zhang et.al. 2604.26311 null
2026-04-29 SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering Tasks Jiao Chen et.al. 2604.26278 null
2026-04-29 Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering Happy Bhati et.al. 2604.26275 null
2026-04-29 Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents Hung Dang et.al. 2604.26274 null
2026-04-29 LATTICE: Evaluating Decision Support Utility of Crypto Agents Aaron Chan et.al. 2604.26235 null
2026-04-29 When Agents Shop for You: Role Coherence in AI-Mediated Markets Soogand Alavi et.al. 2604.26220 null
2026-04-29 Hierarchical Long-Term Semantic Memory for LinkedIn’s Hiring Agent Zhentao Xu et.al. 2604.26197 null
2026-04-28 Lifting Embodied World Models for Planning and Control Alex N. Wang et.al. 2604.26182 null
2026-04-28 Beyond Screenshots: Evaluating VLMs’ Understanding of UI Animations Chen Liang et.al. 2604.26148 null
2026-04-28 I Would If I Could: Reasoning about Dynamics of Actions in Multi-Agent Systems Rustam Galimullin et.al. 2604.26053 null
2026-04-28 Recursive Multi-Agent Systems Xiyuan Yang et.al. 2604.25917 null
2026-04-28 From Threads to Trajectories: A Multi-LLM Pipeline for Community Knowledge Extraction from GitHub Issue Discussions Nazia Shehnaz Joynab et.al. 2604.25880 null
2026-04-28 Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses Jiahang Lin et.al. 2604.25850 null
2026-04-28 Semi-Markov Reinforcement Learning for City-Scale EV Ride-Hailing with Feasibility-Guaranteed Actions An Nguyen et.al. 2604.25848 null
2026-04-28 From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling Jianghao Lin et.al. 2604.25847 null
2026-04-28 Towards Agentic Investigation of Security Alerts Even Eilertsen et.al. 2604.25846 null
2026-04-28 KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning Yixuan Huang et.al. 2604.25788 null
2026-04-28 SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing? Noam Tarshish et.al. 2604.25737 null
2026-04-28 Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study Srikanta Prasad S et.al. 2604.25724 null
2026-04-28 Think Before You Act – A Neurocognitive Governance Model for Autonomous AI Agents Eranga Bandara et.al. 2604.25684 null
2026-04-28 Optimizing ground state preparation protocols with autoresearch Luis Mantilla Calderón et.al. 2604.25610 null
2026-04-28 SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents Mengyao Du et.al. 2604.25562 null
2026-04-28 From CRUD to Autonomous Agents: Formal Validation and Zero-Trust Security for Semantic Gateways in AI-Native Enterprise Systems Ignacio Peyrano et.al. 2604.25555 null
2026-04-28 Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows Shivam Rawat et.al. 2604.25345 null
2026-04-28 Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation Lanshan He et.al. 2604.25318 null
2026-04-28 MARD: A Multi-Agent Framework for Robust Android Malware Detection Xueying Zeng et.al. 2604.25264 null
2026-04-28 AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery Lei Xiong et.al. 2604.25256 null
2026-04-28 Value-Sensitive AI for Prayer: Balancing the Agencies Between Human and AI Agents in Spiritual Context Soonho Kwon et.al. 2604.25230 null
2026-04-28 DATAREEL: Automated Data-Driven Video Story Generation with Animations Ridwan Mahbub et.al. 2604.25220 null
2026-04-28 AgentDID: Trustless Identity Authentication for AI Agents Minghui Xu et.al. 2604.25189 null
2026-04-27 The Last Human-Written Paper: Agent-Native Research Artifacts Jiachen Liu et.al. 2604.24658 null
2026-04-27 AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents Yixiang Zhang et.al. 2604.24657 null
2026-04-27 K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology Soyeon Kim et.al. 2604.24645 null
2026-04-27 Skill Retrieval Augmentation for Agentic AI Weihang Su et.al. 2604.24594 null
2026-04-27 Measuring the Unmeasurable: Markov Chain Reliability for LLM Agents Phat T. Tran-Truong et.al. 2604.24579 null
2026-04-27 FastOMOP: A Foundational Architecture for Reliable Agentic Real-World Evidence Generation on OMOP CDM data Niko Moeller-Grell et.al. 2604.24572 null
2026-04-27 Mono2Sls: Automated Monolith-to-Serverless Migration via Multi-Stage Pipeline with Static Analysis Xingyan Chen et.al. 2604.24550 null
2026-04-27 Beyond the Attention Stability Boundary: Agentic Self-Synthesizing Reasoning Protocols Dahlia Shehata et.al. 2604.24512 null
2026-04-27 GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems Pablo Mateo-Torrejón et.al. 2604.24477 null
2026-04-27 Agentic clinical reasoning over longitudinal myeloma records: a retrospective evaluation against expert consensus Johannes Moll et.al. 2604.24473 null
2026-04-27 On the Footprints of Reviewer Bots Feedback on Agentic Pull Requests in OSS GitHub Repositories Syeda Kaneez Fatima et.al. 2604.24450 null
2026-04-27 PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model Sinin Zhang et.al. 2604.24443 null
2026-04-27 AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark Hongxin Li et.al. 2604.24441 null
2026-04-27 Kwai Summary Attention Technical Report Chenglong Chu et.al. 2604.24432 null
2026-04-27 DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents Junshuo Zhang et.al. 2604.24320 null
2026-04-27 AutoQResearch: LLM-Guided Closed-Loop Policy Search for Adaptive Variational Quantum Optimization Monit Sharma et.al. 2604.24283 null
2026-04-27 RefEvo: Agentic Design with Co-Evolutionary Verification for Agile Reference Model Generation Yifan Zhang et.al. 2604.24218 null
2026-04-27 Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis Jiahong Xiang et.al. 2604.24212 null
2026-04-27 AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization Zonghao Ying et.al. 2604.24118 null
2026-04-27 Closing the Loop: A Software Framework for AI to Support Business Decision Making Jeffrey Wong et.al. 2604.24116 null
2026-04-24 How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks Longju Bai et.al. 2604.22750 null
2026-04-24 A dataset of early blockchain-registered AI agents on Ethereum Yulin Liu et.al. 2604.22652 null
2026-04-24 PASS: A Provenanced Access Subaccount System for Blockchain Wallets Jay Yu et.al. 2604.22602 null
2026-04-24 QuantClaw: Precision Where It Matters for OpenClaw Manyi Zhang et.al. 2604.22577 null
2026-04-24 LARA: Validation-Driven Agentic Supercomputer Workflows for Atomistic Modeling William Dawson et.al. 2604.22571 null
2026-04-24 Relational Archetypes: A Comparative Analysis of AV-Human and Agent-Human Interactions Antoni Lorente et.al. 2604.22564 null
2026-04-24 Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents Xirui Li et.al. 2604.22452 null
2026-04-24 AgentSearchBench: A Benchmark for AI Agent Search in the Wild Bin Wu et.al. 2604.22436 null
2026-04-24 Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation Biagio Andreucci et.al. 2604.22427 null
2026-04-24 Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding Mingchen Shao et.al. 2604.22245 null
2026-04-24 Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QA Zhanli Li et.al. 2604.22239 null
2026-04-24 Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning Chaoran Chen et.al. 2604.22191 null
2026-04-24 Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems Jun He et.al. 2604.22136 null
2026-04-23 Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework Tharindu Kumarage et.al. 2604.22119 null
2026-04-23 Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents Seyed Moein Abtahi et.al. 2604.22085 null
2026-04-23 Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results Benjamin Kohler et.al. 2604.21965 null
2026-04-23 Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models Chee Wei Tan et.al. 2604.21896 null
2026-04-23 Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study Zihan Wang et.al. 2604.21829 null
2026-04-23 Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows Anuj Sadani et.al. 2604.21816 null
2026-04-23 Phenomenological Detector Design and Optimization in Vertically-Integrated Differentiable Full Simulations with Agentic-AI Wonyong Chung et.al. 2604.21804 null
2026-04-23 Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems Ye Yu et.al. 2604.21794 null
2026-04-23 Less Is More: Measuring How LLM Involvement affects Chatbot Accuracy in Static Analysis Krishna Narasimhan et.al. 2604.21746 null
2026-04-23 AEL: Agent Evolving Learning for Open-Ended Environments Wujiang Xu et.al. 2604.21725 null
2026-04-23 Speed-oriented quantum circuit backend Sören Wilkening et.al. 2604.21656 null
2026-04-23 DryRUN: On the Role of Public Tests in LLM-Driven Code Generation Kaushitha Silva et.al. 2604.21598 null
2026-04-23 AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use Yuanjie Lyu et.al. 2604.21590 null
2026-04-23 GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation Yitong Zhou et.al. 2604.21501 null
2026-04-23 MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks Run Hao et.al. 2604.21477 null
2026-04-23 Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision Chentao Li et.al. 2604.21461 null
2026-04-23 The Privacy Guardian Agent: Towards Trustworthy AI Privacy Agents Vincent Freiberger et.al. 2604.21455 null
2026-04-23 AI-Gram: When Visual Agents Interact in a Social Network Andrew Shin et.al. 2604.21446 null
2026-04-23 HiCrew: Hierarchical Reasoning for Long-Form Video Understanding via Question-Aware Multi-Agent Collaboration Yuehan Zhu et.al. 2604.21444 null
2026-04-23 FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation Jinhee Jang et.al. 2604.21420 null
2026-04-23 Privacy-Preserving Distributed Stochastic Optimization with Homomorphic Encryption and Heterogeneous Stepsizes Haoqiang Zhou et.al. 2604.21381 null
2026-04-23 VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation Qijun Han et.al. 2604.21375 null
2026-04-23 CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents Wenjie Fu et.al. 2604.21308 null
2026-04-22 Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem Travis LaCroix et.al. 2604.20805 null
2026-04-22 Synthesizing Multi-Agent Harnesses for Vulnerability Discovery Hanzhi Liu et.al. 2604.20801 null
2026-04-22 SWE-chat: Coding Agent Interactions From Real Users in the Wild Joachim Baumann et.al. 2604.20779 null
2026-04-22 Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization Shan He et.al. 2604.20714 null
2026-04-22 Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows Shivani Kumar et.al. 2604.20658 null
2026-04-22 CHORUS: An Agentic Framework for Generating Realistic Deliberation Data A. Koursaris et.al. 2604.20651 null
2026-04-22 Trust, Lies, and Long Memories: Emergent Social Dynamics and Reputation in Multi-Round Avalon with LLM Agents Suveen Ellawela et.al. 2604.20582 null
2026-04-22 MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills Yingyong Hou et.al. 2604.20441 null
2026-04-22 WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning Juyong Jiang et.al. 2604.20398 null
2026-04-22 R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling Aijia Cheng et.al. 2604.20316 null
2026-04-22 FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory Yingjie Gu et.al. 2604.20300 null
2026-04-22 Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows Hardy Chen et.al. 2604.20200 null
2026-04-22 How is a gas sensor poisoned by volatile methylsiloxanes? Heng Liu et.al. 2604.20197 null
2026-04-22 Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning Ronghao Ni et.al. 2604.20179 null
2026-04-22 Stateless Decision Memory for Enterprise AI Agents Vasundra Srinivasan et.al. 2604.20158 null
2026-04-22 Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models Sachin Kumar et.al. 2604.20148 null
2026-04-22 SAKE: Self-aware Knowledge Exploitation-Exploration for Grounded Multimodal Named Entity Recognition Jielong Tang et.al. 2604.20146 null
2026-04-22 An Agentic Approach to Metadata Reasoning Jiani Zhang et.al. 2604.20144 null
2026-04-22 HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Darsh Kachroo et.al. 2604.20140 null
2026-04-22 AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation Joyjit Roy et.al. 2604.20134 null
2026-04-21 Recent Advances in Causal Analysis of the Stochastic Frontier Model Samuele Centorrino et.al. 2604.19693 null
2026-04-21 InHabit: Leveraging Image Foundation Models for Scalable 3D Human Placement Nikita Kister et.al. 2604.19673 null
2026-04-21 Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Yi Zhong et.al. 2604.19667 null
2026-04-21 ECLASS-Augmented Semantic Product Search for Electronic Components Nico Baumgart et.al. 2604.19664 null
2026-04-21 An AI Agent Execution Environment to Safeguard User Data Robert Stanley et.al. 2604.19657 null
2026-04-21 SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models Josue Torres-Fonseca et.al. 2604.19638 null
2026-04-21 Time Series Augmented Generation for Financial Applications Anton Kolonin et.al. 2604.19633 null
2026-04-21 Goal-Oriented Semantic Communication for Logical Decision Making Ahmet Faruk Saz et.al. 2604.19614 null
2026-04-21 AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories Xue Xia et.al. 2604.19606 null
2026-04-21 A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression Jincheng Ren et.al. 2604.19572 null
2026-04-21 Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment Bobo Li et.al. 2604.19548 null
2026-04-21 Mesh Memory Protocol: Semantic Infrastructure for Multi-Agent LLM Systems Hongwei Xu et.al. 2604.19540 null
2026-04-21 Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps Alankrit Chona et.al. 2604.19533 null
2026-04-21 Revac: A Social Deduction Reasoning Agent Mihir Shriniwas Arya et.al. 2604.19523 null
2026-04-21 From Experience to Skill: Multi-Agent Generative Engine Optimization via Reusable Strategy Learning Beining Wu et.al. 2604.19516 null
2026-04-21 Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies Jess Jones et.al. 2604.19509 null
2026-04-21 Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents Vasundra Srininvasan et.al. 2604.19457 null
2026-04-21 Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture The Flag Challenges Ali Al-Kaswan et.al. 2604.19354 null
2026-04-21 Rethinking Scale: Deployment Trade-offs of Small Language Models under Agent Paradigms Xinlin Wang et.al. 2604.19299 null
2026-04-21 Explicit Trait Inference for Multi-Agent Coordination Suhaib Abdurahman et.al. 2604.19278 null
2026-04-21 iCoRe: An Iterative Correlation-Aware Retriever for Bug Reproduction Test Generation Junyi Wang et.al. 2604.19224 null
2026-04-21 ClawNet: Human-Symbiotic Agent Network for Cross-User Autonomous Cooperation Zhiqin Yang et.al. 2604.19211 null
2026-04-21 RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation Feng Jiang et.al. 2604.19092 null
2026-04-21 Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery Abhinav Agarwal et.al. 2604.19049 null
2026-04-21 Explore Like Humans: Autonomous Exploration with Online SG-Memo Construction for Embodied Agents Xu Chen et.al. 2604.19034 null
2026-04-21 ClawCoin: An Agentic AI-Native Cryptocurrency for Decentralized Agent Economies Shaoyu Li et.al. 2604.19026 null
2026-04-21 On Accelerating Grounded Code Development for Research Santosh Ganji et.al. 2604.19022 null
2026-04-21 Security Is Relative: Training-Free Vulnerability Detection via Multi-Agent Behavioral Contract Synthesis Yongchao Wang et.al. 2604.19012 null
2026-04-21 Debating the Unspoken: Role-Anchored Multi-Agent Reasoning for Half-Truth Detection Yixuan Tang et.al. 2604.19005 null
2026-04-21 A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation Liping Wang et.al. 2604.18988 null
2026-04-21 Gated Coordination for Efficient Multi-Agent Collaboration in Minecraft Game HuaDong Jian et.al. 2604.18975 null
2026-04-21 AutomationBench Daniel Shepard et.al. 2604.18934 null
2026-04-20 AI scientists produce results without reasoning scientifically Martiño Ríos-García et.al. 2604.18805 null
2026-04-20 Mango: Multi-Agent Web Navigation via Global-View Optimization Weixi Tong et.al. 2604.18779 null
2026-04-20 CHICO-Agent: An LLM Agent for the Cross-layer Optimization of 2.5D and 3D Chiplet-based Systems Qihang Wu et.al. 2604.18764 null
2026-04-20 A Scientific Human-Agent Reproduction Pipeline Joschka Birk et.al. 2604.18752 null
2026-04-20 Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs Kevin Murphy et.al. 2604.18576 null
2026-04-20 QRAFTI: An Agentic Framework for Empirical Research in Quantitative Finance Terence Lim et.al. 2604.18500 null
2026-04-20 Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts Jacob Morrison et.al. 2604.18473 null
2026-04-20 TypeScript Repository Indexing for Code Agent Retrieval Junsong Pu et.al. 2604.18413 null
2026-04-20 StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning Daoyu Wang et.al. 2604.18401 null
2026-04-20 Capturing Monetarily Exploitable Vulnerability in Smart Contracts via Auditor Knowledge-Learning Fuzzing Bowen Cai et.al. 2604.18395 null
2026-04-20 OpenGame: Open Agentic Coding for Games Yilei Jiang et.al. 2604.18394 null
2026-04-20 Dissecting AI Trading: Behavioral Finance and Market Bubbles Shumiao Ouyang et.al. 2604.18373 null
2026-04-20 ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship Zhaopei Huang et.al. 2604.18356 null
2026-04-20 HiGMem: A Hierarchical and LLM-Guided Memory System for Long-Term Conversational Agents Shuqi Cao et.al. 2604.18349 null
2026-04-20 Reliability of AI Bots Footprints in GitHub Actions CI/CD Workflows Syed Muhammad Ashhar Shah et.al. 2604.18334 null
2026-04-20 Will People Enjoy a Robot Trainer? A Case Study with Snoopie the Pacerbot Maximilian Du et.al. 2604.18331 null
2026-04-20 EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents Paolo Riva et.al. 2604.18271 null
2026-04-20 Aether: Network Validation Using Agentic AI and Digital Twin Jordan Auge et.al. 2604.18233 null
2026-04-20 AgenTEE: Confidential LLM Agent Execution on Edge Devices Sina Abdollahi et.al. 2604.18231 null
2026-04-20 WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models Xinping Lei et.al. 2604.18224 null
2026-04-20 Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration Qifan Zhang et.al. 2604.18131 null
2026-04-20 Architectural Design Decisions in AI Agent Harnesses Hu Wei et.al. 2604.18071 null
2026-04-20 First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows Sihao Xing et.al. 2604.18038 null
2026-04-20 Topology-Aware LLM-Driven Social Simulation: A Unified Framework for Efficient and Realistic Agent Dynamics Yuwei Xu et.al. 2604.18011 null
2026-04-19 DORA Explorer: Improving the Exploration Ability of LLMs Without Training Priya Gurjar et.al. 2604.17244 null
2026-04-19 A Multi-Agent Approach for Claim Verification from Tabular Data Documents Rudra Ranajee Saha et.al. 2604.17225 null
2026-04-19 Shepherding UAV Swarm with Action Prediction Based on Movement Constraints Yusuke Tsunoda et.al. 2604.17189 null
2026-04-19 BranchBench: Aligning Database Branching with Agentic Demands Elaine Ang et.al. 2604.17180 null
2026-04-18 Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks Tyler H. Merves et.al. 2604.17159 null
2026-04-18 Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM Collaboration Sukwon Yun et.al. 2604.17148 null
2026-04-18 SeekerGym: A Benchmark for Reliable Information Seeking Remy Kim et.al. 2604.17143 null
2026-04-18 HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads Justice Owusu Agyemang et.al. 2604.17111 null
2026-04-18 From Clinical Intent to Clinical Model: An Autonomous Coding-Agent Framework for Clinician-driven AI Development Zihao Zhao et.al. 2604.17110 null
2026-04-18 Live LTL Progress Tracking: Towards Task-Based Exploration Noel Brindise et.al. 2604.17106 null
2026-04-18 GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0) Jiaqing Liang et.al. 2604.17091 null
2026-04-18 From Necklaces to Coalitions: Fair and Self-Interested Distribution of Coalition Value Calculations Terry R. Payne et.al. 2604.17057 null
2026-04-18 Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF) Tianbao Zhang et.al. 2604.17025 null
2026-04-18 Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation Huije Lee et.al. 2604.17020 null
2026-04-18 Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents Sukai Huang et.al. 2604.17019 null
2026-04-18 False Security Confidence in Benign LLM Code Generation Xiaolei Ren et.al. 2604.17014 null
2026-04-18 Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory Poisoning Jiachen Qian et.al. 2604.16966 null
2026-04-18 Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation Minyan Luo et.al. 2604.16958 null
2026-04-18 ClimAgent: LLM as Agents for Autonomous Open-ended Climate Science Analysis Hao Wang et.al. 2604.16922 null
2026-04-18 Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning Weiyu Ma et.al. 2604.16918 null
2026-04-17 ChemGraph-XANES: An Agentic Framework for XANES Simulation and Analysis Vitor F. Grizzi et.al. 2604.16205 null
2026-04-17 MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation Yi Lin et.al. 2604.16175 null
2026-04-17 AstroVLM: Expert Multi-agent Collaborative Reasoning for Astronomical Imaging Quality Diagnosis Yaohui Han et.al. 2604.16024 null
2026-04-17 SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems Hikaru Shindo et.al. 2604.16022 null
2026-04-17 Neurosymbolic Repo-level Code Localization Xiufeng Xu et.al. 2604.16021 null
2026-04-17 AgentV-RL: Scaling Reward Modeling with Agentic Verifier Jiazheng Zhang et.al. 2604.16004 null
2026-04-17 Weak-Link Optimization for Multi-Agent Reasoning and Collaboration Haoyu Bian et.al. 2604.15972 null
2026-04-17 Integrating Graphs, Large Language Models, and Agents: Reasoning and Retrieval Hamed Jelodar et.al. 2604.15951 null
2026-04-17 New Kids: An Architecture and Performance Investigation of Second-Generation Serverless Platforms Trever Schirmer et.al. 2604.15916 null
2026-04-17 Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents Xing Zhang et.al. 2604.15877 null
2026-04-17 CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution Shidong Yang et.al. 2604.15840 null
2026-04-17 Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4 Chengwu Liu et.al. 2604.15839 null
2026-04-17 Exploring Agentic Visual Analytics: A Co-Evolutionary Framework of Roles and Workflows Tianqi Luo et.al. 2604.15813 null
2026-04-17 MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents Weiwei Xie et.al. 2604.15774 null
2026-04-17 The World Leaks the Future: Harness Evolution for Future Prediction Agents Chuyang Wei et.al. 2604.15719 null
2026-04-17 GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows Jize Wang et.al. 2604.15715 null
2026-04-17 Just Type It in Isabelle! AI Agents Drafting, Mechanizing, and Generalizing from Human Hints Kevin Kappelmann et.al. 2604.15713 null
2026-04-17 VoxMind: An End-to-End Agentic Spoken Dialogue System Tianle Liang et.al. 2604.15710 null
2026-04-17 Bilevel Optimization of Agent Skills via Monte Carlo Tree Search Chenyi Huang et.al. 2604.15709 null
2026-04-17 Long-Term Memory for VLA-based Agents in Open-World Task Execution Xu Huang et.al. 2604.15671 null
2026-04-16 MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation Yan Li et.al. 2604.15309 null
2026-04-16 CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas Emanuel Tewolde et.al. 2604.15267 null
2026-04-16 Agentic Microphysics: A Manifesto for Generative AI Safety Federico Pierucci et.al. 2604.15236 null
2026-04-16 RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography Mélanie Roschewitz et.al. 2604.15231 null
2026-04-16 Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines Marcel Wagenländer et.al. 2604.15186 null
2026-04-16 DPC: Training-Free Text-to-SQL Candidate Selection via Dual-Paradigm Consistency Boyan Li et.al. 2604.15163 null
2026-04-16 Autonomous Evolution of EDA Tools: Multi-Agent Self-Evolved ABC Cunxi Yu et.al. 2604.15082 null
2026-04-16 CoGrid & the Multi-User Gymnasium: A Framework for Multi-Agent Experimentation Chase McDonald et.al. 2604.15044 null
2026-04-16 From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench Ke Xu et.al. 2604.15037 null
2026-04-16 Autogenesis: A Self-Evolving Agent Protocol Wentao Zhang et.al. 2604.15034 null
2026-04-16 Dr.~RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement Wenji Fang et.al. 2604.14989 null
2026-04-16 Agentic Explainability at Scale: Between Corporate Fears and XAI Needs Yomna Elsayed et.al. 2604.14984 null
2026-04-16 SAGER: Self-Evolving User Policy Skills for Recommendation Agent Zhen Tao et.al. 2604.14972 null
2026-04-16 RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models Gabriele Mattioli et.al. 2604.14951 null
2026-04-16 IE as Cache: Information Extraction Enhanced Agentic Reasoning Hang Lv et.al. 2604.14930 null
2026-04-16 ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints Pei-An Chen et.al. 2604.14902 null
2026-04-16 Toward Agentic RAG for Ukrainian Marta Sumyk et.al. 2604.14896 null
2026-04-16 Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis Zhiyuan Zhai et.al. 2604.14877 null
2026-04-16 Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-CodeX Zhonghao Yang et.al. 2604.14858 null
2026-04-16 SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling Hao Han et.al. 2604.14820 null
2026-04-16 World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems Runze Li et.al. 2604.14732 null
2026-04-16 Layered Mutability: Continuity and Governance in Persistent Self-Modifying Agents Krti Tallam et.al. 2604.14717 null
2026-04-16 HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks Fan Cui et.al. 2604.14709 null
2026-04-16 CAMO: An Agentic Framework for Automated Causal Discovery from Micro Behaviors to Macro Emergence in LLM Agent Simulations Xiangning Yu et.al. 2604.14691 null
2026-04-16 AIPC: Agent-Based Automation for AI Model Deployment with Qualcomm AI Runtime Jianhao Su et.al. 2604.14661 null
2026-04-15 Enhancing Local Life Service Recommendation with Agentic Reasoning in Large Language Model Shiteng Cao et.al. 2604.14051 null
2026-04-15 A Complete Symmetry Classification of Shallow ReLU Networks Pranavkrishnan Ramakrishnan et.al. 2604.14037 null
2026-04-15 Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents Kangsan Kim et.al. 2604.14004 null
2026-04-15 Acts of Configuration: Rethinking Provenance, Temporality and Legitimacy in Post-Mortem Agents Kellie Yu Hui Sim et.al. 2604.13996 null
2026-04-15 CollabCoder: Plan-Code Co-Evolution via Collaborative Decision-Making for Efficient Code Generation Duy Tung Doan et.al. 2604.13946 null
2026-04-15 AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot Joydeep Biswas et.al. 2604.13940 null
2026-04-15 AI Coding Agents Need Better Compiler Remarks Akash Deo et.al. 2604.13927 null
2026-04-15 DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off Xiaofan Li et.al. 2604.13902 null
2026-04-15 Sandpile Economics: Theory, Identification, and Evidence Diego Vallarino et.al. 2604.13890 null
2026-04-15 ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution Shouzheng Huang et.al. 2604.13787 null
2026-04-15 The cognitive companion: a lightweight parallel monitoring architecture for detecting and recovering from reasoning degradation in LLM agents Rafflesia Khan et.al. 2604.13759 null
2026-04-15 Rethinking AI Hardware: A Three-Layer Cognitive Architecture for Autonomous Agents Li Chen et.al. 2604.13757 null
2026-04-15 Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA Yuanlei Zheng et.al. 2604.13731 null
2026-04-15 Beyond Arrow’s Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration Sayan Kumar Chaki et.al. 2604.13705 null
2026-04-15 IndicDB – Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages Aviral Dawar et.al. 2604.13686 null
2026-04-15 SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment Xixun Lin et.al. 2604.13630 null
2026-04-15 Golden Handcuffs make safer AI agents Aram Ebtekar et.al. 2604.13609 null
2026-04-15 WebMAC: A Multi-Agent Collaborative Framework for Scenario Testing of Web Systems Zhenyu Wan et.al. 2604.13559 null
2026-04-15 AgentComm: Semantic Communication for Embodied Agents Peiwen Jiang et.al. 2604.13558 null
2026-04-15 Don’t Let AI Agents YOLO Your Files: Shifting Information and Control to Filesystems for Agent Safety and Autonomy Shawn et.al. 2604.13536 null
2026-04-13 ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection Wei Zhao et.al. 2604.11790 null
2026-04-13 ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents Fei Tang et.al. 2604.11784 null
2026-04-13 $λ_A$ : A Typed Lambda Calculus for LLM Agent Composition Qin Liu et.al. 2604.11767 null
2026-04-13 Retrieval Is Not Enough: Why Organizational AI Needs Epistemic Infrastructure Federico Bottino et.al. 2604.11759 null
2026-04-13 Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games Keyang Zhong et.al. 2604.11741 null
2026-04-13 SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context Shuquan Lian et.al. 2604.11716 null
2026-04-13 Agentic Driving Coach: Robustness and Determinism of Agentic AI-Powered Human-in-the-Loop Cyber-Physical Systems Deeksha Prahlad et.al. 2604.11705 null
2026-04-13 Towards Autonomous Mechanistic Reasoning in Virtual Cells Yunhui Jang et.al. 2604.11661 null
2026-04-13 CodeTracer: Towards Traceable Agent States Han Li et.al. 2604.11641 null
2026-04-13 Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo Artem Gadzhiev et.al. 2604.11563 null
2026-04-13 UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents Yijuan Liang et.al. 2604.11557 null
2026-04-13 Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Liujie Zhang et.al. 2604.11554 null
2026-04-13 SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering Ningyan Zhu et.al. 2604.11548 null
2026-04-13 Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory Weixian Waylon Li et.al. 2604.11544 null
2026-04-13 Problem Reductions at Scale: Agentic Integration of Computationally Hard Problems Xi-Wei Pan et.al. 2604.11535 null
2026-04-13 PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints Minjun Park et.al. 2604.11523 null
2026-04-13 From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python Jinhua Wang et.al. 2604.11518 null
2026-04-13 OOM-RL: Out-of-Money Reinforcement Learning Market-Driven Alignment for LLM-Based Multi-Agent Systems Kun Liu et.al. 2604.11477 null
2026-04-13 SLALOM: Simulation Lifecycle Analysis via Longitudinal Observation Metrics for Social Simulation Juhoon Lee et.al. 2604.11466 null
2026-04-13 Three Roles, One Model: Role Orchestration at Inference Time to Close the Performance Gap Between Small and Large Agents S. Aaron McClendon et.al. 2604.11465 null
2026-04-12 Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training? Wanyi Chen et.al. 2604.10547 null
2026-04-12 Structure-Grounded Knowledge Retrieval via Code Dependencies for Multi-Step Data Reasoning Xinyi Huang et.al. 2604.10516 null
2026-04-12 Agent Mentor: Framing Agent Knowledge through Semantic Trajectory Analysis Roi Ben-Gigi et.al. 2604.10513 null
2026-04-12 Cooperation in Human and Machine Agents: Promise Theory Considerations M. Burgess et.al. 2604.10505 null
2026-04-12 SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents Mahir Labib Dihan et.al. 2604.10493 null
2026-04-12 Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs Yu Li et.al. 2604.10480 null
2026-04-12 From Query to Counsel: Structured Reasoning with a Multi-Agent Framework and Dataset for Legal Consultation Mingfei Lu et.al. 2604.10470 null
2026-04-12 TrajOnco: a multi-agent framework for temporal reasoning over longitudinal EHR for multi-cancer early detection Sihang Zeng et.al. 2604.10386 null
2026-04-11 Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning Nicolae Cudlenco et.al. 2604.10383 null
2026-04-11 ClawVM: Harness-Managed Virtual Memory for Stateful Tool-Using LLM Agents Mofasshara Rafique et.al. 2604.10352 null
2026-04-11 WaterAdmin: Orchestrating Community Water Distribution Optimization via AI Agents Jiaqi Wen et.al. 2604.10343 null
2026-04-11 From Helpful to Trustworthy: LLM Agents for Pair Programming Ragib Shahariar Ayon et.al. 2604.10300 null
2026-04-11 TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale Malgorzata Gwiazda et.al. 2604.10291 null
2026-04-11 AI Organizations are More Effective but Less Aligned than Individual Agents Judy Hanwen Shen et.al. 2604.10290 null
2026-04-11 The Amazing Agent Race: Strong Tool Users, Weak Navigators Zae Myung Kim et.al. 2604.10261 null
2026-04-11 Credit-Budgeted ICPC-Style Coding: When Agents Must Pay for Every Decision Lingfeng Zhou et.al. 2604.10182 null
2026-04-11 From Speech to Profile: A Protocol-Driven LLM Agent for Psychological Profile Generation Xingjian Yang et.al. 2604.10161 null
2026-04-11 ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based Clarification Zhensheng Wang et.al. 2604.10159 null
2026-04-11 PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification Guangyu Gong et.al. 2604.10134 null
2026-04-11 HARPO: Hierarchical Agentic Reasoning for User-Aligned Conversational Recommendation Subham Raj et.al. 2604.10048 null
2026-04-10 Semantic Rate-Distortion for Bounded Multi-Agent Communication: Capacity-Derived Semantic Spaces and the Communication Cost of Alignment Anthony T. Nixon et.al. 2604.09521 null
2026-04-10 VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning Yucheng Shen et.al. 2604.09508 null
2026-04-10 Strategic Algorithmic Monoculture:Experimental Evidence from Coordination Games Gonzalo Ballestero et.al. 2604.09502 null
2026-04-10 From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models Chenchen Zhang et.al. 2604.09459 null
2026-04-10 E3-TIR: Enhanced Experience Exploitation for Tool-Integrated Reasoning Weiyang Guo et.al. 2604.09455 null
2026-04-10 Many-Tier Instruction Hierarchy in LLM Agents Jingyu Zhang et.al. 2604.09443 null
2026-04-10 Three Modalities, Two Design Probes, One Prototype, and No Vision: Experience-Based Co-Design of a Multi-modal 3D Data Visualization Tool Sanchita S. Kamath et.al. 2604.09426 null
2026-04-10 Do We Really Need to Approach the Entire Pareto Front in Many-Objective Bayesian Optimisation? Chao Jiang et.al. 2604.09417 null
2026-04-10 Do AI Coding Agents Log Like Humans? An Empirical Study Youssef Esseddiq Ouatiti et.al. 2604.09409 null
2026-04-10 HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help? Mohamed Elfeki et.al. 2604.09408 null
2026-04-10 Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation Lingfeng Huang et.al. 2604.09368 null
2026-04-10 SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering Jingzhi Gong et.al. 2604.09297 null
2026-04-10 Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation Haobo Hu et.al. 2604.09195 null
2026-04-10 Order structure and signalling in higher order quantum maps Anna Jenčová et.al. 2604.09192 null
2026-04-10 MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding Henry Zheng et.al. 2604.09167 null
2026-04-10 Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition Peng Wang et.al. 2604.09121 null
2026-04-10 Hierarchical Alignment: Enforcing Hierarchical Instruction-Following in LLMs through Logical Consistency Shu Yang et.al. 2604.09075 null
2026-04-10 V-CAGE: Vision-Closed-Loop Agentic Generation Engine for Robotic Manipulation Yaru Liu et.al. 2604.09036 null
2026-04-10 Generative AI Agent Empowered Power Allocation for HAP Propulsion and Communication Systems Xiaoyu Xing et.al. 2604.09015 null
2026-04-10 ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning Shifeng Liu et.al. 2604.08990 null
2026-04-09 Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models Shilin Yan et.al. 2604.08545 null
2026-04-09 ParseBench: A Document Parsing Benchmark for AI Agents Boyang Zhang et.al. 2604.08538 null
2026-04-09 PSI: Shared State as the Missing Layer for Coherent AI-Generated Instruments in Personal AI Agents Zhiyuan Wang et.al. 2604.08529 null
2026-04-09 ClawBench: Can AI Agents Complete Everyday Online Tasks? Yuxuan Zhang et.al. 2604.08523 null
2026-04-09 MolmoWeb: Open Visual Web Agent and Open Data for the Open Web Tanmay Gupta et.al. 2604.08516 null
2026-04-09 Figures as Interfaces: Toward LLM-Native Artifacts for Scientific Discovery Yifang Wang et.al. 2604.08491 null
2026-04-09 Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain Hanzhi Liu et.al. 2604.08407 null
2026-04-09 Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing Wenhao Yuan et.al. 2604.08401 null
2026-04-09 Awakening the Sleeping Agent: Lean-Specific Agentic Data Reactivates General Tool Use in Goedel Prover Jui-Hui Chung et.al. 2604.08388 null
2026-04-09 SkillClaw: Let Skills Evolve Collectively with Agentic Evolver Ziyu Ma et.al. 2604.08377 null
2026-04-09 Don’t Overthink It: Inter-Rollout Action Agreement as a Free Adaptive-Compute Signal for LLM Agents Khushal Sethi et.al. 2604.08369 null
2026-04-09 A Model Context Protocol Server for Quantum Execution in Hybrid Quantum-HPC Environments Masaki Shiraishi et.al. 2604.08318 null
2026-04-09 ACF: A Collaborative Framework for Agent Covert Communication under Cognitive Asymmetry Wansheng Wu et.al. 2604.08276 null
2026-04-09 Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework Seyed Amir Ahmad Safavi-Naini et.al. 2604.08226 null
2026-04-09 Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Chenyu Zhou et.al. 2604.08224 null
2026-04-09 “Theater of Mind” for LLMs: A Cognitive Architecture Based on Global Workspace Theory Wenlong Shang et.al. 2604.08206 null
2026-04-09 Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling Jiaxuan Wang et.al. 2604.08178 null
2026-04-09 Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning Teng Pang et.al. 2604.08174 null
2026-04-09 Multimodal Latent Reasoning via Predictive Embeddings Ashutosh Adhikari et.al. 2604.08065 null
2026-04-09 ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models Chonghan Qin et.al. 2604.08064 null
2026-04-09 Governed Capability Evolution for Embodied Agents: Safe Upgrade, Compatibility Checking, and Runtime Rollback for Embodied Capability Modules Xue Qin et.al. 2604.08059 null
2026-04-09 PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory Zhifei Xie et.al. 2604.08000 null
2026-04-09 Lighting-grounded Video Generation with Renderer-based Agent Reasoning Ziqi Cai et.al. 2604.07966 null
2026-04-09 TOOLCAD: Exploring Tool-Using Large Language Models in Text-to-CAD Generation with Reinforcement Learning Yifei Gong et.al. 2604.07960 null
2026-04-09 EigentSearch-Q+: Enhancing Deep Research Agents with Structured Reasoning Tools Boer Zhang et.al. 2604.07927 null
2026-04-09 Dynamic Attentional Context Scoping: Agent-Triggered Focus Sessions for Isolated Per-Agent Steering in Multi-Agent LLM Orchestration Nickson Patel et.al. 2604.07911 null
2026-04-09 MemReader: From Passive to Active Extraction for Long-Term Agent Memory Jingyi Kang et.al. 2604.07877 null
2026-04-09 Object-Attribute-Relation Model Driven Adaptive Hierarchical Transmission for Multimodal Semantic Communication Chenxing Li et.al. 2604.07859 null
2026-04-09 We Need Strong Preconditions For Using Simulations In Policy Steven Luo et.al. 2604.07838 null
2026-04-09 Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution Xue Qin et.al. 2604.07833 null
2026-04-09 More Capable, Less Cooperative? When LLMs Fail At Zero-Cost Collaboration Advait Yadav et.al. 2604.07821 null
2026-04-08 ReCodeAgent: A Multi-Agent Workflow for Language-agnostic Translation and Validation of Large-scale Repositories Ali Reza Ibrahimzada et.al. 2604.07341 null
2026-04-08 TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories Yen-Shan Chen et.al. 2604.07223 null
2026-04-08 Agent-Driven Corpus Linguistics: A Framework for Autonomous Linguistic Discovery Jia Yu et.al. 2604.07189 null
2026-04-08 Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering Zhuohong Chen et.al. 2604.07146 null
2026-04-08 AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views Minh Tam Pham et.al. 2604.07041 null
2026-04-08 ReDAct: Uncertainty-Aware Deferral for LLM Agents Dzianis Piatrashyn et.al. 2604.07036 null
2026-04-08 Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argumentation Philipp D. Siedler et.al. 2604.07028 null
2026-04-08 AgentCity: Constitutional Governance for Autonomous Agent Economies via Separation of Power Anbang Ruan et.al. 2604.07007 null
2026-04-08 EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration Yunbo Long et.al. 2604.07003 null
2026-04-08 LungCURE: Benchmarking Multimodal Real-World Clinical Reasoning for Precision Lung Cancer Diagnosis and Treatment Fangyu Hao et.al. 2604.06925 null
2026-04-08 REAgent: Requirement-Driven LLM Agents for Software Issue Resolution Shiqi Kuang et.al. 2604.06861 null
2026-04-08 From Perception to Autonomous Computational Modeling: A Multi-Agent Approach Daniel N. Wilke et.al. 2604.06788 null
2026-04-08 Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Wenhao Yang et.al. 2604.06777 null
2026-04-08 Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents Heng Zhou et.al. 2604.06753 null
2026-04-08 TurboAgent: An LLM-Driven Autonomous Multi-Agent Framework for Turbomachinery Aerodynamic Design Juan Du et.al. 2604.06747 null
2026-04-08 AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents Yujun Cheng et.al. 2604.06696 null
2026-04-08 Aegon: Auditable AI Content Access with Ledger-Bound Tokens and Hardware-Attested Mobile Receipts Amrish Baskaran et.al. 2604.06693 null
2026-04-08 When Agent Markets Arrive Xuan Liu et.al. 2604.06688 null
2026-04-08 Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Minxiao Li et.al. 2604.06683 null
2026-04-08 Argus: Reorchestrating Static Analysis via a Multi-Agent Ensemble for Full-Chain Security Vulnerability Detection Zi Liang et.al. 2604.06633 null
2026-04-08 CCD-CBT: Multi-Agent Therapeutic Interaction for CBT Guided by Cognitive Conceptualization Diagram Chang Liu et.al. 2604.06551 null
2026-04-07 Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems Operating Across Enterprise and Geopolitical Boundaries Andrew Kurtz et.al. 2604.06148 null
2026-04-07 Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents Bowen Ye et.al. 2604.06132 null
2026-04-07 Gym-Anything: Turn any Software into an Agent Environment Pranjal Aggarwal et.al. 2604.06126 null
2026-04-07 ACE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments Wang Yang et.al. 2604.06111 null
2026-04-07 Artificial Intelligence and the Structure of Mathematics Maissam Barkeshli et.al. 2604.06107 null
2026-04-07 Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives Changgeon Ko et.al. 2604.06091 null
2026-04-07 gyaradax: Local Gyrokinetics JAX Code Gianluca Galletti et.al. 2604.06085 null
2026-04-07 CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments Gustav Keppler et.al. 2604.06019 null
2026-04-07 Epistemic Blinding: An Inference-Time Protocol for Auditing Prior Contamination in LLM-Assisted Analysis Michael Cuccarese et.al. 2604.06013 null
2026-04-07 Flowr – Scaling Up Retail Supply Chain Operations Through Agentic AI in Large Scale Supermarket Chains Eranga Bandara et.al. 2604.05987 null
2026-04-07 A Formal Security Framework for MCP-Based AI Agents: Threat Taxonomy, Verification Models, and Defense Mechanisms Nirajan Acharya et.al. 2604.05969 null
2026-04-07 FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures Fan Zhang et.al. 2604.05966 null
2026-04-07 Joint Knowledge Base Completion and Question Answering by Combining Large Language Models and Small Language Models Yinan Liu et.al. 2604.05875 null
2026-04-07 Deep Researcher Agent: An Autonomous Framework for 24/7 Deep Learning Experimentation with Zero-Cost Monitoring Xiangyue Zhang et.al. 2604.05854 null
2026-04-07 Evaluating Learner Representations for Differentiation Prior to Instructional Outcomes Junsoo Park et.al. 2604.05848 null
2026-04-07 AgentGL: Towards Agentic Graph Learning with LLMs via Reinforcement Learning Yuanfu Sun et.al. 2604.05846 null
2026-04-07 Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents Shuai Zhen et.al. 2604.05808 null
2026-04-07 LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo Ojas Jain et.al. 2604.05681 null
2026-04-07 Rectified Schrödinger Bridge Matching for Few-Step Visual Navigation Wuyang Luan et.al. 2604.05673 null
2026-04-07 Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming Baoshun Tong et.al. 2604.05595 null
2026-04-07 Beyond Tools and Persons: Who Are They? Classifying Robots and AI Agents for Proportional Governance Huansheng Ning et.al. 2604.05568 null
2026-04-07 Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents Yanxu Mao et.al. 2604.05549 null
2026-04-07 Experience Transfer for Multimodal LLM Agents in Minecraft Game Chenghao Li et.al. 2604.05533 null
2026-04-07 ActivityEditor: Learning to Synthesize Physically Valid Human Mobility Chenjie Yang et.al. 2604.05529 null
2026-04-07 Market-Bench: Benchmarking Large Language Models on Economic and Trade Competition Yushuo Zheng et.al. 2604.05523 null
2026-04-07 Auditable Agents Yi Nian et.al. 2604.05485 null
2026-04-07 CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment Li Kang et.al. 2604.05484 null
2026-04-07 MA-IDS: Multi-Agent RAG Framework for IoT Network Intrusion Detection with an Experience Library Md Shamimul Islam et.al. 2604.05458 null
2026-04-07 Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use Wuyang Zhang et.al. 2604.05432 null
2026-04-07 CODESTRUCT: Code Agents over Structured Action Spaces Myeongsoo Kim et.al. 2604.05407 null
2026-04-07 Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning Qisheng Su et.al. 2604.05404 null
2026-04-07 Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA Xing Tang et.al. 2604.05387 null
2026-04-05 Schema-Aware Planning and Hybrid Knowledge Toolset for Reliable Knowledge Graph Triple Verification Xinyan Ma et.al. 2604.04190 null
2026-04-05 Readable Minds: Emergent Theory-of-Mind-Like Behavior in LLM Poker Agents Hsieh-Ting Lin et.al. 2604.04157 null
2026-04-05 Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation Peixin Chen et.al. 2604.04108 null
2026-04-05 BAAI Cardiac Agent: An intelligent multimodal agent for automated reasoning and diagnosis of cardiovascular diseases from cardiac magnetic resonance imaging Taiping Qu et.al. 2604.04078 null
2026-04-05 Humans Integrate, Agents Fix: How Agent-Authored Pull Requests Are Referenced in Practice Islem Khemissi et.al. 2604.04059 null
2026-04-05 Causality Laundering: Denial-Feedback Leakage in Tool-Calling LLM Agents Mohammad Hossein Chinaei et.al. 2604.04035 null
2026-04-05 GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces Xinyu Geng et.al. 2604.04017 null
2026-04-05 Quantifying Trust: Financial Risk Management for Trustworthy AI Agents Wenyue Hua et.al. 2604.03976 null
2026-04-05 TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol Khanh Linh Nguyen et.al. 2604.03968 null
2026-04-05 SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources Shuaike Shen et.al. 2604.03964 null
2026-04-05 Symbolic-Vector Attention Fusion for Collective Intelligence Hongwei Xu et.al. 2604.03955 null
2026-04-05 Deploy, Calibrate, Monitor, Heal – No Human Required: An Autonomous AI SRE Agent for Elasticsearch Muhamed Ramees Cheriya Mukkolakkal et.al. 2604.03933 null
2026-04-05 Reimagining RAN Automation in 6G: An Agentic AI Framework with Hierarchical Online Decision Transformer Md Arafat Habib et.al. 2604.03908 null
2026-04-04 LLM-Agent-based Social Simulation for Attitude Diffusion Deepak John Reji et.al. 2604.03898 null
2026-04-04 Enhancing behavioral nudges with large language model-based iterative personalization: A field experiment on electricity and hot-water conservation Zonghan Li et.al. 2604.03881 null
2026-04-04 Your Agent is More Brittle Than You Think: Uncovering Indirect Injection Vulnerabilities in Agentic LLMs Wenhui Zhu et.al. 2604.03870 null
2026-04-04 Beyond Crash-to-Patch: Patch Evolution for Linux Kernel Repair Luyao Bai et.al. 2604.03851 null
2026-04-04 Explainability-Guided Adversarial Attacks on Transformer-Based Malware Detectors Using Control Flow Graphs Andrew Wheeler et.al. 2604.03843 null
2026-04-04 When AI Agents Disagree Like Humans: Reasoning Trace Analysis for Human-AI Collaborative Moderation Michał Wawer et.al. 2604.03796 null
2026-04-04 SoK: Blockchain Agent-to-Agent Payments Yuanzhe Zhang et.al. 2604.03733 null
2026-04-03 From Industry Claims to Empirical Reality: An Empirical Study of Code Review Agents in Pull Requests Kowshik Chowdhury et.al. 2604.03196 null
2026-04-03 Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents Delip Rao et.al. 2604.03173 null
2026-04-03 CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator Yuhan Pu et.al. 2604.03156 null
2026-04-03 A Systematic Security Evaluation of OpenClaw and Its Variants Yuhang Wang et.al. 2604.03131 null
2026-04-03 Co-Evolution of Policy and Internal Reward for Language Agents Xinyu Wang et.al. 2604.03098 null
2026-04-03 SkillRT: Compiling Skills for Efficient Execution Everywhere Le Chen et.al. 2604.03088 null
2026-04-03 Quantitative spectroscopy of single and multiple OB-type stars. Non-LTE spectrum analysis with machine learning P. Aschenbrenner et.al. 2604.03082 null
2026-04-03 Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems Yubin Qu et.al. 2604.03081 null
2026-04-03 Automatic Textbook Formalization Fabian Gloeckle et.al. 2604.03071 null
2026-04-03 Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study Zhihao Chen et.al. 2604.03070 null
2026-04-03 QVAD: A Question-Centric Agentic Framework for Efficient and Training-Free Video Anomaly Detection Lokman Bekit et.al. 2604.03040 null
2026-04-03 Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution KN Ajay Shastry et.al. 2604.03035 null
2026-04-03 InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking Ka Yiu Lee et.al. 2604.02971 null
2026-04-03 AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents Yunhao Feng et.al. 2604.02947 null
2026-04-03 ChatSVA: Bridging SVA Generation for Hardware Verification via Task-Specific LLMs Lik Tung Fu et.al. 2604.02811 null
2026-04-03 Generative AI Use in Professional Graduate Thesis Writing: Adoption, Perceived Outcomes, and the Role of a Research-Specialized Agent Kenji Saito et.al. 2604.02792 null
2026-04-03 Improving Role Consistency in Multi-Agent Collaboration via Quantitative Role Clarity Guoling Zhou et.al. 2604.02770 null
2026-04-03 OMNI-PoseX: A Fast Vision Model for 6D Object Pose Estimation in Embodied Tasks Michael Zhang et.al. 2604.02759 null
2026-04-03 Aligning Progress and Feasibility: A Neuro-Symbolic Dual Memory Framework for Long-Horizon LLM Agents Bin Wen et.al. 2604.02734 null
2026-04-03 GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning DeepReinforce Team et.al. 2604.02721 null
2026-04-02 Novel Memory Forgetting Techniques for Autonomous AI Agents: Balancing Relevance and Efficiency Payal Fofadiya et.al. 2604.02280 null
2026-04-02 The Self Driving Portfolio: Agentic Architecture for Institutional Asset Management Andrew Ang et.al. 2604.02279 null
2026-04-02 SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization Zhengxi Lu et.al. 2604.02268 null
2026-04-02 Quantifying Self-Preservation Bias in Large Language Models Matteo Migliarini et.al. 2604.02174 null
2026-04-02 Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents Xuan Qi et.al. 2604.02155 null
2026-04-02 MTI: A Behavior-Based Temperament Profiling System for AI Agents Jihoon Jeong et.al. 2604.02145 null
2026-04-02 Diff-KD: Diffusion-based Knowledge Distillation for Collaborative Perception under Corruptions Pengcheng Lyu et.al. 2604.02061 null
2026-04-02 APEX: Agent Payment Execution with Policy for Autonomous Agent API Access Mohd Safwan Uddin et.al. 2604.02023 null
2026-04-02 Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning Rafael Pardinas et.al. 2604.02007 null
2026-04-02 ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning Jingyue Gao et.al. 2604.02006 null
2026-04-02 RuleForge: Automated Generation and Validation for Web Vulnerability Detection at Scale Ayush Garg et.al. 2604.01977 null
2026-04-02 AeroTherm-GPT: A Verification-Centered LLM Framework for Thermal Protection System Engineering Workflows Chuhan Qiao et.al. 2604.01738 null
2026-04-02 EvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification Hanrong Zhang et.al. 2604.01687 null
2026-04-02 Hierarchical Memory Orchestration for Personalized Persistent Agents Junming Liu et.al. 2604.01670 null
2026-04-02 AURA: Multimodal Shared Autonomy for Real-World Urban Navigation Yukai Ma et.al. 2604.01659 null
2026-04-02 CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery Ao Qu et.al. 2604.01658 null
2026-04-02 Exploring Robust Multi-Agent Workflows for Environmental Data Management Boyuan Guan et.al. 2604.01647 null
2026-04-02 Seclens: Role-specific Evaluation of LLM’s for security vulnerablity detection Subho Halder et.al. 2604.01637 null
2026-04-02 GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation Taraneh Ghandi et.al. 2604.01610 null
2026-04-02 ByteRover: Agent-Native Memory Through LLM-Curated Hierarchical Context Andy Nguyen et.al. 2604.01599 null
2026-04-02 Read More, Think More: Revisiting Observation Reduction for Web Agents Masafumi Enomoto et.al. 2604.01535 null
2026-04-02 PHMForge: A Scenario-Driven Agentic Benchmark for Industrial Asset Lifecycle Maintenance Ayan Das et.al. 2604.01532 null
2026-04-02 ProdCodeBench: A Production-Derived Benchmark for Evaluating AI Coding Agents Smriti Jha et.al. 2604.01527 null
2026-04-01 HippoCamp: Benchmarking Contextual Agents on Personal Computers Zhe Yang et.al. 2604.01221 null
2026-04-01 $\texttt{YC-Bench}$ : Benchmarking AI Agents for Long-Term Planning and Consistent Execution Muyu He et.al. 2604.01212 null
2026-04-01 CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery Youssef Mroueh et.al. 2604.01210 null
2026-04-01 AgentWatcher: A Rule-based Prompt Injection Monitor Yanting Wang et.al. 2604.01194 null
2026-04-01 Detecting Multi-Agent Collusion Through Multi-Agent Interpretability Aaron Rose et.al. 2604.01151 null
2026-04-01 Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers Atsuyuki Miyai et.al. 2604.01128 null
2026-04-01 CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance Haochen Liu et.al. 2604.01113 null
2026-04-01 OrgAgent: Organize Your Multi-Agent System like a Company Yiru Wang et.al. 2604.01020 null
2026-04-01 AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration Ruhao Liu et.al. 2604.01014 null
2026-04-01 OmniMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory Jiaqi Liu et.al. 2604.01007 null
2026-04-01 Dual Optimal: Make Your LLM Peer-like with Dignity Xiangqi Wang et.al. 2604.00979 null
2026-04-01 A Visionary Look at Vibe Researching Yebo Feng et.al. 2604.00945 null
2026-04-01 Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time Razvan Mihai Popescu et.al. 2604.00917 null
2026-04-01 When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation Henry Peng Zou et.al. 2604.00892 null
2026-04-01 A 4D Representation for Training-Free Agentic Reasoning from Monocular Laparoscopic Video Maximilian Fehrentz et.al. 2604.00867 null
2026-04-01 Agentic Tool Use in Large Language Models Jinchao Hu et.al. 2604.00835 null
2026-04-01 Yet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMs Yang Ye et.al. 2604.00824 null
2026-04-01 UK AISI Alignment Evaluation Case-Study Alexandra Souly et.al. 2604.00788 null
2026-04-01 LangMARL: Natural Language Multi-Agent Reinforcement Learning Huaiyuan Yao et.al. 2604.00722 null
2026-04-01 GRASP: Gradient Realignment via Active Shared Perception for Multi-Agent Collaborative Optimization Sihan Zhou et.al. 2604.00717 null
2026-04-01 AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications Ruozhao Yang et.al. 2604.00704 null
2026-04-01 Internal APIs Are All You Need: Shadow APIs, Shared Discovery, and the Case Against Browser-First Agent Architectures Lewis Tham et.al. 2604.00694 null
2026-04-01 Fluently Lying: Adversarial Robustness Can Be Substrate-Dependent Daye Kang et.al. 2604.00605 null
2026-04-01 Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents Thanh Luong Tuan et.al. 2604.00555 null
2026-04-01 Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding Haibo Wang et.al. 2604.00528 null
2026-04-01 Do Agents Repair When Challenged – or Just Reply? Challenge, Repair, and Public Correction in a Deployed Agent Forum Luyang Zhang et.al. 2604.00518 null
2026-04-01 Executing as You Generate: Hiding Execution Latency in LLM Code Generation Zhensu Sun et.al. 2604.00491 null
2026-04-01 Competition and Cooperation of LLM Agents in Games Jiayi Yao et.al. 2604.00487 null
2026-04-01 The Silicon Mirror: Dynamic Behavioral Gating for Anti-Sycophancy in LLM Agents Harshee Jignesh Shah et.al. 2604.00478 null
2026-04-01 Distributed Safety-Critical Control of Multi-Agent Systems with Time-Varying Communication Topologies Shiyu Cheng et.al. 2604.00429 null
2026-03-31 The Triadic Cognitive Architecture: Bounding Autonomous Action via Spatio-Temporal and Epistemic Friction Davide Di Gioia et.al. 2603.30031 null
2026-03-31 Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks Chong Xiang et.al. 2603.30016 null
2026-03-31 SurgNavAR: An Augmented Reality Surgical Navigation Framework for Optical See-Through Head Mounted Displays Abdullah Thabit et.al. 2603.29990 null
2026-03-31 BayesInsights: Modelling Software Delivery and Developer Experience with Bayesian Networks at Bloomberg Serkan Kirbas et.al. 2603.29929 null
2026-03-31 SkillReducer: Optimizing LLM Agent Skills for Token Efficiency Yudong Gao et.al. 2603.29919 null
2026-03-31 ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation Yinuo Liu et.al. 2603.29902 null
2026-03-31 Owl-AuraID 1.0: An Intelligent System for Autonomous Scientific Instrumentation and Scientific Data Analysis Han Deng et.al. 2603.29828 null
2026-03-31 SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes Léopold Maillard et.al. 2603.29798 null
2026-03-31 CausalPulse: An Industrial-Grade Neurosymbolic Multi-Agent Copilot for Causal Diagnostics in Smart Manufacturing Chathurangi Shyalika et.al. 2603.29755 null
2026-03-31 BotVerse: Real-Time Event-Driven Simulation of Social Agents Edoardo Allegrini et.al. 2603.29741 null
2026-03-31 Latent-Y: A Lab-Validated Autonomous Agent for De Novo Drug Design Latent Labs Team et.al. 2603.29727 null
2026-03-31 Near-Miss: Latent Policy Failure Detection in Agentic Workflows Ella Rabinovich et.al. 2603.29665 null
2026-03-31 CutClaw: Agentic Hours-Long Video Editing via Music Synchronization Shifang Zhao et.al. 2603.29664 null
2026-03-31 6GAgentGym: Tool Use, Data Synthesis, and Agentic Learning for Network Management Jiao Chen et.al. 2603.29656 null
2026-03-31 ASI-Evolve: AI Accelerates AI Weixian Xu et.al. 2603.29640 null
2026-03-31 An Empirical Study of Multi-Agent Collaboration for Automated Research Yang Shen et.al. 2603.29632 null
2026-03-31 Can LLM Agents Identify Spoken Dialects like a Linguist? Tobias Bystrich et.al. 2603.29541 null
2026-03-31 MemFactory: Unified Inference & Training Framework for Agent Memory Ziliang Guo et.al. 2603.29493 null
2026-03-31 ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities Christopher Zanoli et.al. 2603.29399 null
2026-03-31 How and Why Agents Can Identify Bug-Introducing Commits Niklas Risse et.al. 2603.29378 null
2026-03-31 VACP: Visual Analytics Context Protocol Tobias Stähle et.al. 2603.29322 null
2026-03-31 Should I State or Should I Show? Aligning AI with Human Preferences Keaton Ellis et.al. 2603.29317 null
2026-03-31 Beyond pass@1: A Reliability Science Framework for Long-Horizon LLM Agents Aaditya Khanal et.al. 2603.29231 null
2026-03-31 Multi-Layered Memory Architectures for LLM Agents: An Experimental Evaluation of Long-Term Context Retention Sunil Tiwari et.al. 2603.29194 null
2026-03-31 SimMOF: AI agent for Automated MOF Simulations Jaewoong Lee et.al. 2603.29152 null
2026-03-31 Knowledge database development by large language models for countermeasures against viruses and marine toxins Hung N. Do et.al. 2603.29149 null
2026-03-31 SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Kuangshi Ai et.al. 2603.29139 null
2026-03-31 Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation? Wensu Li et.al. 2603.29121 null
2026-03-29 PRBench: End-to-end Paper Reproduction in Physics Research Shi Qiu et.al. 2603.27646 null
2026-03-29 Sci-Mind: Cognitively-Inspired Adversarial Debate for Autonomous Mathematical Modeling Ruiying Sun et.al. 2603.27584 null
2026-03-29 Structured Observation Language for Efficient and Generalizable Vision-Language Navigation Daojie Peng et.al. 2603.27577 null
2026-03-29 SPREAD: Spatial-Physical REasoning via geometry Aware Diffusion Minzhang Li et.al. 2603.27573 null
2026-03-29 RAGent: Physics-Aware Agentic Reasoning for Training-Free mmWave Human Activity Recognition Mingda Han et.al. 2603.27571 null
2026-03-29 Safer Builders, Risky Maintainers: A Comparative Study of Breaking Changes in Human vs Agentic PRs K M Ferdous et.al. 2603.27524 null
2026-03-29 A Systematic Taxonomy of Security Vulnerabilities in the OpenClaw AI Agent Framework Surada Suwansathit et.al. 2603.27517 null
2026-03-29 AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents Zhaopeng Feng et.al. 2603.27490 null
2026-03-29 Multi-Agent Dialectical Refinement for Enhanced Argument Classification Jakub Bąba et.al. 2603.27451 null
2026-03-28 From Tool to Teammate: LLM Coding Agents as Collaborative Partners for Behavioral Labeling in Educational Dialogue Analysis Eason Chen et.al. 2603.27440 null
2026-03-28 The Novelty Bottleneck: A Framework for Understanding Human Effort Scaling in AI-Assisted Work Jacky Liang et.al. 2603.27438 null
2026-03-28 Greedy Is a Strong Default: Agents as Iterative Optimizers Yitao Li et.al. 2603.27415 null
2026-03-28 Heterogeneous Debate Engine: Identity-Grounded Cognitive Architecture for Resilient LLM-Based Ethical Tutoring Jakub Masłowski et.al. 2603.27404 null
2026-03-28 Beyond Completion: Probing Cumulative State Tracking to Predict LLM Agent Performance Dengzhe Hou et.al. 2603.27343 null
2026-03-28 GUIDE: Guided Updates for In-context Decision Evolution in LLM-Driven Spacecraft Operations Alejandro Carrasco et.al. 2603.27306 null
2026-03-28 EpochX: Building the Infrastructure for an Emergent Agent Civilization Huacan Wang et.al. 2603.27304 null
2026-03-28 Self-evolving AI agents for protein discovery and directed evolution Yang Tan et.al. 2603.27303 null
2026-03-28 From Inference Routing to Agent Orchestration: Declarative Policy Compilation with Cross-Layer Verification Huamin Chen et.al. 2603.27299 null
2026-03-28 Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP Martin Vogel et.al. 2603.27277 null
2026-03-28 “Elementary, My Dear Watson.” Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts Shenao Wang et.al. 2603.27204 null
2026-03-26 PSDesigner: Automated Graphic Design with a Human-Like Creative Workflow Xincheng Shuai et.al. 2603.25738 null
2026-03-26 Agent Factories for High Level Synthesis: How Far Can General-Purpose Coding Agents Go in Hardware Optimization? Abhishek Bhandwaldar et.al. 2603.25719 null
2026-03-26 The Kitchen Loop: User-Spec-Driven Development for a Self-Evolving Codebase Yannick Roy et.al. 2603.25697 null
2026-03-26 Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos Abdullah Hamdi et.al. 2603.25645 null
2026-03-26 Designing Any Imaging System from Natural Language: Agent-Constrained Composition over a Finite Primitive Basis Chengshuai Yang et.al. 2603.25636 null
2026-03-26 Social Hippocampus Memory Learning Liping Yi et.al. 2603.25614 null
2026-03-26 EcoThink: A Green Adaptive Inference Framework for Sustainable and Accessible Agents Linxiao Li et.al. 2603.25498 null
2026-03-26 From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the Wild Zhi Zeng et.al. 2603.25423 null
2026-03-26 VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents George Eskandar et.al. 2603.25420 null
2026-03-26 Modernising Reinforcement Learning-Based Navigation for Embodied Semantic Scene Graph Generation Roman Kueble et.al. 2603.25415 null
2026-03-26 SafeGuard ASF: SR Agentic Humanoid Robot System for Autonomous Industrial Safety Thanh Nguyen Canh et.al. 2603.25353 null
2026-03-26 From Intent to Evidence: A Categorical Approach for Structural Evaluation of Deep Research Agents Shuoling Liu et.al. 2603.25342 null
2026-03-26 AD-CARE: A Guideline-grounded, Modality-agnostic LLM Agent for Real-world Alzheimer’s Disease Diagnosis with Multi-cohort Assessment, Fairness Analysis, and Reader Study Wenlong Hou et.al. 2603.25322 null
2026-03-26 FluxEDA: A Unified Execution Infrastructure for Stateful Agentic EDA Zhengrui Chen et.al. 2603.25243 null
2026-03-26 Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills Jingwei Ni et.al. 2603.25158 null
2026-03-26 SEVerA: Verified Synthesis of Self-Evolving Agents Debangshu Banerjee et.al. 2603.25111 null
2026-03-26 OMIND: Framework for Knowledge Grounded Finetuning and Multi-Turn Dialogue Benchmark for Mental Health LLMs Suraj Racha et.al. 2603.25105 null
2026-03-26 From Logic Monopoly to Social Contract: Separation of Power and the Institutional Foundations for Autonomous Agent Economies Anbang Ruan et.al. 2603.25100 null
2026-03-26 Large Language Models as Optimization Controllers: Adaptive Continuation for SIMP Topology Optimization Shaoliang Yang et.al. 2603.25099 null
2026-03-26 ElephantBroker: A Knowledge-Grounded Cognitive Runtime for Trustworthy AI Agents Cristian Lupascu et.al. 2603.25097 null
2026-03-25 Chameleon: Episodic Memory for Long-Horizon Robotic Manipulation Xinying Guo et.al. 2603.24576 null
2026-03-25 Infrastructure for Valuable, Tradable, and Verifiable Agent Memory Mengyuan Li et.al. 2603.24564 null
2026-03-25 The Free-Market Algorithm: Self-Organizing Optimization for Open-Ended Complex Systems Martin Jaraiz et.al. 2603.24559 null
2026-03-25 LensWalk: Agentic Video Understanding by Planning How You See in Videos Keliang Li et.al. 2603.24558 null
2026-03-25 AVO: Agentic Variation Operators for Autonomous Evolutionary Search Terry Chen et.al. 2603.24517 null
2026-03-25 Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs Alexander Panfilov et.al. 2603.24511 null
2026-03-25 Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQA John Ray B. Martinez et.al. 2603.24481 null
2026-03-25 CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents Xiangru Jian et.al. 2603.24440 null
2026-03-25 ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers Songyang Liu et.al. 2603.24414 null
2026-03-25 GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents Yunzhe Wang et.al. 2603.24329 null
2026-03-25 Towards Semantic-based Agent Communication Networks: Vision, Technologies, and Challenges Ping Zhang et.al. 2603.24328 null
2026-03-25 The Specification Gap: Coordination Failure Under Partial Knowledge in Code Agents Camilo Chacón Sartori et.al. 2603.24284 null
2026-03-25 Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning Tommaso Galliena et.al. 2603.24257 null
2026-03-25 Environment-Grounded Multi-Agent Workflow for Autonomous Penetration Testing Michael Somma et.al. 2603.24221 null
2026-03-25 Where Do Your Citations Come From? Citation-Constellation: A Free, Open-Source, No-Code, and Auditable Tool for Citation Network Decomposition with Complementary BARON and HEROCON Scores Mahbub Ul Alam et.al. 2603.24216 null
2026-03-25 CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare Akash Ghosh et.al. 2603.24157 null
2026-03-25 FinToolSyn: A forward synthesis Framework for Financial Tool-Use Dialogue Data with Dynamic Tool Retrieval Caishuang Huang et.al. 2603.24051 null
2026-03-25 ELITE: Experiential Learning and Intent-Aware Transfer for Self-improving Embodied Agents Bingqing Wei et.al. 2603.24018 null
2026-03-25 Language-Grounded Multi-Agent Planning for Personalized and Fair Participatory Urban Sensing Xusen Guo et.al. 2603.24014 null
2026-03-25 From AI Assistant to AI Scientist: Autonomous Discovery of LLM-RL Algorithms with LLM Agents Sirui Xia et.al. 2603.23951 null
2026-03-25 AnalogAgent: Self-Improving Analog Circuit Design Automation with LLM Agents Zhixuan Bao et.al. 2603.23910 null
2026-03-25 Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy Scenarios Li Ma et.al. 2603.23875 null
2026-03-25 See, Remember, Explore: A Benchmark and Baselines for Streaming Spatial Reasoning Yuxi Wei et.al. 2603.23864 null
2026-03-25 BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents Praveen Kumar Myakala et.al. 2603.23848 null
2026-03-25 VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents Yuhao Chen et.al. 2603.23840 null
2026-03-25 AI Fortune-Teller: Juxtaposing Shaman and AI to Reveal Human Agency in the Age of AI Soonho Kwon et.al. 2603.23811 null
2026-03-25 Willful Disobedience: Automatically Detecting Failures in Agentic Traces Reshabh K Sharma et.al. 2603.23806 null
2026-03-25 How are AI agents used? Evidence from 177,000 MCP tools Merlin Stein et.al. 2603.23802 null
2026-03-25 AgentRFC: Security Design Principles and Conformance Testing for Agent Protocols Shenghan Zheng et.al. 2603.23801 null
2026-03-24 Regulating AI Agents Kathrin Gardhouse et.al. 2603.23471 null
2026-03-24 Code Review Agent Benchmark Yuntong Zhang et.al. 2603.23448 null
2026-03-24 Mecha-nudges for Machines Giulio Frey et.al. 2603.23433 null
2026-03-24 Biased Error Attribution in Multi-Agent Human-AI Systems Under Delayed Feedback Teerthaa Parakh et.al. 2603.23419 null
2026-03-24 Designing Agentic AI-Based Screening for Portfolio Investment Mehmet Caner et.al. 2603.23300 null
2026-03-24 Emergence of Fragility in LLM-based Social Networks: the Case of Moltbook Luca Sodano et.al. 2603.23279 null
2026-03-24 PHANTOM Hand Teng Yan et.al. 2603.23152 null
2026-03-24 Polaris: A Gödel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair Aditya Kakade et.al. 2603.23129 null
2026-03-24 AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image Detection Yangxin Yu et.al. 2603.23115 null
2026-03-24 Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution Yechao Zhang et.al. 2603.23064 null
2026-03-24 Minibal: Balanced Game-Playing Without Opponent Modeling Quentin Cohen-Solal et.al. 2603.23059 null
2026-03-24 Knowledge Access Beats Model Size: Memory Augmented Routing for Persistent AI Agents Xunzhuo Liu et.al. 2603.23013 null
2026-03-24 PaperVoyager : Building Interactive Web with Visual Language Models Dasen Dai et.al. 2603.22999 null
2026-03-24 Privacy-Preserving EHR Data Transformation via Geometric Operators: A Human-AI Co-Design Technical Report Maolin Wang et.al. 2603.22954 null
2026-03-24 Task-Aware Positioning for Improvisational Tasks in Mobile Construction Robots via an AI Agent with Multi-LMM Modules Seongju Jang et.al. 2603.22903 null
2026-03-24 Agent-Sentry: Bounding LLM Agents via Execution Provenance Rohan Sequeira et.al. 2603.22868 null
2026-03-24 The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration Haoyuan Xu et.al. 2603.22862 null
2026-03-24 Agent Audit: A Security Analysis System for LLM Agent Applications Haiyue Zhang et.al. 2603.22853 null
2026-03-24 Empirical Comparison of Agent Communication Protocols for Task Orchestration Ivan Dobrovolskyi et.al. 2603.22823 null
2026-03-24 PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding Lirong Che et.al. 2603.22796 null
2026-03-24 Predictive Photometric Uncertainty in Gaussian Splatting for Novel View Synthesis Chamuditha Jayanga Galappaththige et.al. 2603.22786 null
2026-03-24 Can LLM Agents Generate Real-World Evidence? Evaluating Observational Studies in Medical Databases Dubai Li et.al. 2603.22767 null
2026-03-24 CIPL: A Target-Independent Framework for Channel-Inversion Privacy Leakage in Agents Tao Huang et.al. 2603.22751 null
2026-03-24 Why Database Manuals Are Not Enough: Efficient and Reliable Configuration Tuning for DBMSs via Code-Driven LLM Agents Xinyi Zhang et.al. 2603.22708 null
2026-03-23 AwesomeLit: Towards Hypothesis Generation with Agent-Supported Literature Research Zefei Xie et.al. 2603.22648 null
2026-03-23 Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling Yoshiki Masuyama et.al. 2603.22589 null
2026-03-23 UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos Gu Zhang et.al. 2603.22264 null
2026-03-23 Chimera: Latency- and Performance-Aware Multi-agent Serving for Heterogeneous LLMs Kangqi Ni et.al. 2603.22206 null
2026-03-23 Human-Inspired Pavlovian and Instrumental Learning for Autonomous Agent Navigation Jingfeng Shan et.al. 2603.22170 null
2026-03-23 Causal Evidence that Language Models use Confidence to Drive Behavior Dharshan Kumaran et.al. 2603.22161 null
2026-03-23 OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation Sijie Zhao et.al. 2603.22148 null
2026-03-23 StreamingClaw Technical Report Jiawei Chen et.al. 2603.22120 null
2026-03-23 Lemma Discovery in Agentic Program Verification Huan Zhao et.al. 2603.22114 null
2026-03-23 A Context Engineering Framework for Improving Enterprise AI Agents based on Digital-Twin MDP Xi Yang et.al. 2603.22083 null
2026-03-23 Dynamic analysis enhances issue resolution Mingwei Liu et.al. 2603.22048 null
2026-03-23 Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe Xixi Wu et.al. 2603.21972 null
2026-03-23 A Blueprint for Self-Evolving Coding Agents in Vehicle Aerodynamic Drag Prediction Jinhui Ren et.al. 2603.21698 null
2026-03-23 Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execution Traces Neelmani Vispute et.al. 2603.21692 null
2026-03-23 Strategic Infrastructure Design via Multi-Agent Congestion Games with Joint Placement and Pricing Niloofar Aminikalibar et.al. 2603.21691 null
2026-03-23 Optimizing Multi-Agent Weather Captioning via Text Gradient Descent: A Training-Free Approach with Consensus-Aware Gradient Fusion Shixu Liu et.al. 2603.21673 null
2026-03-23 Are AI-assisted Development Tools Immune to Prompt Injection? Charoes Huang et.al. 2603.21642 null
2026-03-23 EnterpriseLab: A Full-Stack Platform for developing and deploying agents in Enterprises Ankush Agarwal et.al. 2603.21630 null
2026-03-23 AgenticRec: End-to-End Tool-Integrated Policy Optimization for Ranking-Oriented Recommender Agents Tianyi Li et.al. 2603.21613 null
2026-03-23 Mind over Space: Can Multimodal Large Language Models Mentally Navigate? Qihui Zhu et.al. 2603.21577 null
2026-03-23 Adaptive Robust Estimator for Multi-Agent Reinforcement Learning Zhongyi Li et.al. 2603.21574 null
2026-03-23 Counterfactual Credit Policy Optimization for Multi-Agent Collaboration Zhongyi Li et.al. 2603.21563 null
2026-03-22 LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning Jianing Wang et.al. 2603.21065 null
2026-03-22 KLDrive: Fine-Grained 3D Scene Reasoning for Autonomous Driving based on Knowledge Graph Ye Tian et.al. 2603.21029 null
2026-03-22 SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration Zihan Guo et.al. 2603.21019 null
2026-03-22 AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation Sukriti Manna et.al. 2603.20986 null
2026-03-21 Detection of adversarial intent in Human-AI teams using LLMs Abed K. Musaffar et.al. 2603.20976 null
2026-03-21 DiscoUQ: Structured Disagreement Analysis for Uncertainty Quantification in LLM Agent Ensembles Bo Jiang et.al. 2603.20975 null
2026-03-21 Learning to Aggregate Zero-Shot LLM Agents for Corporate Disclosure Classification Kemal Kirtac et.al. 2603.20965 null
2026-03-21 Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents Uchi Uchibeke et.al. 2603.20953 null
2026-03-21 User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction Yuren Hao et.al. 2603.20939 null
2026-03-21 AC4A: Access Control for Agents Reshabh K Sharma et.al. 2603.20933 null
2026-03-21 Active Inference for Physical AI Agents – An Engineering Perspective Bert de Vries et.al. 2603.20927 null
2026-03-21 Governance-Aware Vector Subscriptions for Multi-Agent Knowledge Ecosystems Steven Johnson et.al. 2603.20833 null
2026-03-21 GMPilot: An Expert AI Agent For FDA cGMP Compliance Xiaohan Wang et.al. 2603.20815 null
2026-03-21 Agentic Physical-AI for Self-Aware RF Systems Linuka Ratnayake et.al. 2603.20692 null
2026-03-21 Towards Intelligent Geospatial Data Discovery: a knowledge graph-driven multi-agent framework powered by large language models Ruixiang Liu et.al. 2603.20670 null
2026-03-21 ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework Guanzhou Chen et.al. 2603.20644 null
2026-03-21 Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention Manh Nguyen et.al. 2603.20640 null
2026-03-21 AEGIS: From Clues to Verdicts – Graph-Guided Deep Vulnerability Reasoning via Dialectics and Meta-Auditing Sen Fang et.al. 2603.20637 null
2026-03-21 Seed1.8 Model Card: Towards Generalized Real-World Agency Bytedance Seed et.al. 2603.20633 null
2026-03-21 ACRFence: Preventing Semantic Rollback Attacks in Agent Checkpoint-Restore Yusheng Zheng et.al. 2603.20625 null
2026-03-20 AI Agents Can Already Autonomously Perform Experimental High Energy Physics Eric A. Moreno et.al. 2603.20179 null
2026-03-20 Design-OS: A Specification-Driven Framework for Engineering System Design with a Control-Systems Design Case H. Sinan Bank et.al. 2603.20151 null
2026-03-20 Can Large Multimodal Models Inspect Buildings? A Hierarchical Benchmark for Structural Pathology Reasoning Hui Zhong et.al. 2603.20148 null
2026-03-20 Synergistic Perception and Generative Recomposition: A Multi-Agent Orchestration for Expert-Level Building Inspection Hui Zhong et.al. 2603.20143 null
2026-03-20 Reasoning Gets Harder for LLMs Inside A Dialogue Ivan Kartáč et.al. 2603.20133 null
2026-03-20 Revisiting Gene Ontology Knowledge Discovery with Hierarchical Feature Selection and Virtual Study Group of AI Agents Cen Wan et.al. 2603.20132 null
2026-03-20 Agentic Harness for Real-World Compilers Yingwei Zheng et.al. 2603.20075 null
2026-03-20 Orchestrating Human-AI Software Delivery: A Retrospective Longitudinal Field Study of Three Software Modernization Programs Maximiliano Armesto et.al. 2603.20028 null
2026-03-20 RouterKGQA: Specialized–General Model Routing for Constraint-Aware Knowledge Graph Question Answering Bo Yuan et.al. 2603.20017 null
2026-03-20 ReViSQL: Achieving Human-Level Text-to-SQL Yuxuan Zhu et.al. 2603.20004 null
2026-03-20 An Agentic Approach to Generating XAI-Narratives Yifan He et.al. 2603.20003 null
2026-03-20 Trojan’s Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance Fazhong Liu et.al. 2603.19974 null
2026-03-20 Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents Luiz C. Borro et.al. 2603.19935 null
2026-03-20 Robust Beam Codebooks for mmWave/THz Systems: Toward a Stochastic RL Approach Anouar Nechi et.al. 2603.19930 null
2026-03-20 DALI: LLM-Agent Enhanced Dual-Stream Adaptive Leadership Identification for Group Recommendations Boxun Song et.al. 2603.19909 null
2026-03-20 Utility-Guided Agent Orchestration for Efficient LLM Tool Use Boyan Liu et.al. 2603.19896 null
2026-03-20 Beyond detection: cooperative multi-agent reasoning for rapid onboard EO crisis response Alejandro D. Mousist et.al. 2603.19858 null
2026-03-20 Borderless Long Speech Synthesis Xingchen Song et.al. 2603.19798 null
2026-03-20 Text-Based Personas for Simulating User Privacy Decisions Kassem Fawaz et.al. 2603.19791 null
2026-03-20 Embodied Science: Closing the Discovery Loop with Agentic Embodied AI Xiang Zhuang et.al. 2603.19782 null
2026-03-20 A Subgoal-driven Framework for Improving Long-Horizon LLM Agents Taiyi Wang et.al. 2603.19685 null
2026-03-20 GoAgent: Group-of-Agents Communication Topology Generation for LLM-based Multi-Agent Systems Hongjiang Chen et.al. 2603.19677 null
2026-03-20 Semantic Audio-Visual Navigation in Continuous Environments Yichen Zeng et.al. 2603.19660 null
2026-03-20 HyEvo: Self-Evolving Hybrid Agentic Workflows for Efficient Reasoning Beibei Xu et.al. 2603.19639 null
2026-03-20 PowerLens: Taming LLM Agents for Safe and Personalized Mobile Power Management Xingyu Feng et.al. 2603.19584 null
2026-03-20 Skilled AI Agents for Embedded and IoT Systems Development Yiming Li et.al. 2603.19583 null
2026-03-19 A Framework for Formalizing LLM Agent Security Vincent Siu et.al. 2603.19469 null
2026-03-19 Markov Potential Game and Multi-Agent Reinforcement Learning for Autonomous Driving Huiwen Yan et.al. 2603.19188 null
2026-03-19 Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation Swagat Padhan et.al. 2603.19166 null
2026-03-19 From Inference Efficiency to Embodied Efficiency: Revisiting Efficiency Metrics for Vision-Language-Action Models Zhuofan Li et.al. 2603.19131 null
2026-03-19 SignAgent: Agentic LLMs for Linguistically-Grounded Sign Language Annotation and Dataset Curation Oliver Cory et.al. 2603.19059 null
2026-03-19 The Simplicity of the Hodge Bundle Anand Patel et.al. 2603.19052 null
2026-03-19 LLMs Aren’t Human: A Critical Perspective on LLM Personality Kim Zierahn et.al. 2603.19030 null
2026-03-19 Security awareness in LLM agents: the NDAI zone case Enrico Bottazzi et.al. 2603.19011 null
2026-03-19 AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science An Luo et.al. 2603.19005 null
2026-03-19 Agentic Business Process Management: A Research Manifesto Diego Calvanese et.al. 2603.18916 null
2026-03-19 Security, privacy, and agentic AI in a regulatory view: From definitions and distinctions to provisions and reflections Shiliang Zhang et.al. 2603.18914 null
2026-03-19 Act While Thinking: Accelerating LLM Agents via Pattern-Aware Speculative Tool Execution Yifan Sui et.al. 2603.18897 null
2026-03-19 I Can’t Believe It’s Corrupt: Evaluating Corruption in Multi-Agent Governance Systems Vedanta S P et.al. 2603.18894 null
2026-03-19 RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models Xiao Feng et.al. 2603.18859 null
2026-03-19 Agent Control Protocol: Admission Control for Agent Actions Marcelo Fernandez et.al. 2603.18829 null
2026-03-19 ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents Hao Zhang et.al. 2603.18815 null
2026-03-19 Mi:dm K 2.5 Pro KT Tech innovation Group et.al. 2603.18788 null
2026-03-19 ClawTrap: A MITM-Based Red-Teaming Framework for Real-World OpenClaw Security Evaluation Haochen Zhao et.al. 2603.18762 null
2026-03-19 Memento-Skills: Let Agents Design Agents Huichi Zhou et.al. 2603.18743 null
2026-03-19 Measuring and Exploiting Confirmation Bias in LLM-Assisted Security Code Review Dimitris Mitropoulos et.al. 2603.18740 null
2026-03-19 MemMA: Coordinating the Memory Cycle through Multi-Agent Reasoning and In-Situ Self-Evolution Minhua Lin et.al. 2603.18718 null
2026-03-19 A more accurate rational non-commutative algorithm for multiplying 4x4 matrices using 48 multiplications Jean-Guillaume Dumas et.al. 2603.18699 null
2026-03-19 Multimodal Model for Computational Pathology:Representation Learning and Image Compression Peihang Wu et.al. 2603.18660 null
2026-03-19 D-Mem: A Dual-Process Memory System for LLM Agents Zhixing You et.al. 2603.18631 null
2026-03-19 ZEBRAARENA: A Diagnostic Simulation Environment for Studying Reasoning-Action Coupling in Tool-Augmented LLMs Wanjia Zhao et.al. 2603.18614 null
2026-03-19 Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provably Enoch Hyunwook Kang et.al. 2603.18563 null
2026-03-19 Robotic Agentic Platform for Intelligent Electric Vehicle Disassembly Zachary Allen et.al. 2603.18520 null
2026-03-19 CyberJustice Tutor: An Agentic AI Framework for Cybersecurity Learning via Think-Plan-Act Reasoning and Pedagogical Scaffolding Baiqiang Wang et.al. 2603.18470 null
2026-03-19 SODIUM: From Open Web Data to Queryable Databases Chuxuan Hu et.al. 2603.18447 null
2026-03-19 TopoChunker: Topology-Aware Agentic Document Chunking Framework Xiaoyu Liu et.al. 2603.18409 null
2026-03-19 From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents Myeongseob Ko et.al. 2603.18382 null
2026-03-19 PlanTwin: Privacy-Preserving Planning Abstractions for Cloud-Assisted LLM Agents Guangsheng Yu et.al. 2603.18377 null
2026-03-19 Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care Shivam Shukla et.al. 2603.18375 null
2026-03-18 TDAD: Test-Driven Agentic Development - Reducing Code Regressions in AI Coding Agents via Graph-Based Impact Analysis Pepe Alonso et.al. 2603.17973 null
2026-03-18 Interpretable Traffic Responsibility from Dashcam Video via Legal Multi Agent Reasoning Jingchun Yang et.al. 2603.17930 null
2026-03-18 Differential Privacy in Generative AI Agents: Analysis and Optimal Tradeoffs Ya-Ting Yang et.al. 2603.17902 null
2026-03-18 ArchBench: Benchmarking Generative-AI for Software Architecture Tasks Bassam Adnan et.al. 2603.17833 null
2026-03-18 RPMS: Enhancing LLM-Based Embodied Planning through Rule-Augmented Memory Synergy Zhenhang Yuan et.al. 2603.17831 null
2026-03-18 CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents Lintang Sutawika et.al. 2603.17829 null
2026-03-18 Governed Memory: A Production Architecture for Multi-Agent Workflows Hamed Taheri et.al. 2603.17787 null
2026-03-18 MALLES: A Multi-agent LLMs-based Economic Sandbox with Consumer Preference Alignment Yusen Wu et.al. 2603.17694 null
2026-03-18 Can Blindfolded LLMs Still Trade? An Anonymization-First Framework for Portfolio Optimization Joohyoung Jeon et.al. 2603.17692 null
2026-03-18 Sensi: Learn One Thing at a Time – Curriculum-Based Test-Time Learning for LLM Game Agents Mohsen Arjmandi et.al. 2603.17683 null
2026-03-18 Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards Philipp Normann et.al. 2603.17673 null
2026-03-18 AgentVLN: Towards Agentic Vision-and-Language Navigation Zihao Xin et.al. 2603.17670 null
2026-03-18 VeriGrey: Greybox Agent Validation Yuntong Zhang et.al. 2603.17639 null
2026-03-18 Complementary Reinforcement Learning Dilxat Muhtar et.al. 2603.17621 null
2026-03-18 VeriAgent: A Tool-Integrated Multi-Agent System with Evolving Memory for PPA-Aware RTL Code Generation Yaoxiang Wang et.al. 2603.17613 null
2026-03-18 In Trust We Survive: Emergent Trust Learning Qianpu Chen et.al. 2603.17564 null
2026-03-18 DustNET: enabling machine learning and AI models of dusty plasmas Zhehui Wang et.al. 2603.17493 null
2026-03-18 Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare Saikat Maiti et.al. 2603.17419 null
2026-03-18 Is Your LLM-as-a-Recommender Agent Trustable? LLMs’ Recommendation is Easily Hacked by Biases (Preferences) Zichen Tang et.al. 2603.17417 null
2026-03-18 Bootstrapping Coding Agents: The Specification Is the Program Martin Monperrus et.al. 2603.17399 null
2026-03-18 Agentic Cognitive Profiling: Realigning Automated Alzheimer’s Disease Detection with Clinical Construct Validity Jiawen Kang et.al. 2603.17392 null
2026-03-18 An Auditable AI Agent Loop for Empirical Economics: A Case Study in Forecast Combination Minchul Shin et.al. 2603.17381 null
2026-03-18 EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image Detection Chenyang Zhu et.al. 2603.17343 null
2026-03-18 Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures Risham Sidhu et.al. 2603.17333 null
2026-03-18 Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress Yuelin Zhang et.al. 2603.17312 null
2026-03-18 IEMAS: An Incentive-Efficiency Routing Framework for Open Agentic Web Ecosystems Hongze Liu et.al. 2603.17302 null
2026-03-18 SEAL-Tag: Self-Tag Evidence Aggregation with Probabilistic Circuits for PII-Safe Retrieval-Augmented Generation Jin Xie et.al. 2603.17292 null
2026-03-18 Graph-Native Cognitive Memory for AI Agents: Formal Belief Revision Semantics for Versioned Memory Architectures Young Bin Park et.al. 2603.17244 null
2026-03-17 AI Scientist via Synthetic Task Scaling Ziyang Cai et.al. 2603.17216 null
2026-03-17 CODMAS: A Dialectic Multi-Agent Collaborative Framework for Structured RTL Optimization Che-Ming Chang et.al. 2603.17204 null
2026-03-17 MetaClaw: Just Talk – An Agent That Meta-Learns and Evolves in the Wild Peng Xia et.al. 2603.17187 null
2026-03-17 PAuth - Precise Task-Scoped Authorization For Agents Reshabh K Sharma et.al. 2603.17170 null
2026-03-17 How Clued up are LLMs? Evaluating Multi-Step Deductive Reasoning in a Text-Based Game Environment Rebecca Ansell et.al. 2603.17169 null
2026-03-17 Intent Formalization: A Grand Challenge for Reliable Coding in the Age of AI Agents Shuvendu K. Lahiri et.al. 2603.17150 null
2026-03-17 When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents Lu Yan et.al. 2603.17104 null
2026-03-17 OpenQlaw: An Agentic AI Assistant for Analysis of 2D Quantum Materials Sankalp Pandey et.al. 2603.17043 null
2026-03-17 Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory Sahil Sen et.al. 2603.16862 null
2026-03-17 Internalizing Agency from Reflective Experience Rui Ge et.al. 2603.16843 null
2026-03-17 Learning to Present: Inverse Specification Rewards for Agentic Slide Generation Karthik Ragunath Ananda Kumar et.al. 2603.16839 null
2026-03-17 Anticipatory Planning for Multimodal AI Agents Yongyuan Liang et.al. 2603.16777 null
2026-03-17 Nonstandard Errors in AI Agents Ruijiang Gao et.al. 2603.16744 null
2026-03-17 Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure Caglar Yildirim et.al. 2603.16734 null
2026-03-17 IQuest-Coder-V1 Technical Report Jian Yang et.al. 2603.16733 null
2026-03-17 When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making Jun Liu et.al. 2603.16673 null
2026-03-17 Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation Jiawei Mao et.al. 2603.16664 null
2026-03-17 When Openclaw Agents Learn from Each Other: Insights from Emergent AI Agent Communities for Human-AI Partnership in Education Eason Chen et.al. 2603.16663 null
2026-03-17 Bio-inspired metaheuristic optimization for hierarchical architecture design of industrial control systems Ruslan Zakirzyanov et.al. 2603.16617 null
2026-03-17 Runtime Governance for AI Agents: Policies on Paths Maurits Kaptein et.al. 2603.16586 null
2026-03-17 Malicious Or Not: Adding Repository Context to Agent Skill Classification Florian Holzbauer et.al. 2603.16572 null
2026-03-17 DanceHA: A Multi-Agent Framework for Document-Level Aspect-Based Sentiment Analysis Lei Wang et.al. 2603.16546 null
2026-03-17 AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents Shannan Yan et.al. 2603.16496 null
2026-03-17 RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments Linghua Zhang et.al. 2603.16453 null
2026-03-17 TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas Ai Jian et.al. 2603.16448 null
2026-03-17 Visual Distraction Undermines Moral Reasoning in Vision-Language Models Xinyi Yang et.al. 2603.16445 null
2026-03-17 RECOVER: Robust Entity Correction via agentic Orchestration of hypothesis Variants for Evidence-based Recovery Abhishek Kumar et.al. 2603.16411 null
2026-03-17 Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits Jia Qing Yap et.al. 2603.16335 null
2026-03-17 VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents Zhengbo Zhang et.al. 2603.16289 null
2026-03-17 CoMAI: A Collaborative Multi-Agent Framework for Robust and Equitable Interview Evaluation Gengxin Sun et.al. 2603.16215 null
2026-03-17 Proactive Rejection and Grounded Execution: A Dual-Stage Intent Analysis Paradigm for Safe and Efficient AIoT Smart Homes Xinxin Jin et.al. 2603.16207 null
2026-03-17 Parametric Social Identity Injection and Diversification in Public Opinion Simulation Hexi Wang et.al. 2603.16142 null
2026-03-17 Social Simulacra in the Wild: AI Agent Communities on Moltbook Agam Goyal et.al. 2603.16128 null
2026-03-17 SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding Songcheng Cai et.al. 2603.16124 null
2026-03-17 Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective Noppanat Wadlom et.al. 2603.16104 null
2026-03-17 Occupation-Measure Mean-Field Control: Optimization over Measures and Frank-Wolfe Methods Di Yu et.al. 2603.16094 null
2026-03-17 Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation Chang Nie et.al. 2603.16086 null
2026-03-17 ARISE: Agent Reasoning with Intrinsic Skill Evolution in Hierarchical Reinforcement Learning Yu Li et.al. 2603.16060 null
2026-03-17 Enhancing Linguistic Generalization of VLA: Fine-Tuning OpenVLA via Synthetic Instruction Augmentation Dongik Shin et.al. 2603.16044 null
2026-03-17 Speak, Segment, Track, Navigate: An Interactive System for Video-Guided Skull-Base Surgery Jecia Z. Y. Mao et.al. 2603.16024 null
2026-03-17 Interpretable Context Methodology: Folder Structure as Agentic Architecture Jake Van Clief et.al. 2603.16021 null
2026-03-16 Evaluating Agentic Optimization on Large Codebases Atharva Sehgal et.al. 2603.16011 null
2026-03-16 CoDesignAI: An AI-Enabled Multi-Agent, Multi-User System for Collaborative Urban Design at the Conceptual Stage Zhaoxi Zhang et.al. 2603.16008 null
2026-03-16 From Workflow Automation to Capability Closure: A Formal Framework for Safe and Revenue-Aware Customer Service AI Cosimo Spera et.al. 2603.15978 null
2026-03-16 OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data Yuwen Du et.al. 2603.15594 null
2026-03-16 Lore: Repurposing Git Commit Messages as a Structured Knowledge Protocol for AI Coding Agents Ivan Stetsenko et.al. 2603.15566 null
2026-03-16 InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems Shaojie Shi et.al. 2603.15542 null
2026-03-16 QiboAgent: a practitioner’s guideline to open source assistants for Quantum Computing code development Lorenzo Esposito et.al. 2603.15538 null
2026-03-16 Beyond the Covariance Trap: Unlocking Generalization in Same-Subject Knowledge Editing for Large Language Models Xiyu Liu et.al. 2603.15518 null
2026-03-16 Agentic workflow enables the recovery of critical materials from complex feedstocks via selective precipitation Andrew Ritchhart et.al. 2603.15491 null
2026-03-16 Agent Lifecycle Toolkit (ALTK): Reusable Middleware Components for Robust AI Agents Zidane Wright et.al. 2603.15473 null
2026-03-16 Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents Simone Aonzo et.al. 2603.15457 null
2026-03-16 TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems Kai Wang et.al. 2603.15408 null
2026-03-16 SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering? Tingxu Han et.al. 2603.15401 null
2026-03-16 RieMind: Geometry-Grounded Spatial Agent for Scene Understanding Fernando Ropero et.al. 2603.15386 null
2026-03-16 SKILLS: Structured Knowledge Injection for LLM-Driven Telecommunications Operations Ivo Brett et.al. 2603.15372 null
2026-03-16 Brain-Inspired Graph Multi-Agent Systems for LLM Reasoning Guangfu Hao et.al. 2603.15371 null
2026-03-16 PMAx: An Agentic Framework for AI-Driven Process Mining Anton Antonov et.al. 2603.15351 null
2026-03-16 Intelligent Co-Design: An Interactive LLM Framework for Interior Spatial Design via Multi-Modal Agents Ren Jian Lim et.al. 2603.15341 null
2026-03-16 CCTU: A Benchmark for Tool Use under Complex Constraints Junjie Ye et.al. 2603.15309 null
2026-03-16 The Impact of AI-Assisted Development on Software Security: A Study of Gemini and Developer Experience Nadine Jost et.al. 2603.15298 null
2026-03-16 Evolutionary Transfer Learning for Dragonchess Jim O’Connor et.al. 2603.15297 null
2026-03-16 Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory Rongjie Jiang et.al. 2603.15280 null
2026-03-16 Mechanistic Foundations of Goal-Directed Control Alma Lago et.al. 2603.15248 null
2026-03-14 LegacyTranslate: LLM-based Multi-Agent Method for Legacy Code Translation Zahra Moti et.al. 2603.14054 null
2026-03-14 A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning Yichang Xu et.al. 2603.14052 null
2026-03-14 Sovereign-OS: A Charter-Governed Operating System for Autonomous AI Agents with Verifiable Fiscal Discipline Aojie Yuan et.al. 2603.14011 null
2026-03-14 ToolFlood: Beyond Selection – Hiding Valid Tools from LLM Agents via Semantic Covering Hussein Jawad et.al. 2603.13950 null
2026-03-14 AI Agents in Financial Markets: Architecture, Applications, and Systemic Implications Hui Gong et.al. 2603.13942 null
2026-03-14 APEX-Searcher: Augmenting LLMs’ Search Capabilities through Agentic Planning and Execution Kun Chen et.al. 2603.13853 null
2026-03-14 ClimateAgents: A Multi-Agent Research Assistant for Social-Climate Dynamics Analysis Shan Shan et.al. 2603.13840 null
2026-03-14 Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space Quoc-Huy Trinh et.al. 2603.13800 null
2026-03-14 DeceptGuard :A Constitutional Oversight Framework For Detecting Deception in LLM Agents Snehasis Mukhopadhyay et.al. 2603.13791 null
2026-03-14 LiveWeb-IE: A Benchmark For Online Web Information Extraction Seungbin Yang et.al. 2603.13773 null
2026-03-14 Retrieve, Schedule, Reflect: LLM Agents for Chip QoR Optimization Yikang ouyang et.al. 2603.13767 null
2026-03-14 Six Interventions for the Responsible and Ethical Implementation of Medical AI Agents Tom Bisson et.al. 2603.13743 null
2026-03-14 Testing with AI Agents: An Empirical Study of Test Generation Frequency, Quality, and Coverage Suzuka Yoshimoto et.al. 2603.13724 null
2026-03-14 Do AI Agents Really Improve Code Readability? Kyogo Horikawa et.al. 2603.13723 null
2026-03-14 InterventionLens: A Multi-Agent Framework for Detecting ASD Intervention Strategies in Parent-Child Shared Reading Xiao Wang et.al. 2603.13710 null
2026-03-14 TheraAgent: Multi-Agent Framework with Self-Evolving Memory and Evidence-Calibrated Reasoning for PET Theranostics Zhihao Chen et.al. 2603.13676 null
2026-03-14 Audo-Sight: AI-driven Ambient Perception Across Edge-Cloud for Blind and Low Vision Users Jacob Bradshaw et.al. 2603.13668 null
2026-03-13 SemRep: Generative Code Representation Learning with Code Transformations Weichen Li et.al. 2603.13640 null
2026-03-13 Design and evaluation of an agentic workflow for crisis-related synthetic tweet datasets Roben Delos Reyes et.al. 2603.13625 null
2026-03-13 EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings Shiva Krishna Reddy Malay et.al. 2603.13594 null
2026-03-13 From Experiments to Expertise: Scientific Knowledge Consolidation for AI-Driven Computational Research Haonan Huang et.al. 2603.13191 null
2026-03-13 Semantic Invariance in Agentic AI I. de Zarzà et.al. 2603.13173 null
2026-03-13 Steve-Evolving: Open-World Embodied Self-Evolution via Fine-Grained Diagnosis and Dual-Track Knowledge Distillation Zhengwei Xie et.al. 2603.13131 null
2026-03-13 AgentRM: An OS-Inspired Resource Manager for LLM Agent Systems Jianshu She et.al. 2603.13110 null
2026-03-13 Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence Seunghwan Bang et.al. 2603.13091 null
2026-03-13 PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses Chenlong Yin et.al. 2603.13026 null
2026-03-13 ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning Bangjun Xiao et.al. 2603.13019 null
2026-03-13 Structured Distillation for Personalized Agent Memory: 11x Token Reduction with Retrieval Preservation Sydney Lewis et.al. 2603.13017 null
2026-03-13 Generative Horcrux: Designing AI Carriers for Afterlife Selves Zhen-Chi Lai et.al. 2603.12971 null
2026-03-13 Efficient and Interpretable Multi-Agent LLM Routing via Ant Colony Optimization Xudong Wang et.al. 2603.12933 null
2026-03-13 ToolTree: Efficient LLM Agent Tool Planning via Dual-Feedback Monte Carlo Tree Search and Bidirectional Pruning Shuo Yang et.al. 2603.12740 null
2026-03-13 AI Planning Framework for LLM-Based Web Agents Orit Shahnovsky et.al. 2603.12710 null
2026-03-13 Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw Zonghao Ying et.al. 2603.12644 null
2026-03-13 Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents Yushu Li et.al. 2603.12634 null
2026-03-13 Collaborative Multi-Agent Optimization for Personalized Memory System Wenyu Mao et.al. 2603.12631 null
2026-03-13 AEGIS: No Tool Call Left Unchecked – A Pre-Execution Firewall and Audit Layer for AI Agents Aojie Yuan et.al. 2603.12621 null
2026-03-13 Human-AI Collaborative Autonomous Experimentation With Proxy Modeling for Comparative Observation Arpan Biswas et.al. 2603.12618 null
2026-03-13 ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents Jiangrong Wu et.al. 2603.12614 null
2026-03-13 InterDeepResearch: Enabling Human-Agent Collaborative Information Seeking through Interactive Deep Research Bo Pan et.al. 2603.12608 null
2026-03-13 AgentDrift: Unsafe Recommendation Drift Under Tool Corruption Hidden by Ranking Metrics in LLM Agents Zekun Wu et.al. 2603.12564 null
2026-03-13 Large Language Models as Delivery Rider: Generating Instant Food Delivery Riders’ Routing Decision with LLM Agent Framework Chengbo Zhang et.al. 2603.12559 null
2026-03-12 Generating Expressive and Customizable Evals for Timeseries Data Analysis Agents with AgentFuel Aadyaa Maddi et.al. 2603.12483 null
2026-03-12 OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams Yibin Yan et.al. 2603.12265 null
2026-03-12 Security Considerations for Artificial Intelligence Agents Ninghui Li et.al. 2603.12230 null
2026-03-12 IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse Yushi Bai et.al. 2603.12201 null
2026-03-12 GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows Zexuan Yan et.al. 2603.12155 null
2026-03-12 Automatic Generation of High-Performance RL Environments Seth Karten et.al. 2603.12145 null
2026-03-12 O3N: Omnidirectional Open-Vocabulary Occupancy Prediction Mengfei Duan et.al. 2603.12144 null
2026-03-12 Increasing intelligence in AI agents can worsen collective outcomes Neil F. Johnson et.al. 2603.12129 null
2026-03-12 On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents Deyu Zou et.al. 2603.12109 null
2026-03-12 XSkill: Continual Learning from Experience and Skills in Multimodal Agents Guanyu Jiang et.al. 2603.12056 null
2026-03-12 Kinetic SIS opinion-driven models with asymmetric awareness feedback: macroscopic limit and polarization Juan Pablo Pinasco et.al. 2603.12041 null
2026-03-12 Cascade: Composing Software-Hardware Attack Gadgets for Adversarial Threat Amplification in Compound AI Systems Sarbartha Banerjee et.al. 2603.12023 null
2026-03-12 Can RL Improve Generalization of LLM Agents? An Empirical Study Zhiheng Xi et.al. 2603.12011 null
2026-03-12 LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories Qianpu Sun et.al. 2603.11987 null
2026-03-12 HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios Jiayue Pu et.al. 2603.11975 null
2026-03-12 Normative Common Ground Replication (NormCoRe): Replication-by-Translation for Studying Norms in Multi-agent AI Luca Deck et.al. 2603.11974 null
2026-03-12 PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents Minjia Wang et.al. 2603.11955 null
2026-03-12 CogSearch: A Cognitive-Aligned Multi-Agent Framework for Proactive Decision Support in E-Commerce Search Zhouwei Zhai et.al. 2603.11927 null
2026-03-12 QUARE: Multi-Agent Negotiation for Balancing Quality Attributes in Requirements Engineering Haowei Cheng et.al. 2603.11890 null
2026-03-12 ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics Omar Coser et.al. 2603.11872 null
2026-03-12 Derain-Agent: A Plug-and-Play Agent Framework for Rainy Image Restoration Zhaocheng Yu et.al. 2603.11866 null
2026-03-12 Social, Legal, Ethical, Empathetic and Cultural Norm Operationalisation for AI Agents Radu Calinescu et.al. 2603.11864 null
2026-03-12 You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents Ching-Yu Kao et.al. 2603.11862 null
2026-03-12 OpenClaw PRISM: A Zero-Fork, Defense-in-Depth Runtime Security Layer for Tool-Augmented LLM Agents Frank Li et.al. 2603.11853 null
2026-03-12 Hybrid Human-Agent Social Dilemmas in Energy Markets Isuri Perera et.al. 2603.11834 null
2026-03-12 Large language models for optical network O&M: Agent-embedded workflow for automation Shengnan Li et.al. 2603.11828 null
2026-03-12 DocSage: An Information Structuring Agent for Multi-Doc Multi-Entity Question Answering Teng Lin et.al. 2603.11798 null
2026-03-12 Disentangled Representation Learning through Unsupervised Symmetry Group Discovery Dang-Nhu Barthélémy et.al. 2603.11790 null
2026-03-11 LLMGreenRec: LLM-Based Multi-Agent Recommender System for Sustainable E-Commerce Hao N. Nguyen et.al. 2603.11025 null
2026-03-11 Task-Aware Delegation Cues for LLM Agents Xingrui Gu et.al. 2603.11011 null
2026-03-11 Bio-Inspired Self-Supervised Learning for Wrist-worn IMU Signals Prithviraj Tarale et.al. 2603.10961 null
2026-03-11 UltrasoundAgents: Hierarchical Multi-Agent Evidence-Chain Reasoning for Breast Ultrasound Diagnosis Yali Zhu et.al. 2603.10852 null
2026-03-11 Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis Yujie Zheng et.al. 2603.10846 null
2026-03-11 Nurture-First Agent Development: Building Domain-Expert AI Agents Through Conversational Knowledge Crystallization Linghao Zhang et.al. 2603.10808 null
2026-03-11 Re-Evaluating EVMBench: Are AI Agents Ready for Smart Contract Security? Chaoyuan Peng et.al. 2603.10795 null
2026-03-11 A Control-Theoretic Foundation for Agentic Systems Ali Eslami et.al. 2603.10779 null
2026-03-11 HeartAgent: An Autonomous Agent System for Explainable Differential Diagnosis in Cardiology Shuang Zhou et.al. 2603.10764 null
2026-03-11 AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations Yu He et.al. 2603.10749 null
2026-03-11 Pneuma-Seeker: A Relational Reification Mechanism to Align AI Agents with Human Work over Relational Data Muhammad Imam Luthfi Balaka et.al. 2603.10747 null
2026-03-11 FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model Xiaoxu Xu et.al. 2603.10712 null
2026-03-11 Structured Linked Data as a Memory Layer for Agent-Orchestrated Retrieval Andrea Volpini et.al. 2603.10700 null
2026-03-11 Cybo-Waiter: A Physical Agentic Framework for Humanoid Whole-Body Locomotion-Manipulation Peng Ren et.al. 2603.10675 null
2026-03-11 Breaking User-Centric Agency: A Tri-Party Framework for Agent-Based Recommendation Yaxin Gong et.al. 2603.10673 null
2026-03-11 Terminal Is All You Need: Design Properties for Human-AI Agent Collaboration Alexandre De Masi et.al. 2603.10664 null
2026-03-11 ESG Reporting Lifecycle Management with Large Language Models and AI Agents Thong Hoang et.al. 2603.10646 null
2026-03-11 Trajectory-Informed Memory Generation for Self-Improving Agent Systems Gaodan Fang et.al. 2603.10600 null
2026-03-11 DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtime Julian Lorenz et.al. 2603.10538 null
2026-03-11 Safe and Scalable Web Agent Learning via Recreated Websites Hyungjoo Chae et.al. 2603.10505 null
2026-03-11 World2Act: Latent Action Post-Training via Skill-Compositional World Models An Dinh Vuong et.al. 2603.10422 null
2026-03-11 Don’t Let the Claw Grip Your Hand: A Security Analysis and Defense Framework for OpenClaw Zhengyang Shan et.al. 2603.10387 null
2026-03-11 AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU Yuning Zhang et.al. 2603.10342 null
2026-03-10 PRECEPT: Planning Resilience via Experience, Context Engineering & Probing Trajectories A Unified Framework for Test-Time Adaptation with Compositional Rule Learning and Pareto-Guided Prompt Evolution Arash Shahmansoori et.al. 2603.09641 null
2026-03-10 Context Engineering: From Prompts to Corporate Multi-Agent Architecture Vera V. Vishnyakova et.al. 2603.09619 null
2026-03-10 Vibe-Creation: The Epistemology of Human-AI Emergent Cognition Ilya Levin et.al. 2603.09486 null
2026-03-10 A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation Yoon Jo Kim et.al. 2603.09448 null
2026-03-10 ProvAgent: Threat Detection Based on Identity-Behavior Binding and Multi-Agent Collaborative Attack Investigation Wenhao Yan et.al. 2603.09358 null
2026-03-10 Beyond Scaling: Assessing Strategic Reasoning and Rapid Decision-Making Capability of LLMs in Zero-sum Environments Yang Li et.al. 2603.09337 null
2026-03-10 Reward-Zero: Language Embedding Driven Implicit Reward Mechanisms for Reinforcement Learning Heng Zhang et.al. 2603.09331 null
2026-03-10 TA-Mem: Tool-Augmented Autonomous Memory Retrieval for LLM in Long-Term Conversational QA Mengwei Yuan et.al. 2603.09297 null
2026-03-10 Abundant Intelligence and Deficient Demand: A Macro-Financial Stress Test of Rapid AI Adoption Xupeng Chen et.al. 2603.09209 null
2026-03-10 MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data Zongxia Li et.al. 2603.09206 null
2026-03-10 Latent-DARM: Bridging Discrete Diffusion And Autoregressive Models For Reasoning Lina Berrayana et.al. 2603.09184 null
2026-03-10 Real-Time Trust Verification for Safe Agentic Actions using TrustBench Tavishi Sharma et.al. 2603.09157 null
2026-03-10 DataFactory: Collaborative Multi-Agent Framework for Advanced Table Question Answering Tong Wang et.al. 2603.09152 null
2026-03-10 Deep Tabular Research via Continual Experience-Driven Execution Junnan Dong et.al. 2603.09151 null
2026-03-10 From Days to Minutes: An Autonomous AI Agent Achieves Reliable Clinical Triage in Remote Patient Monitoring Seunghwan Kim et.al. 2603.09052 null
2026-03-10 EPOCH: An Agentic Protocol for Multi-Round System Optimization Zhanlin Liu et.al. 2603.09049 null
2026-03-10 FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Yinpeng Wu et.al. 2603.09046 null
2026-03-10 Time, Identity and Consciousness in Language Model Agents Elija Perrier et.al. 2603.09043 null
2026-03-09 Meissa: Multi-modal Medical Agentic Intelligence Yixiong Chen et.al. 2603.09018 null
2026-03-09 Can AI Agents Generate Microservices? How Far are We? Bassam Adnan et.al. 2603.09004 null
2026-03-09 Agentic Critical Training Weize Liu et.al. 2603.08706 null
2026-03-09 Predicting Conflict Impact on Performance in O-RAN Pietro Brach del Prever et.al. 2603.08685 null
2026-03-09 OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning Krista Opsahl-Ong et.al. 2603.08655 null
2026-03-09 PostTrainBench: Can LLM Agents Automate LLM Post-Training? Ben Rank et.al. 2603.08640 null
2026-03-09 Reachability-based Temporal Logic Verification for Reliable LLM-guided Human-Autonomy Teaming Joonwon Choi et.al. 2603.08633 null
2026-03-09 Trust via Reputation of Conviction Aravind R. Iyengar et.al. 2603.08575 null
2026-03-09 SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement Yi Chen et.al. 2603.08520 null
2026-03-09 Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA Ummar Abbas et.al. 2603.08501 null
2026-03-09 Towards Modeling Cybersecurity Behavior of Humans in Organizations Klaas Ole Kürtz et.al. 2603.08484 null
2026-03-09 Behavioral Generative Agents for Power Dispatch and Auction Shaoze Li et.al. 2603.08477 null
2026-03-09 One Model Is Enough: Native Retrieval Embeddings from LLM Agent Hidden States Bo Jiang et.al. 2603.08429 null
2026-03-09 IronEngine: Towards General AI Assistant Xi Mo et.al. 2603.08425 null
2026-03-09 A Recipe for Stable Offline Multi-agent Reinforcement Learning Dongsu Lee et.al. 2603.08399 null
2026-03-09 A Hierarchical Error-Corrective Graph Framework for Autonomous Agents with LLM-Based Action Generation Cong Cao et.al. 2603.08388 null
2026-03-09 M $^3$ -ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering Peijin Xie et.al. 2603.08369 null
2026-03-09 SPD-RAG: Sub-Agent Per Document Retrieval-Augmented Generation Yagiz Can Akay et.al. 2603.08329 null
2026-03-09 Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design Hai Xia et.al. 2603.08322 null
2026-03-09 FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use Jiaxuan Lu et.al. 2603.08262 null
2026-03-09 SplitAgent: A Privacy-Preserving Distributed Architecture for Enterprise-Cloud Agent Collaboration Jianshu She et.al. 2603.08221 null
2026-03-09 RexDrug: Reliable Multi-Drug Combination Extraction through Reasoning-Enhanced LLMs Zhijun Wang et.al. 2603.08166 null
2026-03-07 The Yerkes-Dodson Curve for AI Agents: Emergent Cooperation Under Environmental Pressure in Multi-Agent LLM Simulations Ivan Pasichnyk et.al. 2603.07360 null
2026-03-07 LLM-FK: Multi-Agent LLM Reasoning for Foreign Key Detection in Large-Scale Complex Databases Zijian Tang et.al. 2603.07278 null
2026-03-07 Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing Arash Marioriyad et.al. 2603.07202 null
2026-03-07 Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice Yuxu Ge et.al. 2603.07191 null
2026-03-07 Agentic Planning with Reasoning for Image Styling via Offline RL Subhojyoti Mukherjee et.al. 2603.07148 null
2026-03-07 aCAPTCHA: Verifying That an Entity Is a Capable Agent via Asymmetric Hardness Zuyao Xu et.al. 2603.07116 null
2026-03-07 Enhancing Consistency of Werewolf AI through Dialogue Summarization and Persona Information Yoshiki Tanaka et.al. 2603.07111 null
2026-03-07 AutoUE: Automated Generation of 3D Games in Unreal Engine via Multi-Agent Systems Lei Yin et.al. 2603.07106 null
2026-03-07 Exploring the Reasoning Depth of Small Language Models in Software Architecture: A Multidimensional Evaluation Framework Towards Software Engineering 2.0 Ha Vo et.al. 2603.07091 null
2026-03-07 Enhancing Web Agents with a Hierarchical Memory Tree Yunteng Tan et.al. 2603.07024 null
2026-03-07 SuperSkillsStack: Agency, Domain Knowledge, Imagination, and Taste in Human-AI Design Education Qian Huang et.al. 2603.07016 null
2026-03-06 LLM2SMT: Building an SMT Solver with Zero Human-Written Code Mikoláš Janota et.al. 2603.06931 null
2026-03-06 A Contrastive Fewshot RGBD Traversability Segmentation Framework for Indoor Robotic Navigation Qiyuan An et.al. 2603.06927 null
2026-03-06 T2Nav Algebraic Topology Aware Temporal Graph Memory and Loop Detection for ZeroShot Visual Navigation Quang-Anh N. D. et.al. 2603.06918 null
2026-03-06 Empowering Locally Deployable Medical Agent via State Enhanced Logical Skills for FHIR-based Clinical Tasks Wanrong Yang et.al. 2603.06902 null
2026-03-06 Distributed Legal Infrastructure for a Trustworthy Agentic Web Tomer Jordi Chaffer et.al. 2603.06884 null
2026-03-06 LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models Matthew Lyle Olson et.al. 2603.06874 null
2026-03-06 AIMD-L: An automated laboratory for high-throughput characterization of structural materials for extreme environments Todd C. Hufnagel et.al. 2603.06835 null
2026-03-06 Stability-Guided Exploration for Diverse Motion Generation Eckart Cobo-Briesewitz et.al. 2603.06773 null
2026-03-06 Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing Anmol Gulati et.al. 2603.06503 null
2026-03-06 REACT++: Efficient Cross-Attention for Real-Time Scene Graph Generation Maëlic Neau et.al. 2603.06386 null
2026-03-06 Provuse: Platform-Side Function Fusion for Performance and Efficiency in FaaS Environments Niklas Kowallik et.al. 2603.06170 null
2026-03-06 Spatial Colour Mixing Illusions as a Perception Stress Test for Vision-Language Models Nicoleta-Nina Basoc et.al. 2603.06141 null
2026-03-06 Agentic LLM Planning via Step-Wise PDDL Simulation: An Empirical Characterisation Kai Göbel et.al. 2603.06064 null
2026-03-06 A LINDDUN-based Privacy Threat Modeling Framework for GenAI Qianying Liao et.al. 2603.06051 null
2026-03-06 Pre-AI Baseline: Developer IDE Satisfaction and Tool Autonomy in 2022 Nikola Balić et.al. 2603.06050 null
2026-03-06 THETA: A Textual Hybrid Embedding-based Topic Analysis Framework and AI Scientist Agent for Scalable Computational Social Science Zhenke Duan et.al. 2603.05972 null
2026-03-06 XAI for Coding Agent Failures: Transforming Raw Execution Traces into Actionable Insights Arun Joshi et.al. 2603.05941 null
2026-03-06 DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality Yukun Huang et.al. 2603.05912 null
2026-03-06 Evolving Deception: When Agents Evolve, Deception Wins Zonghao Ying et.al. 2603.05872 null
2026-03-06 Evolving Medical Imaging Agents via Experience-driven Self-skill Discovery Lin Fan et.al. 2603.05860 null
2026-03-06 Proof-of-Guardrail in AI Agents and What (Not) to Trust from It Xisen Jin et.al. 2603.05786 null
2026-03-05 TML-Bench: Benchmark for Data Science Agents on Tabular ML Tasks Mykola Pinchuk et.al. 2603.05764 null
2026-03-05 Agentic AI – Physicist Collaboration in Experimental Particle Physics: A Proof-of-Concept Measurement with LEP Open Data Anthony Badea et.al. 2603.05735 null
2026-03-05 Real-Time AI Service Economy: A Framework for Agentic Computing Across the Continuum Lauri Lovén et.al. 2603.05614 null
2026-03-05 Building Enterprise Realtime Voice Agents from Scratch: A Technical Tutorial Jielin Qiu et.al. 2603.05413 null
2026-03-05 Building AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learned Nghi D. Q. Bui et.al. 2603.05344 null
2026-03-05 WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces Sicheng Fan et.al. 2603.05295 null
2026-03-05 STRUCTUREDAGENT: Planning with AND/OR Trees for Long-Horizon Web Tasks ELita Lobo et.al. 2603.05294 null
2026-03-05 KARL: Knowledge Agents via Reinforcement Learning Jonathan D. Chang et.al. 2603.05218 null
2026-03-05 Escaping the Hydrolysis Trap: An Agentic Workflow for Inverse Design of Durable Photocatalytic Covalent Organic Frameworks Iman Peivaste et.al. 2603.05188 null
2026-03-05 MedCoRAG: Interpretable Hepatology Diagnosis via Hybrid Evidence Retrieval and Multispecialty Consensus Zheng Li et.al. 2603.05129 null
2026-03-05 Bidirectional Curriculum Generation: A Multi-Agent Framework for Data-Efficient Mathematical Reasoning Boren Hu et.al. 2603.05120 null
2026-03-05 Jagarin: A Three-Layer Architecture for Hibernating Personal Duty Agents on Mobile Ravi Kiran Kadaboina et.al. 2603.05069 null
2026-03-05 WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents Sicheng Fan et.al. 2603.05044 null
2026-03-05 AegisUI: Behavioral Anomaly Detection for Structured User Interface Protocols in AI Agent Systems Mohd Safwan Uddin et.al. 2603.05031 null
2026-03-05 RepoLaunch: Automating Build&Test Pipeline of Code Repositories on ANY Language and ANY Platform Kenan Li et.al. 2603.05026 null
2026-03-05 BioLLMAgent: A Hybrid Framework with Enhanced Structural Interpretability for Simulating Human Decision-Making in Computational Psychiatry Zuo Fei et.al. 2603.05016 null
2026-03-05 Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video Understanding Zheng Wang et.al. 2603.04977 null
2026-03-05 TimeWarp: Evaluating Web Agents by Revisiting the Past Md Farhan Ishmam et.al. 2603.04949 null
2026-03-05 EVMbench: Evaluating AI Agents on Smart Contract Security Justin Wang et.al. 2603.04915 null
2026-03-05 AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows Ivoline C. Ngong et.al. 2603.04902 null
2026-03-05 EvoTool: Self-Evolving Tool-Use Policy Optimization in LLM Agents via Blame-Aware Mutation and Diversity-Aware Selection Shuo Yang et.al. 2603.04900 null
2026-03-05 FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications Yunfan Zhang et.al. 2603.04857 null
2026-03-05 EchoGuard: An Agentic Framework with Knowledge-Graph Memory for Detecting Manipulative Communication in Longitudinal Dialogue Ratna Kandala et.al. 2603.04815 null
2026-03-04 AgentIR: Reasoning-Aware Retrival for Deep Research Agents Zijian Chen et.al. 2603.04384 null
2026-03-04 $τ$ -Knowledge: Evaluating Conversational Agents over Unstructured Knowledge Quan Shi et.al. 2603.04370 null
2026-03-04 Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks Haoyu Liu et.al. 2603.04364 null
2026-03-04 LabelBuddy: An Open Source Music and Audio Language Annotation Tagging Tool Using AI Assistance Ioannis Prokopiou et.al. 2603.04293 null
2026-03-04 Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory Zhenting Wang et.al. 2603.04257 null
2026-03-04 CodeTaste: Can LLMs Generate Human-Level Code Refactorings? Alex Thillen et.al. 2603.04177 null
2026-03-04 A Multi-Agent Framework for Interpreting Multivariate Physiological Time Series Davide Gabrielli et.al. 2603.04142 null
2026-03-04 Right in Time: Reactive Reasoning in Regulated Traffic Spaces Simon Kohaut et.al. 2603.03977 null
2026-03-04 From Threat Intelligence to Firewall Rules: Semantic Relations in Hybrid AI Agent and Expert System Architectures Chiara Bonfanti et.al. 2603.03911 null
2026-03-04 A Rubric-Supervised Critic from Sparse Real-World Outcomes Xingyao Wang et.al. 2603.03800 null
2026-03-04 LifeBench: A Benchmark for Long-Horizon Multi-Source Memory Zihao Cheng et.al. 2603.03781 null
2026-03-04 MACC: Multi-Agent Collaborative Competition for Scientific Exploration Satoshi Oyama et.al. 2603.03780 null
2026-03-04 Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding Junhan Chen et.al. 2603.03762 null
2026-03-04 AgentSelect: Benchmark for Narrative Query-to-Agent Recommendation Yunxiao Shi et.al. 2603.03761 null
2026-03-04 Agentic Peer-to-Peer Networks: From Content Distribution to Capability and Action Sharing Taotao Wang et.al. 2603.03753 null
2026-03-04 AI4S-SDS: A Neuro-Symbolic Solvent Design System via Sparse MCTS and Differentiable Physics Alignment Jiangyu Chen et.al. 2603.03686 null
2026-03-04 MAGE: Meta-Reinforcement Learning for Language Agents toward Strategic Exploration and Exploitation Lu Yang et.al. 2603.03680 null
2026-03-04 Mozi: Governed Autonomy for Drug Discovery LLM Agents He Cao et.al. 2603.03655 null
2026-03-04 Goal-Driven Risk Assessment for LLM-Powered Systems: A Healthcare Case Study Neha Nagaraja et.al. 2603.03633 null
2026-03-04 Behind the Prompt: The Agent-User Problem in Information Retrieval Saber Zerhoudi et.al. 2603.03630 null
2026-03-03 Conversational Learning Diagnosis via Reasoning Multi-Turn Interactive Learning Fangzhou Yao et.al. 2603.03236 null
2026-03-03 AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework Zihang Zeng et.al. 2603.03233 null
2026-03-03 Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use Aradhye Agarwal et.al. 2603.03205 null
2026-03-03 Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration? Dadi Guo et.al. 2603.03202 null
2026-03-03 BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? Guoxin Chen et.al. 2603.03194 null
2026-03-03 Saarthi for AGI: Towards Domain-Specific General Intelligence for Formal Verification Aman Kumar et.al. 2603.03175 null
2026-03-03 From Language to Action: Can LLM-Based Agents Be Used for Embodied Robot Cognition? Shinas Shaji et.al. 2603.03148 null
2026-03-03 How to Model AI Agents as Personas?: Applying the Persona Ecosystem Playground to 41,300 Posts on Moltbook for Behavioral Insights Danial Amin et.al. 2603.03140 null
2026-03-03 Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation Hongliu Cao et.al. 2603.03116 null
2026-03-03 RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy Optimization Siwei Zhang et.al. 2603.03078 null
2026-03-03 MA-CoNav: A Master-Slave Multi-Agent Framework with Hierarchical Collaboration and Dual-Level Reflection for Long-Horizon Embodied VLN Ling Luo et.al. 2603.03024 null
2026-03-03 Contextualized Privacy Defense for LLM Agents Yule Wen et.al. 2603.02983 null
2026-03-03 Architecting Trust in Artificial Epistemic Agents Nahema Marchal et.al. 2603.02960 null
2026-03-03 Changing Pedagogical Paradigms: Integrating Generative AI in Mathematics to Enhance Digital Literacy through ‘Mathematical Battles with AI’ Maria Moskalenko et.al. 2603.02955 null
2026-03-03 Learning to Generate and Extract: A Multi-Agent Collaboration Framework For Zero-shot Document-level Event Arguments Extraction Guangjun Zhang et.al. 2603.02909 null
2026-03-03 Speech recognition assisted by large language models to command software orally – Application to an augmented and virtual reality web app for immersive molecular graphics Fabio Cortes Rodriguez et.al. 2603.02901 null
2026-03-03 SpecLoop: An Agentic RTL-to-Specification Framework with Formal Verification Feedback Loop Fu-Chieh Chang et.al. 2603.02895 null
2026-03-03 LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval Minh-Chi Phung et.al. 2603.02888 null
2026-03-03 BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation Zihao Zhu et.al. 2603.02816 null
2026-03-03 VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning Ruiyang Zhang et.al. 2603.02795 null
2026-03-02 GenDB: The Next Generation of Query Processing – Synthesized, Not Engineered Jiale Lao et.al. 2603.02081 null
2026-03-02 MMNavAgent: Multi-Magnification WSI Navigation Agent for Clinically Consistent Whole-Slide Analysis Zhengyang Xu et.al. 2603.02079 null
2026-03-02 When an AI Judges Your Work: The Hidden Costs of Algorithmic Assessment David Almog et.al. 2603.02076 null
2026-03-02 Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning Guilhem Fouilhé et.al. 2603.02070 null
2026-03-02 “When to Hand Off, When to Work Together”: Expanding Human-Agent Co-Creative Collaboration through Concurrent Interaction Kihoon Son et.al. 2603.02050 null
2026-03-02 Expanding LLM Agent Boundaries with Strategy-Guided Exploration Andrew Szot et.al. 2603.02045 null
2026-03-02 CHOP: Counterfactual Human Preference Labels Improve Obstacle Avoidance in Visuomotor Navigation Policies Gershom Seneviratne et.al. 2603.02004 null
2026-03-02 LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations Viet-Thanh Pham et.al. 2603.01952 null
2026-03-02 CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification Jinpeng Chen et.al. 2603.01940 null
2026-03-02 A Hetero-functional Graph State Estimator for Watershed Systems: Application to the Chesapeake Bay Megan S. Harris et.al. 2603.01931 null
2026-03-02 Demonstrating ViviDoc: Generating Interactive Documents through Human-Agent Collaboration Yinghao Tang et.al. 2603.01912 null
2026-03-02 Agentic Code Reasoning Shubham Ugare et.al. 2603.01896 null
2026-03-02 Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition Mingwei Liu et.al. 2603.01814 null
2026-03-02 What Papers Don’t Tell You: Recovering Tacit Knowledge for Automated Paper Reproduction Lehui Li et.al. 2603.01801 null
2026-03-02 TopoCurate:Modeling Interaction Topology for Tool-Use Agent Training Jinluan Yang et.al. 2603.01714 null
2026-03-02 CeProAgents: A Hierarchical Agents System for Automated Chemical Process Development Yuhang Yang et.al. 2603.01654 null
2026-03-02 LexChronos: An Agentic Framework for Structured Event Timeline Extraction in Indian Jurisprudence Anka Chandrahas Tummepalli et.al. 2603.01651 null
2026-03-02 QCAgent: An agentic framework for quality-controllable pathology report generation from whole slide image Rundong Wang et.al. 2603.01647 null
2026-03-02 SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing Anjali Parashar et.al. 2603.01630 null
2026-03-02 Evaluating and Understanding Scheming Propensity in LLM Agents Mia Hopman et.al. 2603.01608 null
2026-02-27 CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation Weinan Dai et.al. 2602.24286 null
2026-02-27 UXSim: Towards a Hybrid User Search Simulation Saber Zerhoudi et.al. 2602.24241 null
2026-02-27 Controllable Reasoning Models Are Private Thinkers Haritz Puerto et.al. 2602.24210 null
2026-02-27 Agentic AI-RAN: Enabling Intent-Driven, Explainable and Self-Evolving Open RAN Intelligence Zhizhou He et.al. 2602.24115 null
2026-02-27 A Novel Hierarchical Multi-Agent System for Payments Using LLMs Joon Kiat Chua et.al. 2602.24068 null
2026-02-27 Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking Zhicheng Fang et.al. 2602.24009 null
2026-02-27 Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments Florent Delgrange et.al. 2602.23997 null
2026-02-27 Novice Developers Produce Larger Review Overhead for Project Maintainers while Vibe Coding Syed Ammar Asdaque et.al. 2602.23905 null
2026-02-27 Experience-Guided Self-Adaptive Cascaded Agents for Breast Cancer Screening and Diagnosis with Reduced Biopsy Referrals Pramit Saha et.al. 2602.23899 null
2026-02-27 AoE: Always-on Egocentric Human Video Collection for Embodied AI Bowen Yang et.al. 2602.23893 null
2026-02-27 RUMAD: Reinforcement-Unifying Multi-Agent Debate Chao Wang et.al. 2602.23864 null
2026-02-27 CLFEC: A New Task for Unified Linguistic and Factual Error Correction in paragraph-level Chinese Professional Writing Jian Kai et.al. 2602.23845 null
2026-02-27 OPTIAGENT: A Physics-Driven Agentic Framework for Automated Optical Design Yuyu Geng et.al. 2602.23761 null
2026-02-27 U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation Xiang Deng et.al. 2602.23739 null
2026-02-27 From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning Seungdong Yoa et.al. 2602.23729 null
2026-02-27 The Auton Agentic AI Framework Sheng Cao et.al. 2602.23720 null
2026-02-27 ProductResearch: Training E-Commerce Deep Research Agents via Multi-Agent Synthetic Trajectory Distillation Jiangyuan Wang et.al. 2602.23716 null
2026-02-27 PseudoAct: Leveraging Pseudocode Synthesis for Flexible Planning and Action Control in Large Language Model Agents Yihan et.al. 2602.23668 null
2026-02-27 SGAgent: Suggestion-Guided LLM-Based Multi-Agent Framework for Repository-Level Software Repair Quanjun Zhang et.al. 2602.23647 null
2026-02-27 Toward E2E Intelligence in 6G Networks: An AI Agent-Based RAN-CN Converged Intelligence Framework Youbin Han et.al. 2602.23623 null
2026-02-26 Toward Expert Investment Teams:A Multi-Agent LLM System with Fine-Grained Trading Tasks Kunihiro Miyazaki et.al. 2602.23330 null
2026-02-26 ParamMem: Augmenting Language Agents with Parametric Reflective Memory Tianjun Yao et.al. 2602.23320 null
2026-02-26 EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents Wenjia Wang et.al. 2602.23205 null
2026-02-26 ESAA: Event Sourcing for Autonomous Agents in LLM-Based Software Engineering Elzo Brito dos Santos Filho et.al. 2602.23193 null
2026-02-26 AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios Zhaochen Su et.al. 2602.23166 null
2026-02-26 Three AI-agents walk into a bar . . . . `Lord of the Flies’ tribalism emerges among smart AI-Agents Dhwanil M. Mori et.al. 2602.23093 null
2026-02-26 Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy Matthew Sutton et.al. 2602.23088 null
2026-02-26 Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent Boyang Zhang et.al. 2602.23079 null
2026-02-26 Accelerated Online Risk-Averse Policy Evaluation in POMDPs with Theoretical Guarantees and Novel CVaR Bounds Yaacov Pariente et.al. 2602.23073 null
2026-02-26 Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization Zeyuan Liu et.al. 2602.23008 null
2026-02-26 FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning Zehao Li et.al. 2602.22963 null
2026-02-26 Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study Zihao Zhao et.al. 2602.22959 null
2026-02-26 OmniGAIA: Towards Native Omni-Modal AI Agents Xiaoxi Li et.al. 2602.22897 null
2026-02-26 Decentralized Ranking Aggregation: Gossip Algorithms for Borda and Copeland Consensus Anna Van Elst et.al. 2602.22847 null
2026-02-26 DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation Hao Zheng et.al. 2602.22839 null
2026-02-26 Endogenous Poverty Traps in Continuous Time: A Signaling Approach Massimo Giannini et.al. 2602.22836 null
2026-02-26 Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks Shuo He et.al. 2602.22817 null
2026-02-26 MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasks Shiqian Su et.al. 2602.22808 null
2026-02-26 AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications Yujie Zhao et.al. 2602.22769 null
2026-02-26 Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents Jiahong Xiang et.al. 2602.22764 null
2026-02-25 An Empirical Study of Bugs in Modern LLM Agent Frameworks Xinxue Zhu et.al. 2602.21806 null
2026-02-25 Two-Stage Active Distribution Network Voltage Control via LLM-RL Collaboration: A Hybrid Knowledge-Data-Driven Approach Xu Yang et.al. 2602.21715 null
2026-02-25 Hierarchical LLM-Based Multi-Agent Framework with Prompt Optimization for Multi-Robot Task Planning Tomoya Kawabe et.al. 2602.21670 null
2026-02-25 Towards Autonomous Graph Data Analytics with Analytics-Augmented Generation Qiange Wang et.al. 2602.21604 null
2026-02-25 Power and Limitations of Aggregation in Compound AI Systems Nivasini Ananthakrishnan et.al. 2602.21556 null
2026-02-25 Which Tool Response Should I Trust? Tool-Expertise-Aware Chest X-ray Agent with Multimodal Agentic Learning Zheang Huai et.al. 2602.21517 null
2026-02-25 Both Ends Count! Just How Good are LLM Agents at “Text-to-Big SQL”? Germán T. Eizaguirre et.al. 2602.21480 null
2026-02-25 Pancake: Hierarchical Memory System for Multi-Agent LLM Serving Zhengding Hu et.al. 2602.21477 null
2026-02-25 SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards Dengjia Zhang et.al. 2602.21158 null
2026-02-24 MemoPhishAgent: Memory-Augmented Multi-Modal LLM Agent for Phishing URL Detection Xuan Chen et.al. 2602.21394 null
2026-02-24 Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibration Charafeddine Mouzouni et.al. 2602.21368 null
2026-02-24 A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives Dmitrii Pantiukhin et.al. 2602.21351 null
2026-02-24 Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data Emre Can Acikgoz et.al. 2602.21320 null
2026-02-24 Region of Interest Segmentation and Morphological Analysis for Membranes in Cryo-Electron Tomography Xingyi Cheng et.al. 2602.21195 null
2026-02-24 A Benchmark for Deep Information Synthesis Debjit Paul et.al. 2602.21143 null
2026-02-24 “Are You Sure?”: An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems Xinfeng Li et.al. 2602.21127 null
2026-02-24 Cooperative-Competitive Team Play of Real-World Craft Robots Rui Zhao et.al. 2602.21119 null
2026-02-24 From Perception to Action: An Interactive Benchmark for Vision Reasoning Yuhao Wu et.al. 2602.21015 null
2026-02-24 Toward an Agentic Infused Software Ecosystem Mark Marron et.al. 2602.20979 null
2026-02-24 Airavat: An Agentic Framework for Internet Measurement Alagappan Ramanathan et.al. 2602.20924 null
2026-02-24 SoK: Agentic Skills – Beyond Tool Use in LLM Agents Yanna Jiang et.al. 2602.20867 null
2026-02-24 Body-Reservoir Governance in Repeated Games: Embodied Decision-Making, Dynamic Sentinel Adaptation, and Complexity-Regularized Optimization Yuki Nakamura et.al. 2602.20846 null
2026-02-24 Pipeline for Verifying LLM-Generated Mathematical Solutions Varvara Sazonova et.al. 2602.20770 null
2026-02-24 PyVision-RL: Forging Open Agentic Vision Models via RL Shitian Zhao et.al. 2602.20739 null
2026-02-24 AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs Che Wang et.al. 2602.20720 null
2026-02-24 ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction Che Wang et.al. 2602.20708 null
2026-02-24 How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective Bo Peng et.al. 2602.20687 null
2026-02-24 Agile V: A Compliance-Ready Framework for AI-Augmented Engineering – From Concept to Audit-Ready Delivery Christopher Koch et.al. 2602.20684 null
2026-02-24 Grid-Mind: An LLM-Orchestrated Multi-Fidelity Agent for Automated Connection Impact Assessment Mohamed Shamseldein et.al. 2602.20683 null
2026-02-24 AnimeAgent: Is the Multi-Agent via Image-to-Video models a Good Disney Storytelling Artist? Hailong Yan et.al. 2602.20664 null
2026-02-24 Grounding LLMs in Scientific Discovery via Embodied Actions Bo Zhang et.al. 2602.20639 null
2026-02-24 Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing Yuejiao Su et.al. 2602.20597 null
2026-02-23 Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks David Schmotz et.al. 2602.20156 null
2026-02-23 The LLMbda Calculus: AI Agents, Conversations, and Information Flow Zac Garby et.al. 2602.20064 null
2026-02-23 Interaction Theater: A case of LLM Agents Interacting at Scale Sarath Shekkizhar et.al. 2602.20059 null
2026-02-23 Let There Be Claws: An Early Social Network Analysis of AI Agents on Moltbook H. C. W. Price et.al. 2602.20044 null
2026-02-23 AgenticSum: An Agentic Inference-Time Framework for Faithful Clinical Text Summarization Fahmida Liza Piya et.al. 2602.20040 null
2026-02-23 Agents of Chaos Natalie Shapira et.al. 2602.20021 null
2026-02-23 Assessing Risks of Large Language Models in Mental Health Support: A Framework for Automated Clinical AI Red Teaming Ian Steenstra et.al. 2602.19948 null
2026-02-23 OpenClaw, Moltbook, and ClawdLab: From Agent-Only Social Networks to Autonomous Scientific Research Lukas Weidener et.al. 2602.19810 null
2026-02-23 Janus-Faced Technological Progress and the Arms Race in the Education of Humans and Chatbots Wolfgang Kuhle et.al. 2602.19783 null
2026-02-23 TAPE: Tool-Guided Adaptive Planning and Constrained Execution in Language Model Agents Jongwon Jeong et.al. 2602.19633 null
2026-02-23 ISO-Bench: Can Coding Agents Optimize Real-World Inference Workloads? Ayush Nangia et.al. 2602.19594 null
2026-02-23 Vinedresser3D: Agentic Text-guided 3D Editing Yankuan Chi et.al. 2602.19542 null
2026-02-23 Cost-Aware Diffusion Active Search Arundhati Banerjee et.al. 2602.19538 null
2026-02-23 An LLM-Enabled Frequency-Aware Flow Diffusion Model for Natural-Language-Guided Power System Scenario Generation Zhenghao Zhou et.al. 2602.19522 null
2026-02-23 Ada-RS: Adaptive Rejection Sampling for Selective Thinking Yirou Ge et.al. 2602.19519 null
2026-02-23 Pixel2Phys: Distilling Governing Laws from Visual Dynamics Ruikun Li et.al. 2602.19516 null
2026-02-23 Security Risks of AI Agents Hiring Humans: An Empirical Marketplace Study Pulak Mehta et.al. 2602.19514 null
2026-02-23 Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark Lalitha Pranathi Pulavarthy et.al. 2602.19502 null
2026-02-23 ComplLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making Ziyang Guo et.al. 2602.19458 null
2026-02-23 When AI Teammates Meet Code Review: Collaboration Signals Shaping the Integration of Agent-Authored Pull Requests Costain Nachuma et.al. 2602.19441 null
2026-02-20 SARAH: Spatially Aware Real-time Agentic Humans Evonne Ng et.al. 2602.18432 null
2026-02-20 Towards More Standardized AI Evaluation: From Models to Agents Ali El Filali et.al. 2602.18029 null
2026-02-20 Mean-Field Reinforcement Learning without Synchrony Shan Yang et.al. 2602.18026 null
2026-02-20 NIMMGen: Learning Neural-Integrated Mechanistic Digital Twins with LLMs Zihan Guan et.al. 2602.18008 null
2026-02-20 WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics Madhav Kanda et.al. 2602.17990 null
2026-02-20 Mining Type Constructs Using Patterns in AI-Generated Code Imgyeong Lee et.al. 2602.17955 null
2026-02-20 Alignment in Time: Peak-Aware Orchestration for Long-Horizon Agentic Systems Hanjing Shi et.al. 2602.17910 null
2026-02-19 El Agente Gráfico: Structured Execution Graphs for Scientific Agents Jiaru Bai et.al. 2602.17902 null
2026-02-19 El Agente Sólido: A New Age(nt) for Solid State Simulations Sai Govind Hari Kumar et.al. 2602.17886 null
2026-02-19 Exploring The Impact Of Proactive Generative AI Agent Roles In Time-Sensitive Collaborative Problem-Solving Tasks Anirban Mukhopadhyay et.al. 2602.17864 null
2026-02-19 The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems Leon Staufer et.al. 2602.17753 null
2026-02-19 FAMOSE: A ReAct Approach to Automated Feature Discovery Keith Burghardt et.al. 2602.17641 null
2026-02-19 What Makes a Good LLM Agent for Real-world Penetration Testing? Gelei Deng et.al. 2602.17622 null
2026-02-19 Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs Luke Huang et.al. 2602.17616 null
2026-02-19 AutoNumerics: An Autonomous, PDE-Agnostic Multi-Agent Pipeline for Scientific Computing Jianda Du et.al. 2602.17607 null
2026-02-19 Modeling Distinct Human Interaction in Web Agents Faria Huq et.al. 2602.17588 null
2026-02-19 RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward Qiucheng Wu et.al. 2602.17558 null
2026-02-19 KLong: Training LLM Agent for Extremely Long-horizon Tasks Yue Liu et.al. 2602.17547 null
2026-02-19 MedClarify: An information-seeking AI agent for medical diagnosis with case-specific follow-up questions Hui Min Wong et.al. 2602.17308 null
2026-02-19 Web Verbs: Typed Abstractions for Reliable Task Composition on the Agentic Web Linxi Jiang et.al. 2602.17245 null
2026-02-19 From Labor to Collaboration: A Methodological Experiment Using AI Agents to Augment Research Perspectives in Taiwan’s Humanities and Social Sciences Yi-Chih Huang et.al. 2602.17221 null
2026-02-19 AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation Siyu Wang et.al. 2602.17100 null
2026-02-19 What to Cut? Predicting Unnecessary Methods in Agentic Code Generation Kan Watanabe et.al. 2602.17091 null
2026-02-19 How AI Coding Agents Communicate: A Study of Pull Request Description Characteristics and Human Review Responses Kan Watanabe et.al. 2602.17084 null
2026-02-19 Phase-Aware Mixture of Experts for Agentic Reinforcement Learning Shengtian Yang et.al. 2602.17038 null
2026-02-19 Wink: Recovering from Misbehaviors in Coding Agents Rahul Nanda et.al. 2602.17037 null
2026-02-19 M2F: Automated Formalization of Mathematical Literature at Scale Zichen Wang et.al. 2602.17016 null
2026-02-19 Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History Serin Kim et.al. 2602.17003 null
2026-02-18 Automating Agent Hijacking via Structural Template Injection Xinhao Deng et.al. 2602.16958 null
2026-02-18 LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation Hejia Zhang et.al. 2602.16953 null
2026-02-18 Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents Arnold Cartagena et.al. 2602.16943 null
2026-02-18 Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents Wenxuan Ding et.al. 2602.16699 null
2026-02-18 Towards a Science of AI Agent Reliability Stephan Rabanser et.al. 2602.16666 null
2026-02-18 Evaluating Collective Behaviour of Hundreds of LLM Agents Richard Willis et.al. 2602.16662 null
2026-02-18 DataJoint 2.0: A Computational Substrate for Agentic Scientific Workflows Dimitri Yatsenko et.al. 2602.16585 null
2026-02-18 MerLean: An Agentic Framework for Autoformalization in Quantum Computation Yuanjie Ren et.al. 2602.16554 null
2026-02-18 Agentic AI, Medical Morality, and the Transformation of the Patient-Physician Relationship Robert Ranisch et.al. 2602.16553 null
2026-02-18 Automated Extraction of Mechanical Constitutive Models from Scientific Literature using Large Language Models: Applications in Cultural Heritage Conservation Rui Hu et.al. 2602.16551 null
2026-02-18 Estimation of Conformal Metrics Jérôme Taupin et.al. 2602.16466 null
2026-02-18 RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation Yixue Zhang et.al. 2602.16444 null
2026-02-18 Verifiable Semantics for Agent-to-Agent Communication Philipp Schoenegger et.al. 2602.16424 null
2026-02-18 Label-Consistent Data Generation for Aspect-Based Sentiment Analysis Using LLM Agents Mohammad H. A. Monfared et.al. 2602.16379 null
2026-02-18 Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents Nivya Talokar et.al. 2602.16346 null
2026-02-18 Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents Yun-Shiuan Chuang et.al. 2602.16246 null
2026-02-18 Submodular Maximization under Supermodular Constraint: Greedy Guarantees Ajitesh Srivastava et.al. 2602.16240 null
2026-02-18 EnterpriseGym Corecraft: Training Generalizable Agents on High-Fidelity RL Environments Sushant Mehta et.al. 2602.16179 null
2026-02-18 Learning Personalized Agents from Human Feedback Kaiqu Liang et.al. 2602.16173 null
2026-02-18 HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents Jiangweizhi Peng et.al. 2602.16165 null
2026-02-18 Empirical Cumulative Distribution Function Clustering for LLM-based Agent System Analysis Chihiro Watanabe et.al. 2602.16131 null
2026-02-18 GPSBench: Do Large Language Models Understand GPS Coordinates? Thinh Hung Truong et.al. 2602.16105 null
2026-02-17 The Limits of Long-Context Reasoning in Automated Bug Fixing Ravi Raju et.al. 2602.16069 null
2026-02-17 Developing AI Agents with Simulated Data: Why, what, and how? Xiaoran Liu et.al. 2602.15816 null
2026-02-17 Decision Quality Evaluation Framework at Pinterest Yuqi Tian et.al. 2602.15809 null
2026-02-17 GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent Systems Yiqin Yang et.al. 2602.15776 null
2026-02-17 GLM-5: from Vibe Coding to Agentic Engineering GLM-5 Team et.al. 2602.15763 null
2026-02-17 Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections Xianglin Yang et.al. 2602.15654 null
2026-02-17 Improving MLLMs in Embodied Exploration and Question Answering with Human-Inspired Memory Modeling Ji Li et.al. 2602.15513 null
2026-02-17 In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations Mohammad Aflah Khan et.al. 2602.15456 null
2026-02-17 World-Model-Augmented Web Agents with Action Correction Zhouzhou Shen et.al. 2602.15384 null
2026-02-17 EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use Siwei Wen et.al. 2602.15329 null
2026-02-17 AgriWorld:A World Tools Protocol Framework for Verifiable Agricultural Reasoning with Code-Executing LLM Agents Zhixing Zhang et.al. 2602.15325 null
2026-02-17 Visual Persuasion: What Influences Decisions of Vision-Language Models? Manuel Cherep et.al. 2602.15278 null
2026-02-17 Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence Alisa Vinogradova et.al. 2602.15019 null
2026-02-16 Knowing Isn’t Understanding: Re-grounding Generative Proactivity with Epistemic and Behavioral Insight Kirandeep Kaur et.al. 2602.15259 null
2026-02-16 Secure and Energy-Efficient Wireless Agentic AI Networks Yuanyan Song et.al. 2602.15212 null
2026-02-16 Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems Mason Nakamura et.al. 2602.15198 null
2026-02-16 OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction Skyler Hallinan et.al. 2602.15197 null
2026-02-16 Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs Luise Ge et.al. 2602.15173 null
2026-02-16 ResearchGym: Evaluating Language Model Agents on Real-World AI Research Aniketh Garikaparthi et.al. 2602.15112 null
2026-02-16 PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement Yian Wang et.al. 2602.14968 null
2026-02-16 Sovereign Agents: Towards Infrastructural Sovereignty and Diffused Accountability in Decentralized AI Botao Amber Hu et.al. 2602.14951 null
2026-02-16 MAC-AMP: A Closed-Loop Multi-Agent Collaboration System for Multi-Objective Antimicrobial Peptide Design Gen Zhou et.al. 2602.14926 null
2026-02-16 Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions Mohammed Mehedi Hasan et.al. 2602.14878 null
2026-02-16 EmbeWebAgent: Embedding Web Agents into Any Customized UI Chenyang Ma et.al. 2602.14865 null
2026-02-16 Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows Bardia Mohammadi et.al. 2602.14849 null
2026-02-16 On Convergence Analysis of Network-GIANT: An approximate Hessian-based fully distributed optimization algorithm Souvik Das et.al. 2602.14830 null
2026-02-16 Overthinking Loops in Agents: A Structural Risk via MCP Tools Yohan Lee et.al. 2602.14798 null
2026-02-16 WebWorld: A Large-Scale World Model for Web Agent Training Zikai Xiao et.al. 2602.14721 null
2026-02-16 Removing Planner Bias in Goal Recognition Through Multi-Plan Dataset Generation Mustafa F. Abdelwahed et.al. 2602.14691 null
2026-02-16 ST-EVO: Towards Generative Spatio-Temporal Evolution of Multi-Agent Communication Topologies Xingjian Wu et.al. 2602.14681 null
2026-02-16 FactorMiner: A Self-Evolving Agent with Skills and Experience Memory for Financial Alpha Discovery Yanlong Wang et.al. 2602.14670 null
2026-02-16 Towards Selection as Power: Bounding Decision Authority in Autonomous Agents Jose Manuel de la Chica Rodriguez et.al. 2602.14606 null
2026-02-16 MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs Gabriel Roccabruna et.al. 2602.14589 null
2026-02-16 Efficient Multi-round LLM Inference over Disaggregated Serving Wenhao He et.al. 2602.14516 null
2026-02-16 When OpenClaw AI Agents Teach Each Other: Peer Learning Patterns in the Moltbook Community Eason Chen et.al. 2602.14477 null
2026-02-16 Socially-Weighted Alignment: A Game-Theoretic Framework for Multi-Agent LLM Systems Furkan Mumcu et.al. 2602.14471 null
2026-02-16 Traceable Latent Variable Discovery Based on Multi-Agent Collaboration Huaming Du et.al. 2602.14456 null
2026-02-16 RoboSolver: A Multi-Agent Large Language Model Framework for Solving Robotic Arm Problems Hamid Khabazi et.al. 2602.14438 null
2026-02-13 Asynchronous Verified Semantic Caching for Tiered LLM Architectures Asmit Kumar Singh et.al. 2602.13165 null
2026-02-13 In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach Yiran Gao et.al. 2602.13156 null
2026-02-13 Automating UI Optimization through Multi-Agentic Reasoning Zhipeng Li et.al. 2602.13126 null
2026-02-13 TraceBack: Multi-Agent Decomposition for Fine-Grained Table Attribution Tejas Anvekar et.al. 2602.13059 null
2026-02-13 Quantization-Aware Collaborative Inference for Large Embodied AI Models Zhonghao Lyu et.al. 2602.13052 null
2026-02-13 SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Yujiong Shen et.al. 2602.12984 null
2026-02-13 Human Tool: An MCP-Style Framework for Human-Agent Collaboration Yuanrong Tang et.al. 2602.12953 null
2026-02-13 BrowseComp- $V^3$ : A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents Huanyao Zhang et.al. 2602.12876 null
2026-02-13 WebClipper: Efficient Evolution of Web Agents with Graph-based Trajectory Pruning Junjie Wang et.al. 2602.12852 null
2026-02-13 SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Xiangyi Li et.al. 2602.12670 null
2026-02-13 Think Fast and Slow: Step-Level Cognitive Depth Adaptation for LLM Agents Ruihan Yang et.al. 2602.12662 null
2026-02-13 The Rise of AI Agent Communities: Large-Scale Analysis of Discourse and Interaction on Moltbook Lingyao Li et.al. 2602.12634 null
2026-02-13 AI Agents for Inventory Control: Human-LLM-OR Complementarity Jackie Baek et.al. 2602.12631 null
2026-02-13 Opinion dynamics and mutual influence with LLM agents through dialog simulation Yulong He et.al. 2602.12583 null
2026-02-13 Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation Lajanugen Logeswaran et.al. 2602.12544 null
2026-02-13 Favia: Forensic Agent for Vulnerability-fix Identification and Analysis André Storhaug et.al. 2602.12500 null
2026-02-12 Existence Results and KKT Optimality Conditions for Generalized Quasiconvex Functions M. H. Alizadeh et.al. 2602.12455 null
2026-02-12 Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward Renjun Xu et.al. 2602.12430 null
2026-02-12 Provably Convergent Actor-Critic in Risk-averse MARL Yizhou Zhang et.al. 2602.12386 null
2026-02-12 Agentic Test-Time Scaling for WebAgents Nicholas Lee et.al. 2602.12276 null
2026-02-12 CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use Zhen Zhang et.al. 2602.12268 null
2026-02-12 Think like a Scientist: Physics-guided LLM Agent for Equation Discovery Jianke Yang et.al. 2602.12259 null
2026-02-12 VIRENA: Virtual Arena for Research, Education, and Democratic Innovation Emma Hoes et.al. 2602.12207 null
2026-02-12 Visual Reasoning Benchmark: Evaluating Multimodal LLMs on Classroom-Authentic Visual Problems from Primary Education Mohamed Huti et.al. 2602.12196 null
2026-02-12 MalTool: Malicious Tool Attacks on LLM Agents Yuepeng Hu et.al. 2602.12194 null
2026-02-12 On the Adoption of AI Coding Agents in Open-source Android and iOS Development Muhammad Ahmad Khan et.al. 2602.12144 null
2026-02-12 STAR : Bridging Statistical and Agentic Reasoning for Large Model Performance Prediction Xiaoxiao Wang et.al. 2602.12143 null
2026-02-12 Embodied AI Agents for Team Collaboration in Co-located Blue-Collar Work Kaisa Vaananen et.al. 2602.12136 null
2026-02-12 Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation Kehang Zhu et.al. 2602.12089 null
2026-02-12 Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? Thibaud Gloaguen et.al. 2602.11988 null
2026-02-12 Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments Romain Froger et.al. 2602.11964 null
2026-02-12 AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection Pretam Ray et.al. 2602.11931 null
2026-02-12 Agentic AI for Cybersecurity: A Meta-Cognitive Architecture for Governable Autonomy Andrei Kojukhov et.al. 2602.11897 null
2026-02-12 Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems Wanxing Wu et.al. 2602.11877 null
2026-02-12 Intelligent AI Delegation Nenad Tomašev et.al. 2602.11865 null
2026-02-12 Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception Lai Wei et.al. 2602.11858 null
2026-02-12 FlowMind: Execute-Summarize for Structured Workflow Generation from LLM Reasoning Yihao Liu et.al. 2602.11782 null
2026-02-12 TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents Aladin Djuhera et.al. 2602.11767 null
2026-02-12 Cooperation Breakdown in LLM Agents Under Communication Delays Keita Nishimoto et.al. 2602.11754 null
2026-02-11 Learning to Compose for Cross-domain Agentic Workflow Generation Jialiang Wang et.al. 2602.11114 null
2026-02-11 GameDevBench: Evaluating Agentic Capabilities Through Game Development Wayne Chi et.al. 2602.11103 null
2026-02-11 Fine-Tuning GPT-5 for GPU Kernel Generation Ali Tehrani et.al. 2602.11000 null
2026-02-11 TVCACHE: A Stateful Tool-Value Cache for Post-Training LLM Agents Abhishek Vijaya Kumar et.al. 2602.10986 null
2026-02-11 Blind Gods and Broken Screens: Architecting a Secure, Intent-Centric Mobile Agent Operating System Zhenhua Zou et.al. 2602.10915 null
2026-02-11 See, Plan, Snap: Evaluating Multimodal GUI Agents in Scratch Xingyi Zhang et.al. 2602.10814 null
2026-02-11 DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories Chenlong Deng et.al. 2602.10809 null
2026-02-11 Hidden Licensing Risks in the LLMware Ecosystem Bo Wang et.al. 2602.10758 null
2026-02-11 Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents Yifei Li et.al. 2602.10715 null
2026-02-11 UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory Yongshi Ye et.al. 2602.10652 null
2026-02-11 ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents YoungHoon Jeon et.al. 2602.10620 null
2026-02-11 Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters Ailin Huang et.al. 2602.10604 null
2026-02-11 When Skills Lie: Hidden-Comment Injection in LLM Agents Qianli Wang et.al. 2602.10498 null
2026-02-11 From Prompt-Response to Goal-Directed Systems: The Evolution of Agentic AI Software Architecture Mamdouh Alenezi et.al. 2602.10479 null
2026-02-11 The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis Peiran Wang et.al. 2602.10453 null
2026-02-11 Distributed Online Convex Optimization with Nonseparable Costs and Constraints Zhaoye Pan et.al. 2602.10452 null
2026-02-11 AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning Liyang Chen et.al. 2602.10439 null
2026-02-11 AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles Wenkai Fan et.al. 2602.10429 null
2026-02-10 Frame-Level Internal Tool Use for Temporal Grounding in Audio LMs Joesph An et.al. 2602.10230 null
2026-02-10 Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents Haochen Wang et.al. 2602.10226 null
2026-02-10 SAGE: Scalable Agentic 3D Scene Generation for Embodied AI Hongchi Xia et.al. 2602.10116 null
2026-02-10 DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos Juncheng Mu et.al. 2602.10105 null
2026-02-10 Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning Zhaoyang Wang et.al. 2602.10090 null
2026-02-10 Anagent For Enhancing Scientific Table & Figure Analysis Xuehang Guo et.al. 2602.10081 null
2026-02-10 Chain of Mindset: Reasoning with Adaptive Cognitive Modes Tianyi Jiang et.al. 2602.10063 null
2026-02-10 Artisan: Agentic Artifact Evaluation Doehyun Baek et.al. 2602.10046 null
2026-02-10 Discovering High Level Patterns from Simulation Traces Sean Memery et.al. 2602.10009 null
2026-02-10 Human-AI Synergy Supports Collective Creative Search Chenyi Li et.al. 2602.10001 null
2026-02-10 Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis? Taeyoon Kim et.al. 2602.09937 null
2026-02-10 Focus Session: LLM4PQC – An Agentic Framework for Accurate and Efficient Synthesis of PQC Cores Buddhi Perera et.al. 2602.09919 null
2026-02-10 Immersion in the GitHub Universe: Scaling Coding Agents to Mastery Jiale Zhao et.al. 2602.09892 null
2026-02-10 The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies Chenxu Wang et.al. 2602.09877 null
2026-02-10 BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation Yucheng Hu et.al. 2602.09849 null
2026-02-10 Generative AI Adoption in an Energy Company: Exploring Challenges and Use Cases Malik Abdul Sami et.al. 2602.09846 null
2026-02-10 Internalizing Multi-Agent Reasoning for Accurate and Efficient LLM-based Recommendation Yang Wu et.al. 2602.09829 null
2026-02-10 AnalyticsGPT: An LLM Workflow for Scientometric Question Answering Khang Ly et.al. 2602.09817 null
2026-02-10 Tiny Moves: Game-based Hypothesis Refinement Agnieszka Dobrowolska et.al. 2602.09801 null
2026-02-10 QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery George Tsigkourakos et.al. 2602.09774 null
2026-02-10 TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution Deyang Jiang et.al. 2602.09662 null
2026-02-10 MATA: Multi-Agent Framework for Reliable and Flexible Table Question Answering Sieun Hyeon et.al. 2602.09642 null
2026-02-09 InternAgent-1.5: A Unified Agentic Framework for Long-Horizon Autonomous Scientific Discovery Shiyang Feng et.al. 2602.08990 null
2026-02-09 A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents Raghu Arghal et.al. 2602.08964 null
2026-02-09 Digital Twin and Agentic AI for Wild Fire Disaster Management: Intelligent Virtual Situation Room Mohammad Morsali et.al. 2602.08949 null
2026-02-09 Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance Giovanni Pinna et.al. 2602.08915 null
2026-02-09 Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems Lang Feng et.al. 2602.08847 null
2026-02-09 AMEM4Rec: Leveraging Cross-User Similarity for Memory Evolution in Agentic LLM Recommenders Minh-Duc Nguyen et.al. 2602.08837 null
2026-02-09 LLM-Enhanced Wearables for Comprehensible Health Guidance in LMICs Mohammad Shaharyar Ahsan et.al. 2602.08701 null
2026-02-09 6G-Bench: An Open Benchmark for Semantic Communication and Network-Level Reasoning with Foundation Models in AI-Native 6G Networks Mohamed Amine Ferrag et.al. 2602.08675 null
2026-02-09 PRISM: A Principled Framework for Multi-Agent Reasoning via Gain Decomposition Yiming Yang et.al. 2602.08586 null
2026-02-09 Agent-Supported Foresight for AI Systemic Risks: AI Agents for Breadth, Experts for Judgment Leon Fröhling et.al. 2602.08565 null
2026-02-09 Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs Ahmed Salem et.al. 2602.08563 null
2026-02-09 Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approaches Syed Mehtab Hussain Shah et.al. 2602.08561 null
2026-02-09 Three Lessons from Citizen-Centric Participatory AI Design Eike Schneiders et.al. 2602.08554 null
2026-02-09 EvoCorps: An Evolutionary Multi-Agent Framework for Depolarizing Online Discourse Ning Lin et.al. 2602.08529 null
2026-02-09 SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios Tian Gao et.al. 2602.08440 null
2026-02-09 Decentralized Intent-Based Multi-Robot Task Planner with LLM Oracles on Hyperledger Fabric Farhad Keramat et.al. 2602.08421 null
2026-02-09 From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent Yuhang Wang et.al. 2602.08412 null
2026-02-09 On Protecting Agentic Systems’ Intellectual Property via Watermarking Liwen Wang et.al. 2602.08401 null
2026-02-09 Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AI Feiyu Wu et.al. 2602.08373 null
2026-02-09 Toward Formalizing LLM-Based Agent Designs through Structural Context Modeling and Semantic Dynamics Analysis Haoyu Jia et.al. 2602.08276 null
2026-02-06 Agentic Uncertainty Reveals Agentic Overconfidence Jean Kaddour et.al. 2602.06948 null
2026-02-06 Directing Space: Rehearsing Architecture as Performer with Explainable AI Pavlos Panagiotidis et.al. 2602.06915 null
2026-02-06 TraceCoder: A Trace-Driven Multi-Agent Framework for Automated Debugging of LLM-Generated Code Jiangping Huang et.al. 2602.06875 null
2026-02-06 AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents Alisia Lupidi et.al. 2602.06855 null
2026-02-06 LLM Active Alignment: A Nash Equilibrium Perspective Tonghan Wang et.al. 2602.06836 null
2026-02-06 ScaleEnv: Scaling Environment Synthesis from Scratch for Generalist Interactive Tool-Use Agent Training Dunwei Tu et.al. 2602.06820 null
2026-02-06 Table-as-Search: Formulate Long-Horizon Agentic Information Seeking as Table Completion Tian Lan et.al. 2602.06724 null
2026-02-06 The Law of Task-Achieving Body Motion: Axiomatizing Success of Robot Manipulation Actions Malte Huerkamp et.al. 2602.06572 null
2026-02-06 SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees Tianyi Hu et.al. 2602.06554 null
2026-02-06 Completing Missing Annotation: Multi-Agent Debate for Accurate and Scalable Relevant Assessment for IR Benchmarks Minjeong Ban et.al. 2602.06526 null
2026-02-06 Evolutionary Generation of Multi-Agent Systems Yuntong Hu et.al. 2602.06511 null
2026-02-06 TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents Yibing Liu et.al. 2602.06443 null
2026-02-06 Towards Adaptive Environment Generation for Training Embodied Agents Teresa Yeo et.al. 2602.06366 null
2026-02-06 Nipping the Drift in the Bud: Retrospective Rectification for Robust Vision-Language Navigation Gang He et.al. 2602.06356 null
2026-02-06 Zero-Trust Runtime Verification for Agentic Payment Protocols: Mitigating Replay and Context-Binding Failures in AP2 Qianlong Lan et.al. 2602.06345 null
2026-02-06 Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent Zhou Xuan et.al. 2602.06325 null
2026-02-06 Trustworthy AI Software Engineers Aldeida Aleti et.al. 2602.06310 null
2026-02-05 RuleSmith: Multi-Agent LLMs for Automated Game Balancing Ziyao Zeng et.al. 2602.06232 null
2026-02-05 M3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual Reasoning Bangji Yang et.al. 2602.06166 null
2026-02-05 DyTopo: Dynamic Topology Routing for Multi-Agent Reasoning via Semantic Matching Yuxing Lu et.al. 2602.06039 null
2026-02-05 V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval Dongyang Chen et.al. 2602.06034 null
2026-02-05 PhysicsAgentABM: Physics-Guided Generative Agent-Based Modeling Kavana Venkatesh et.al. 2602.06030 null
2026-02-05 Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory Haozhen Zhang et.al. 2602.06025 null
2026-02-05 From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents Bingsheng Yao et.al. 2602.05987 null
2026-02-05 SAGE: Benchmarking and Improving Retrieval for Deep Research Agents Tiansheng Hu et.al. 2602.05975 null
2026-02-05 Learning to Share: Selective Memory for Efficient Parallel Agentic Systems Joseph Fioresi et.al. 2602.05965 null
2026-02-05 AgenticTagger: Structured Item Representation for Recommendation with LLM Agents Zhouhang Xie et.al. 2602.05945 null
2026-02-05 ContextBench: A Benchmark for Context Retrieval in Coding Agents Han Li et.al. 2602.05892 null
2026-02-05 Agent2Agent Threats in Safety-Critical LLM Assistants: A Human-Centric Taxonomy Lukas Stappen et.al. 2602.05877 null
2026-02-05 DuoDrama: Supporting Screenplay Refinement Through LLM-Assisted Human Reflection Yuying Tang et.al. 2602.05854 null
2026-02-05 OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions Fangzhi Xu et.al. 2602.05843 null
2026-02-05 Bifrost: Steering Strategic Trajectories to Bridge Contextual Gaps for Self-Improving Agents Quan M. Tran et.al. 2602.05810 null
2026-02-05 RocqSmith: Can Automatic Optimization Forge Better Proof Agents? Andrei Kozyrev et.al. 2602.05762 null
2026-02-05 Task-Oriented Robot-Human Handovers on Legged Manipulators Andreea Tulbure et.al. 2602.05760 null
2026-02-05 Learning to Inject: Automated Prompt Injection via Reinforcement Learning Xin Chen et.al. 2602.05746 null
2026-02-05 A Dual-Loop Agent Framework for Automated Vulnerability Reproduction Bin Liu et.al. 2602.05721 null
2026-02-05 Making AI Agents Evaluate Misleading Charts without Nudging Swaroop Panda et.al. 2602.05662 null
2026-02-05 Reactive Knowledge Representation and Asynchronous Reasoning Simon Kohaut et.al. 2602.05625 null
2026-02-05 On the computational properties of ambivalent sets and functions Dag Normann et.al. 2602.05620 null
2026-02-04 Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing Zhaotian Weng et.al. 2602.04837 null
2026-02-04 Active Asymmetric Multi-Agent Multimodal Learning under Uncertainty Rui Liu et.al. 2602.04763 null
2026-02-04 SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation David F. Ramirez et.al. 2602.04712 null
2026-02-04 WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learning Zelai Xu et.al. 2602.04634 null
2026-02-04 VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration Jaeyoon Jung et.al. 2602.04587 null
2026-02-04 Vibe AIGC: A New Paradigm for Content Generation via Agentic Orchestration Jiaheng Liu et.al. 2602.04575 null
2026-02-04 OSCAgent: Accelerating the Discovery of Organic Solar Cells with LLM Agents Zhaolin Hu et.al. 2602.04510 null
2026-02-04 ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control Zhentao Tang et.al. 2602.04496 null
2026-02-04 EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL Lunjun Zhang et.al. 2602.04417 null
2026-02-04 From Assumptions to Actions: Turning LLM Reasoning into Uncertainty-Aware Planning for Embodied Agents SeungWon Seo et.al. 2602.04326 null
2026-02-04 Agent-Omit: Training Efficient LLM Agents for Adaptive Thought and Observation Omission via Agentic Reinforcement Learning Yansong Ning et.al. 2602.04284 null
2026-02-04 Data Agents: Levels, State of the Art, and Open Problems Yuyu Luo et.al. 2602.04261 null
2026-02-04 Why Agentic-PRs Get Rejected: A Comparative Study of Coding Agents Sota Nakashima et.al. 2602.04226 null
2026-02-04 InterPReT: Interactive Policy Restructuring and Training Enable Effective Imitation Learning from Laypersons Feiyu Gavin Zhu et.al. 2602.04213 null
2026-02-04 From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents Xinyue Wang et.al. 2602.04197 null
2026-02-04 A Modern System Recipe for Situated Embodied Human-Robot Conversation with Real-Time Multimodal LLMs and Tool-Calling Dong Won Lee et.al. 2602.04157 null
2026-02-04 OMG-Agent: Toward Robust Missing Modality Generation with Decoupled Coarse-to-Fine Agentic Workflows Ruiting Dai et.al. 2602.04144 null
2026-02-04 DELTA: Deliberative Multi-Agent Reasoning with Reinforcement Learning for Multimodal Psychological Counseling Jiangnan Yang et.al. 2602.04112 null
2026-02-03 Active Epistemic Control for Query-Efficient Verified Planning Shuhui Qu et.al. 2602.03974 null
2026-02-03 AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent Yinyi Luo et.al. 2602.03955 null
2026-02-03 AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations Minjun Zhu et.al. 2602.03828 null
2026-02-03 FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation Zimu Lu et.al. 2602.03798 null
2026-02-03 WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents Xilong Wang et.al. 2602.03792 null
2026-02-03 Efficient Estimation of Kernel Surrogate Models for Task Attribution Zhenshuo Zhang et.al. 2602.03783 null
2026-02-03 An Empirical Study of Collective Behaviors and Social Dynamics in Large Language Model Agents Farnoosh Hashemi et.al. 2602.03775 null
2026-02-03 Training Multi-Turn Search Agent via Contrastive Dynamic Branch Sampling Yubao Zhao et.al. 2602.03719 null
2026-02-03 OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering Yifan Zhu et.al. 2602.03707 null
2026-02-03 Cognitively Diverse Multiple-Choice Question Generation: A Hybrid Multi-Agent Framework with Large Language Models Yu Tian et.al. 2602.03704 null
2026-02-03 BIRDTurk: Adaptation of the BIRD Text-to-SQL Dataset to Turkish Burak Aktaş et.al. 2602.03633 null
2026-02-03 Can LLMs Do Rocket Science? Exploring the Limits of Complex Reasoning with GTOC 12 Iñaki del Campo et.al. 2602.03630 null
2026-02-03 Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants Valerie Chen et.al. 2602.03593 null
2026-02-03 Don’t believe everything you read: Understanding and Measuring MCP Behavior under Misleading Tool Descriptions Zhihao Li et.al. 2602.03580 null
2026-02-03 Game-Theoretic and Algorithmic Analyses of Multi-Agent Routing under Crossing Costs Tesshu Hanaka et.al. 2602.03455 null
2026-02-03 CRL-VLA: Continual Vision-Language-Action Learning Qixin Zeng et.al. 2602.03445 null
2026-02-03 A-RAG: Scaling Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces Mingxuan Du et.al. 2602.03442 null
2026-02-03 Ontology-to-tools compilation for executable semantic constraint enforcement in LLM agents Xiaochi Zhou et.al. 2602.03439 null
2026-02-03 Failure is Feedback: History-Aware Backtracking for Agentic Traversal in Multimodal Graphs Joohyung Yun et.al. 2602.03432 null
2026-02-03 Verified Critical Step Optimization for LLM Agents Mukai Li et.al. 2602.03412 null
2026-02-03 MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning Shengyuan Liu et.al. 2602.03320 null
2026-02-03 MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research Yifan Shi et.al. 2602.03318 null
2026-02-02 RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents Jialiang Zhu et.al. 2602.02486 null
2026-02-02 AgentRx: Diagnosing AI Agent Failures from Execution Trajectories Shraddha Barke et.al. 2602.02475 null
2026-02-02 MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents Haozhen Zhang et.al. 2602.02474 null
2026-02-02 Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts Aiden Yiliu Li et.al. 2602.02468 null
2026-02-02 Relationship-Aware Hierarchical 3D Scene Graph for Task Reasoning Albert Gassol Puigjaner et.al. 2602.02456 null
2026-02-02 Drift-Bench: Diagnosing Cooperative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction Han Bao et.al. 2602.02455 null
2026-02-02 Multi-Agent Monte Carlo Tree Search for Makespan-Efficient Object Rearrangement in Cluttered Spaces Hanwen Ren et.al. 2602.02411 null
2026-02-02 David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning Samuel Nellessen et.al. 2602.02395 null
2026-02-02 Live-Evo: Online Evolution of Agentic Memory from Continuous Feedback Yaolun Zhang et.al. 2602.02369 null
2026-02-02 SWE-Universe: Scale Real-World Verifiable Environments to Millions Mouxiang Chen et.al. 2602.02361 null
2026-02-02 A Task-Level Evaluation of AI Agents in Open-Source Projects Shojibur Rahman et.al. 2602.02345 null
2026-02-02 Statistical Learning Theory in Lean 4: Empirical Processes from Scratch Yuanhe Zhang et.al. 2602.02285 null
2026-02-02 OmniCode: A Benchmark for Evaluating Software Engineering Agents Atharv Sonwane et.al. 2602.02262 null
2026-02-02 Online Fine-Tuning of Pretrained Controllers for Autonomous Driving via Real-Time Recurrent RL Julian Lemmel et.al. 2602.02236 null
2026-02-02 Fat-Cat: Document-Driven Metacognitive Multi-Agent System for Complex Reasoning Tong Yang et.al. 2602.02206 null
2026-02-02 TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improvement in LLM Agents Hang Yan et.al. 2602.02196 null
2026-02-02 Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents Pengfei He et.al. 2602.02164 null
2026-02-02 D-CORE: Incentivizing Task Decomposition in Large Reasoning Models for Complex Tool Use Bowen Xu et.al. 2602.02160 null
2026-02-02 SIDiffAgent: Self-Improving Diffusion Agent Shivank Garg et.al. 2602.02051 null
2026-02-02 Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents Zeping Li et.al. 2602.02050 null
2026-01-30 PaperBanana: Automating Academic Illustration for AI Scientists Dawei Zhu et.al. 2601.23265 null
2026-01-30 Eigenweights for arithmetic Hirzebruch Proportionality Tony Feng et.al. 2601.23245 null
2026-01-30 Multi-Agent Systems Should be Treated as Principal-Agent Problems Paulius Rauba et.al. 2601.23211 null
2026-01-30 Greedy Routing Reachability Games Pascal Lenzner et.al. 2601.23126 null
2026-01-30 From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching Zhixiang Zhang et.al. 2601.23088 null
2026-01-30 Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning Siyu Gong et.al. 2601.23032 null
2026-01-30 SolAgent: A Specialized Multi-Agent Framework for Solidity Code Generation Wei Chen et.al. 2601.23009 null
2026-01-30 MiTa: A Hierarchical Multi-Agent Collaboration Framework with Memory-integrated and Task Allocation XiaoJie Zhang et.al. 2601.22974 null
2026-01-30 Sifting the Noise: A Comparative Study of LLM Agents in Vulnerability False Positive Filtering Yunpeng Xiong et.al. 2601.22952 null
2026-01-30 MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering Chuanzhe Guo et.al. 2601.22859 null
2026-01-30 FACET: Multi-Agent AI Supporting Teachers in Scaling Differentiated Learning for Diverse Students Jana Gonnermann-Müller et.al. 2601.22788 null
2026-01-30 AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement Libin Qiu et.al. 2601.22758 null
2026-01-30 Test-Time Mixture of World Models for Embodied Agents in Dynamic Environments Jinwoo Jang et.al. 2601.22647 null
2026-01-30 ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review Palash Goyal et.al. 2601.22638 null
2026-01-30 MCP-Diag: A Deterministic, Protocol-Driven Architecture for AI-Native Network Diagnostics Devansh Lodha et.al. 2601.22633 null
2026-01-30 SYMPHONY: Synergistic Multi-agent Planning with Heterogeneous Language Model Assembly Wei Zhu et.al. 2601.22623 null
2026-01-30 From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents Jiaxuan Gao et.al. 2601.22607 null
2026-01-30 TimeMachine-bench: A Benchmark for Evaluating Model Capabilities in Repository-Level Migration Tasks Ryo Fujii et.al. 2601.22597 null
2026-01-30 PerfGuard: A Performance-Aware Agent for Visual Content Generation Zhipeng Chen et.al. 2601.22571 null
2026-01-30 LEAP – Live Experiments for Active Pedagogy Sumedh Karajagi et.al. 2601.22534 null
2026-01-29 Exploring Reasoning Reward Model for Agents Kaixuan Fan et.al. 2601.22154 null
2026-01-29 DynaWeb: Model-Based Reinforcement Learning of Web Agents Hang Ding et.al. 2601.22149 null
2026-01-29 StepShield: When, Not Whether to Intervene on Rogue Agents Gloria Felicia et.al. 2601.22136 null
2026-01-29 World of Workflows: a Benchmark for Bringing World Models to Enterprise Systems Lakshya Gupta et.al. 2601.22130 null
2026-01-29 SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents Yifeng Ding et.al. 2601.22129 null
2026-01-29 Optimizing Agentic Workflows using Meta-tools Sami Abuzakuk et.al. 2601.22037 null
2026-01-29 CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty Johannes Kirmayr et.al. 2601.22027 null
2026-01-29 When “Better” Prompts Hurt: Evaluation-Driven Iteration for LLM Applications Daniel Commey et.al. 2601.22025 null
2026-01-29 Heterogeneous Computing: The Key to Powering the Future of AI Agent Inference Yiren Zhao et.al. 2601.22001 null
2026-01-29 Liquid Interfaces: A Dynamic Ontology for the Interoperability of Autonomous Systems Dhiogo de Sá et.al. 2601.21993 null
2026-01-29 How do Visual Attributes Influence Web Agents? A Comprehensive Evaluation of User Interface Design Factors Kuai Yu et.al. 2601.21961 null
2026-01-29 ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models Bowen Fang et.al. 2601.21947 null
2026-01-29 AgenticSimLaw: A Juvenile Courtroom Multi-Agent Debate Simulation for Explainable High-Stakes Tabular Decision Making Jon Chun et.al. 2601.21936 null
2026-01-29 Self-Compression of Chain-of-Thought via Multi-Agent Reinforcement Learning Yiqun Chen et.al. 2601.21919 null
2026-01-29 JADE: Bridging the Strategic-Operational Gap in Dynamic Agentic RAG Yiqun Chen et.al. 2601.21916 null
2026-01-29 WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agents Yao Zhang et.al. 2601.21872 null
2026-01-29 Embodied Task Planning via Graph-Informed Action Generation with Large Lanaguage Model Xiang Li et.al. 2601.21841 null
2026-01-29 CORE:Toward Ubiquitous 6G Intelligence Through Collaborative Orchestration of Large Language Model Agents Over Hierarchical Edge Zitong Yu et.al. 2601.21822 null
2026-01-29 BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics Dionizije Fa et.al. 2601.21800 null
2026-01-29 E-mem: Multi-agent based Episodic Context Reconstruction for LLM Agent Memory Kaixiang Wang et.al. 2601.21714 null
2026-01-28 MemCtrl: Using MLLMs as Active Memory Controllers on Embodied Agents Vishnu Sashank Dorbala et.al. 2601.20831 null
2026-01-28 Reinforcement Learning via Self-Distillation Jonas Hübotter et.al. 2601.20802 null
2026-01-28 SERA: Soft-Verified Efficient Repository Agents Ethan Shen et.al. 2601.20789 null
2026-01-28 The Monotone Priority System: Foundations of Contract-Specific Sequencing Naveen Durvasula et.al. 2601.20783 null
2026-01-28 Agentic Fog: A Policy-driven Framework for Distributed Intelligence in Fog Computing Saeed Akbar et.al. 2601.20764 null
2026-01-28 AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts Shicheng Fang et.al. 2601.20730 null
2026-01-28 MedViz: An Agent-based, Visual-guided Research Assistant for Navigating Biomedical Literature Huan He et.al. 2601.20709 null
2026-01-28 Investigating the Development of Task-Oriented Communication in Vision-Language Models Boaz Carmeli et.al. 2601.20641 null
2026-01-28 Agent Benchmarks Fail Public Sector Requirements Jonathan Rystrøm et.al. 2601.20617 null
2026-01-28 AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios Kaiyuan Chen et.al. 2601.20613 null
2026-01-28 PathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMs Oguzhan Gungordu et.al. 2601.20539 null
2026-01-28 Normative Equivalence in human-AI Cooperation: Behaviour, Not Identity, Drives Cooperation in Mixed-Agent Groups Nico Mutzner et.al. 2601.20487 null
2026-01-28 Piloting Planetarium Visualizations with LLMs during Live Events in Science Centers Mathis Brossier et.al. 2601.20466 null
2026-01-28 PEARL: Plan Exploration and Adaptive Reinforcement Learning for Multihop Tool Use Qihao Wang et.al. 2601.20439 null
2026-01-28 Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents Qihao Wang et.al. 2601.20412 null
2026-01-28 On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents Jai Lal Lulla et.al. 2601.20404 null
2026-01-28 OmegaUse: Building a General-Purpose GUI Agent for Autonomous Task Execution Le Zhang et.al. 2601.20380 null
2026-01-28 LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning Wei Huang et.al. 2601.20375 null
2026-01-28 AMA: Adaptive Memory via Multi-Agent Collaboration Weiquan Huang et.al. 2601.20352 null
2026-01-28 Demonstration-Free Robotic Control via LLM Agents Brian Y. Tsui et.al. 2601.20334 null
2026-01-27 Towards a complete characterization of indicator variograms and madograms Xavier Emery et.al. 2601.19800 null
2026-01-27 LVLMs and Humans Ground Differently in Referential Communication Peter Zeng et.al. 2601.19792 null
2026-01-27 Agentic Design Patterns: A System-Theoretic Framework Minh-Dung Dao et.al. 2601.19752 null
2026-01-27 Veri-Sure: A Contract-Aware Multi-Agent Framework with Temporal Tracing and Formal Verification for Correct RTL Code Generation Jiale Liu et.al. 2601.19747 null
2026-01-27 Who Said CVE? How Vulnerability Identifiers Are Mentioned by Humans, Bots, and Agents in Pull Requests Pien Rooijendijk et.al. 2601.19636 null
2026-01-27 ComAgent: Multi-LLM based Agentic AI Empowered Intelligent Wireless Networks Haoyun Li et.al. 2601.19607 null
2026-01-27 Toward Architecture-Aware Evaluation Metrics for LLM Agents Débora Souza et.al. 2601.19583 null
2026-01-27 Yunque DeepResearch Technical Report Yuxuan Cai et.al. 2601.19578 null
2026-01-27 ALRM: Agentic LLM for Robotic Manipulation Vitor Gaboardi dos Santos et.al. 2601.19510 null
2026-01-27 Understanding Dominant Themes in Reviewing Agentic AI-authored Code Md. Asif Haider et.al. 2601.19287 null
2026-01-27 LLM-based Vulnerability Detection at Project Scale: An Empirical Study Fengjie Li et.al. 2601.19239 null
2026-01-27 MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution Libo Sun et.al. 2601.19199 null
2026-01-27 Multi-Agent Procedural Graph Extraction with Structural and Logical Refinement Wangyang Ying et.al. 2601.19170 null
2026-01-27 AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection Wachiraphan Charoenwet et.al. 2601.19138 null
2026-01-27 Exploring Weaknesses in Function Call Models via Reinforcement Learning: An Adversarial Data Augmentation Approach Weiran Guo et.al. 2601.19122 null
2026-01-27 LLMs as Orchestrators: Constraint-Compliant Multi-Agent Optimization for Recommendation Systems Guilin Zhang et.al. 2601.19121 null
2026-01-27 Agree to Disagree: Consensus-Free Flocking under Constraints Peter Travis Jardine et.al. 2601.19119 null
2026-01-27 Reward Engineering for Reinforcement Learning in Software Tasks Md Rayhanul Masud et.al. 2601.19100 null
2026-01-27 More at Stake: How Payoff and Language Shape LLM Agent Strategies in Cooperation Dilemmas Trung-Kiet Huynh et.al. 2601.19082 null
2026-01-27 Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward Dipendra Misra et.al. 2601.19055 null
2026-01-26 Are Conversational AI Agents the Way Out? Co-Designing Reader-Oriented News Experiences with Immigrants and Journalists Yongle Zhang et.al. 2601.18772 null
2026-01-26 Let’s Make Every Pull Request Meaningful: An Empirical Analysis of Developer and Agentic Pull Requests Haruhiko Yoshioka et.al. 2601.18749 null
2026-01-26 Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge Li Kang et.al. 2601.18733 null
2026-01-26 TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent Xingyu Sui et.al. 2601.18700 null
2026-01-26 FadeMem: Biologically-Inspired Forgetting for Efficient Agent Memory Lei Wei et.al. 2601.18642 null
2026-01-26 AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning Mingyang Song et.al. 2601.18631 null
2026-01-26 An LLM-Agent-Based Framework for Age of Information Optimization in Heterogeneous Random Access Networks Fang Liu et.al. 2601.18563 null
2026-01-26 GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning Kaixun Jiang et.al. 2601.18543 null
2026-01-26 Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates Yibo Li et.al. 2601.18510 null
2026-01-26 DEEPMED: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & Inference Zihan wang et.al. 2601.18496 null
2026-01-26 DV-VLN: Dual Verification for Reliable LLM-Based Vision-and-Language Navigation Zijun Li et.al. 2601.18492 null
2026-01-26 AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security Dongrui Liu et.al. 2601.18491 null
2026-01-26 daVinci-Dev: Agent-native Mid-training for Software Engineering Ji Zeng et.al. 2601.18418 null
2026-01-26 ARMOR: Agentic Reasoning for Methods Orchestration and Reparameterization for Robust Adversarial Attacks Gabriel Lee Jun Rong et.al. 2601.18386 null
2026-01-26 AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito Yinghan Hou et.al. 2601.18381 null
2026-01-26 Promises, Perils, and (Timely) Heuristics for Mining Coding Agent Activity Romain Robes Théo Matricon et.al. 2601.18345 null
2026-01-26 Agentic Much? Adoption of Coding Agents on GitHub Romain Robbes et.al. 2601.18341 null
2026-01-26 MultiVis-Agent: A Multi-Agent Framework with Logic Rules for Reliable and Comprehensive Cross-Modal Data Visualization Jinwei Lu et.al. 2601.18320 null
2026-01-26 Temp-R1: A Unified Autonomous Agent for Complex Temporal KGQA via Reverse Curriculum Reinforcement Learning Zhaoyan Gong et.al. 2601.18296 null
2026-01-26 Reinforcement Learning with Distributed MPC for Fuel-Efficient Platoon Control with Discrete Gear Transitions Samuel Mallick et.al. 2601.18294 null
2026-01-23 Spatial-Agent: Agentic Geo-spatial Reasoning with Scientific Core Concepts Riyang Bao et.al. 2601.16965 null
2026-01-23 AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems Mohamed Amine Ferrag et.al. 2601.16964 null
2026-01-23 AI builds, We Analyze: An Empirical Study of AI-Generated Build Code Quality Anwar Ghammam et.al. 2601.16839 null
2026-01-23 Will It Survive? Deciphering the Fate of AI-Generated Code in Open Source Musfiqur Rahman et.al. 2601.16809 null
2026-01-23 SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents Yuhang Wang et.al. 2601.16746 null
2026-01-23 LongCat-Flash-Thinking-2601 Technical Report Meituan LongCat Team et.al. 2601.16725 null
2026-01-23 Adoption of Generative Artificial Intelligence in the German Software Engineering Industry: An Empirical Study Ludwig Felder et.al. 2601.16700 null
2026-01-23 AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning Suzhong Fu et.al. 2601.16685 null
2026-01-23 LUMINA: Long-horizon Understanding for Multi-turn Interactive Agents Amin Rakhsha et.al. 2601.16649 null
2026-01-23 A Cognitive Framework for Autonomous Agents: Toward Human-Inspired Design Francesco Guidi et.al. 2601.16648 null
2026-01-23 Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification Gaurav Maheshwari et.al. 2601.16530 null
2026-01-23 REprompt: Prompt Generation for Intelligent Software Development Guided by Requirements Engineering Junjie Shi et.al. 2601.16507 null
2026-01-23 EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration Xinshuai Guo et.al. 2601.16489 null
2026-01-23 Introducing the Generative Application Firewall (GAF) Joan Vendrell Farreny et.al. 2601.15824 null
2026-01-22 The Behavioral Fabric of LLM-Powered GUI Agents: Human Values and Interaction Outcomes Simret Araya Gebreegziabher et.al. 2601.16356 null
2026-01-22 When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems Donghao Huang et.al. 2601.16280 null
2026-01-22 Controlling Long-Horizon Behavior in Language Model Agents with Explicit State Dynamics Sukesh Subaharan et.al. 2601.16087 null
2026-01-22 Agentic Confidence Calibration Jiaxin Zhang et.al. 2601.15778 null
2026-01-22 UXCascade: Scalable Usability Testing with Simulated User Agents Steffen Holter et.al. 2601.15777 null
2026-01-22 VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning Chenglin Li et.al. 2601.15724 null
2026-01-22 AgentSM: Semantic Memory for Agentic Text-to-SQL Asim Biswal et.al. 2601.15709 null
2026-01-22 Agentic Uncertainty Quantification Jiaxin Zhang et.al. 2601.15703 null
2026-01-22 From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models Jiaxin Zhang et.al. 2601.15690 null
2026-01-22 Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats Ee Wei Seah et.al. 2601.15679 null
2026-01-22 Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors Zhiwei Zhang et.al. 2601.15625 null
2026-01-22 Autonomous Business System via Neuro-symbolic AI Cecil Pang et.al. 2601.15599 null
2026-01-22 Emerging from Ground: Addressing Intent Deviation in Tool-Using Agents via Deriving Real Calls into Virtual Trajectories Qian Xiong et.al. 2601.15120 null
2026-01-21 TransportAgents: a multi-agents LLM framework for traffic accident severity prediction Zhichao Yang et.al. 2601.15519 null
2026-01-21 Vibe Coding Kills Open Source Miklós Koren et.al. 2601.15494 null
2026-01-21 Taxonomy-Aligned Risk Extraction from 10-K Filings with Autonomous Improvement Using LLMs Rian Dolphin et.al. 2601.15247 null
2026-01-21 When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling Niful Islam et.al. 2601.15232 null
2026-01-21 Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub Ramtin Ehsani et.al. 2601.15195 null
2026-01-21 Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems Yinzhu Chen et.al. 2601.15161 null
2026-01-21 How to Build AI Agents by Augmenting LLMs with Codified Human Expert Domain Knowledge? A Software Engineering Framework Choro Ulan uulu et.al. 2601.15153 null
2026-01-21 CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Tianshi Xu et.al. 2601.15141 null
2026-01-21 From Who They Are to How They Act: Behavioral Traits in Generative Agent-Based Models of Social Media Valerio La Gatta et.al. 2601.15114 null
2026-01-21 Facilitating Proactive and Reactive Guidance for Decision Making on the Web: A Design Probe with WebSeek Yanwei Huang et.al. 2601.15100 null
2026-01-21 The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution Chen Qian et.al. 2601.15075 null
2026-01-21 Game-Theoretic Lens on LLM-based Multi-Agent Systems Jianing Hao et.al. 2601.15047 null
2026-01-21 Interoperable Architecture for Digital Identity Delegation for AI Agents with Blockchain Integration David Ricardo Saavedra et.al. 2601.14982 null
2026-01-21 CodeDelegator: Mitigating Context Pollution via Role Separation in Code-as-Action Agents Tianxiang Fei et.al. 2601.14914 null
2026-01-21 Optimizing FaaS Platforms for MCP-enabled Agentic Workflows Varad Kulkarni et.al. 2601.14735 null
2026-01-21 Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation Muhammad Khalifa et.al. 2601.14691 null
2026-01-21 INFA-Guard: Mitigating Malicious Propagation via Infection-Aware Safeguarding in LLM-Based Multi-Agent Systems Yijin Zhou et.al. 2601.14667 null
2026-01-21 NeuroFilter: Privacy Guardrails for Conversational LLM Agents Saswat Das et.al. 2601.14660 null
2026-01-21 MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks Zixuan Ke et.al. 2601.14652 null
2026-01-21 An LLM Agent-based Framework for Whaling Countermeasures Daisuke Miyamoto et.al. 2601.14606 null
2026-01-21 Holmes: An Evidence-Grounded LLM Agent for Auditable DDoS Investigation in Cloud Networks Haodong Chen et.al. 2601.14601 null
2026-01-20 XR: Cross-Modal Agents for Composed Image Retrieval Zhongyu Yang et.al. 2601.14245 null
2026-01-20 APEX-Agents Bertie Vidgen et.al. 2601.14242 null
2026-01-20 Toward Efficient Agents: Memory, Tool learning, and Planning Xiaofang Yang et.al. 2601.14192 null
2026-01-20 Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance Qianli Ma et.al. 2601.14171 null
2026-01-20 CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems Tong Xie et.al. 2601.14140 null
2026-01-20 Zero-shot adaptable task planning for autonomous construction robots: a comparative study of lightweight single and multi-AI agent systems Hossein Naderi et.al. 2601.14091 null
2026-01-20 LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems Badri N. Patro et.al. 2601.14053 null
2026-01-20 Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics Junqi Liu et.al. 2601.14027 null
2026-01-20 VirtualCrime: Evaluating Criminal Potential of Large Language Models via Sandbox Simulation Yilin Tang et.al. 2601.13981 null
2026-01-20 FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation Jing Zuo et.al. 2601.13976 null
2026-01-20 Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval Joaquín Polonuer et.al. 2601.13969 null
2026-01-20 VulnResolver: A Hybrid Agent Framework for LLM-Based Automated Vulnerability Issue Resolution Mingming Zhang et.al. 2601.13933 null
2026-01-20 Understanding Human-Multi-Agent Team Formation for Creative Work Hyunseung Lim et.al. 2601.13865 null
2026-01-20 Small Models, Big Impact: Tool-Augmented AI Agents for Wireless Network Planning Yongqiang Zhang et.al. 2601.13843 null
2026-01-20 HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction Yuhua Jin et.al. 2601.13801 null
2026-01-20 On Autopilot? An Empirical Study of Human-AI Teaming and Review Practices in Open Source Haoyu Gao et.al. 2601.13754 null
2026-01-20 SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories Aditya Bharat Soni et.al. 2601.13713 null
2026-01-20 Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games Christopher Kao et.al. 2601.13709 null
2026-01-20 Generative Intent Prediction Agentic AI empowered Edge Service Function Chain Orchestration Yan Sun et.al. 2601.13694 null
2026-01-20 Toward Agentic AI: Task-Oriented Communication for Hierarchical Planning of Long-Horizon Tasks Sin-Yu Huang et.al. 2601.13685 null
2026-01-16 Applying Formal Methods Tools to an Electronic Warfare Codebase (Experience report) Letitia W. Li et.al. 2601.11510 null
2026-01-16 The Poisoned Apple Effect: Strategic Manipulation of Mediated Markets via Technology Expansion of AI Agents Eilam Shapira et.al. 2601.11496 null
2026-01-16 The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents Ziyu Wang et.al. 2601.11421 null
2026-01-16 Can Small Agent Collaboration Beat a Single Big LLM? Agata Żywot et.al. 2601.11327 null
2026-01-16 Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation Pingzhi Tang et.al. 2601.11258 null
2026-01-16 Bayesian optimisation for Bayesian evidence (BOBE) – a fast and efficient likelihood emulator for model selection Nathan Cohen et.al. 2601.11150 null
2026-01-16 Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems Zixu Wang et.al. 2601.11147 null
2026-01-16 ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development Jie Yang et.al. 2601.11077 null
2026-01-16 AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts Keyu Li et.al. 2601.11044 null
2026-01-16 AJAR: Adaptive Jailbreak Architecture for Red-teaming Yipu Dou et.al. 2601.10971 null
2026-01-16 Modeling Multi-Party Interaction in Couples Therapy: A Multi-Agent Simulation Approach Canwen Wang et.al. 2601.10970 null
2026-01-16 Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents Kaiyu Zhou et.al. 2601.10955 null
2026-01-15 Multi-Agent Taint Specification Extraction for Vulnerability Detection Jonah Ghebremichael et.al. 2601.10865 null
2026-01-15 Towards Reliable ML Feature Engineering via Planning in Constrained-Topology of LLM Agents Himanshu Thakur et.al. 2601.10820 null
2026-01-15 Institutional AI: A Governance Framework for Distributional AGI Safety Federico Pierucci et.al. 2601.10599 null
2026-01-15 Mitigating GIL Bottlenecks in Edge AI Systems Mridankan Mandal et.al. 2601.10582 null
2026-01-15 From Single to Multi-Agent Reasoning: Advancing GeneGPT for Genomics QA Kimia Abedini et.al. 2601.10581 null
2026-01-15 Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment Felix Jahn et.al. 2601.10520 null
2026-01-15 AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior Nadya Abaev et.al. 2601.10440 null
2026-01-15 Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering Xinyu Zhu et.al. 2601.10402 null
2026-01-15 Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text Zhihao Xu et.al. 2601.10355 null
2026-01-15 OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding Deming Ding et.al. 2601.10343 null
2026-01-15 Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale Yi Liu et.al. 2601.10338 null
2026-01-15 coTherapist: A Behavior-Aligned Small Language Model to Support Mental Healthcare Experts Prottay Kumar Adhikary et.al. 2601.10246 null
2026-01-15 STEAMROLLER: A Multi-Agent System for Inclusive Automatic Speech Recognition for People who Stutter Ziqi Xu et.al. 2601.10223 null
2026-01-15 Autonomous Quantum Simulation through Large Language Model Agents Weitang Li et.al. 2601.10194 null
2026-01-15 ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback Yutao Mou et.al. 2601.10156 null
2026-01-15 M^4olGen: Multi-Agent, Multi-Stage Molecular Generation under Precise Multi-Property Constraints Yizhan Li et.al. 2601.10131 null
2026-01-15 Role-Playing Agents Driven by Large Language Models: Current Status, Challenges, and Future Trends Ye Wang et.al. 2601.10122 null
2026-01-15 Repository Intelligence Graph: Deterministic Architectural Map for LLM Code Assistants Tsvi Cherny-Shahar et.al. 2601.10112 null
2026-01-15 Collective behavior based on agent-environment interactions Gaston Briozzo et.al. 2601.10046 null
2026-01-15 PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization Tingyue Pan et.al. 2601.10029 null
2026-01-15 Structured Personality Control and Adaptation for LLM Agents Jinpeng Wang et.al. 2601.10025 null
2026-01-15 EHRNavigator: A Multi-Agent System for Patient-Level Clinical Question Answering over Heterogeneous Electronic Health Records Lingfei Qian et.al. 2601.10020 null
2026-01-14 LLMs can Compress LLMs: Adaptive Pruning by Agents Sai Varun Kodathala et.al. 2601.09694 null
2026-01-14 Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning Zhiyuan Hu et.al. 2601.09667 null
2026-01-14 LLM for Large-Scale Optimization Model Auto-Formulation: A Lightweight Few-Shot Learning Approach Kuo Liang et.al. 2601.09635 null
2026-01-14 The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multi-Step Malware Ben Nassi et.al. 2601.09625 null
2026-01-14 What Do LLM Agents Know About Their World? Task2Quiz: A Paradigm for Studying Environment Understanding Siyuan Liu et.al. 2601.09503 null
2026-01-14 Dissecting Judicial Reasoning in U.S. Copyright Damage Awards Pei-Chi Lo et.al. 2601.09459 null
2026-01-14 SC-MAS: Constructing Cost-Efficient Multi-Agent Systems with Edge-Level Heterogeneous Collaboration Di Zhao et.al. 2601.09434 null
2026-01-14 Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception Zhen Wan et.al. 2601.09413 null
2026-01-14 Improving Implicit Hate Speech Detection via a Community-Driven Multi-Agent Framework Ewelina Gajewska et.al. 2601.09342 null
2026-01-14 MACRO-LLM: LLM-Empowered Multi-Agent Collaborative Reasoning under Spatiotemporal Partial Observability Handi Chen et.al. 2601.09295 null
2026-01-14 Blue Teaming Function-Calling Agents Greta Dolcetti et.al. 2601.09292 null
2026-01-14 M $^3$ Searcher: Modular Multimodal Information Seeking Agency with Retrieval-Oriented Reasoning Xiaohan Yu et.al. 2601.09278 null
2026-01-14 Coordinated Pandemic Control with Large Language Model Agents as Policymaking Assistants Ziyi Shi et.al. 2601.09264 null
2026-01-14 MAXS: Meta-Adaptive Exploration with LLM Agents Jian Zhang et.al. 2601.09259 null
2026-01-14 Hybrid guided variational autoencoder for visual place recognition Ni Wang et.al. 2601.09248 null
2026-01-14 Honesty-Aware Multi-Agent Framework for High-Fidelity Synthetic Data Generation in Digital Psychiatric Intake Doctor-Patient Interactions Xinyuan Zhang et.al. 2601.09216 null
2026-01-14 PrivacyReasoner: Can LLM Emulate a Human-like Privacy Mind? Yiwen Tu et.al. 2601.09152 null
2026-01-14 World Craft: Agentic Framework to Create Visualizable Worlds via Text Jianwen Sun et.al. 2601.09150 null
2026-01-14 KryptoPilot: An Open-World Knowledge-Augmented LLM Agent for Automated Cryptographic Exploitation Xiaonan Liu et.al. 2601.09129 null
2026-01-14 The AI Hippocampus: How Far are We From Human Memory? Zixia Jia et.al. 2601.09113 null
2026-01-13 Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System Hsiang-Wei Huang et.al. 2601.08829 null
2026-01-13 Inferring Latent Intentions: Attributional Natural Language Inference in LLM Agents Xin Quan et.al. 2601.08742 null
2026-01-13 Data Product MCP: Chat with your Enterprise Data Marco Tonnarelli et.al. 2601.08687 null
2026-01-13 GraphSearch: Agentic Search-Augmented Reasoning for Zero-Shot Graph Learning Jiajin Liu et.al. 2601.08621 null
2026-01-13 ExpSeek: Self-Triggered Experience Seeking for Web Agents Wenyuan Zhang et.al. 2601.08605 null
2026-01-13 M3-BENCH: Process-Aware Evaluation of LLM Agents Social Behaviors in Mixed-Motive Games Sixiong Xie et.al. 2601.08462 null
2026-01-13 Hybrid Distillation with CoT Guidance for Edge-Drone Control Code Generation Yizhan Feng et.al. 2601.08412 null
2026-01-13 WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents Xinyi Wu et.al. 2601.08406 null
2026-01-13 A New Tool to Find Lightweight (And, Xor) Implementations of Quadratic Vectorial Boolean Functions up to Dimension 9 Marie Bolzer et.al. 2601.08368 null
2026-01-13 Semantic Laundering in AI Agent Architectures: Why Tool Boundaries Do Not Confer Epistemic Warrant Oleg Romanchuk et.al. 2601.08333 null
2026-01-13 Safe Heterogeneous Multi-Agent RL with Communication Regularization for Coordinated Target Acquisition Gabriele Calzolari et.al. 2601.08327 null
2026-01-13 AgriAgent: Contract-Driven Planning and Capability-Aware Tool Orchestration in Real-World Agriculture Bo Yang et.al. 2601.08308 null
2026-01-13 ToolACE-MCP: Generalizing History-Aware Routing from MCP Tools to the Agent Web Zhiyuan Yao et.al. 2601.08276 null
2026-01-13 Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees Kun Li et.al. 2601.08274 null
2026-01-13 User-Oriented Multi-Turn Dialogue Generation with Tool Use at scale Jungho Cho et.al. 2601.08225 null
2026-01-13 Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis Da Song et.al. 2601.08196 null
2026-01-13 Route, Retrieve, Reflect, Repair: Self-Improving Agentic Framework for Visual Detection and Linguistic Reasoning in Medical Imaging Md. Faiyaz Abdullah Sayeedi et.al. 2601.08192 null
2026-01-13 SwiftMem: Fast Agentic Memory via Query-aware Indexing Anxin Tian et.al. 2601.08160 null
2026-01-13 Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions Arin Gopalan Yadav et.al. 2601.08156 null
2026-01-12 MemoBrain: Executive Memory as an Agentic Brain for Reasoning Hongjin Qian et.al. 2601.08079 null
2026-01-12 Beyond Single-Shot: Multi-step Tool Retrieval via Query Planning Wei Fang et.al. 2601.07782 null
2026-01-12 Semisimple algebraic groups over real closed fields Raphael Appenzeller et.al. 2601.07732 null
2026-01-12 DIAGPaper: Diagnosing Valid and Specific Weaknesses in Scientific Papers via Multi-Agent Reasoning Zhuoyang Zou et.al. 2601.07611 null
2026-01-12 Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments Bingyang Ye et.al. 2601.07606 null
2026-01-12 VirtualEnv: A Platform for Embodied AI Research Kabir Swain et.al. 2601.07553 null
2026-01-12 FROAV: A Framework for RAG Observation and Agent Verification – Lowering the Barrier to LLM Agent Research Tzu-Hsuan Lin et.al. 2601.07504 null
2026-01-12 R3-RECON: Radiance-Field-Free Active Reconstruction via Renderability Xiaofeng Jin et.al. 2601.07484 null
2026-01-12 JudgeFlow: Agentic Workflow Optimization via Block Judge Zihan Ma et.al. 2601.07477 null
2026-01-12 Learning How to Remember: A Meta-Cognitive Management Method for Structured and Transferable Agent Memory Sirui Liang et.al. 2601.07470 null
2026-01-12 Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents Miao Su et.al. 2601.07468 null
2026-01-12 MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP Ruiqi Li et.al. 2601.07395 null
2026-01-12 OpenTinker: Separating Concerns in Agentic Reinforcement Learning Siqi Zhu et.al. 2601.07376 null
2026-01-12 Agentic Diagnostic Reasoning over Telecom and Datacenter Infrastructure Nicolas Tacheny et.al. 2601.07342 null
2026-01-12 ARM: Role-Conditioned Neuron Transplantation for Training-Free Generalist LLM Agent Merging Zhuoka Feng et.al. 2601.07309 null
2026-01-12 The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents Weihao Xuan et.al. 2601.07264 null
2026-01-12 When Bots Take the Bait: Exposing and Mitigating the Emerging Social Engineering Attack in Web Automation Agent Xinyi Wu et.al. 2601.07263 null
2026-01-12 ColorBrowserAgent: An Intelligent GUI Agent for Complex Long-Horizon Web Automation Jiamu Zhou et.al. 2601.07262 null
2026-01-12 Lost in the Noise: How Reasoning Models Fail with Contextual Distractors Seongyun Lee et.al. 2601.07226 null
2026-01-12 Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration Yang Zhao et.al. 2601.07224 null
2026-01-12 Active Context Compression: Autonomous Memory Management in LLM Agents Nikhil Verma et.al. 2601.07190 null
2026-01-09 Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards Jiajie Zhang et.al. 2601.06021 null
2026-01-09 Don’t Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks Elias Lumer et.al. 2601.06007 null
2026-01-09 Agentic LLMs as Powerful Deanonymizers: Re-identification of Participants in the Anthropic Interviewer Dataset Tianshi Li et.al. 2601.05918 null
2026-01-09 TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents Dawei Wang et.al. 2601.05899 null
2026-01-09 StackPlanner: A Centralized Hierarchical Multi-Agent System with Task-Experience Memory Management Ruizhe Zhang et.al. 2601.05890 null
2026-01-09 Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense Víctor Mayoral-Vilches et.al. 2601.05887 null
2026-01-09 EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis Xiaoshuai Song et.al. 2601.05808 null
2026-01-09 VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit Junda Lin et.al. 2601.05755 null
2026-01-09 LIDL: LLM Integration Defect Localization via Knowledge Graph-Enhanced Multi-Agent Analysis Gou Tan et.al. 2601.05539 null
2026-01-09 Task Cascades for Efficient Unstructured Data Processing Shreya Shankar et.al. 2601.05536 null
2026-01-09 CHisAgent: A Multi-Agent Framework for Event Taxonomy Construction in Ancient Chinese Cultural Systems Xuemei Tang et.al. 2601.05520 null
2026-01-09 Memory Poisoning Attack and Defense on Memory Based LLM-Agents Balachandra Devarangadi Sunil et.al. 2601.05504 null
2026-01-09 EvidFuse: Writing-Time Evidence Learning for Consistent Text-Chart Data Reporting Huanxiang Lin et.al. 2601.05487 null
2026-01-09 MMUEChange: A Generalized LLM Agent Framework for Intelligent Multi-Modal Urban Environment Change Analysis Zixuan Xiao et.al. 2601.05483 null
2026-01-09 STELP: Secure Transpilation and Execution of LLM-Generated Programs Swapnil Shinde et.al. 2601.05467 null
2026-01-09 MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents Tamil Sudaravan Mohan Doss et.al. 2601.05215 null
2026-01-08 Conformity and Social Impact on AI Agents Alessandro Bellina et.al. 2601.05384 null
2026-01-08 Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models Zheng Luo et.al. 2601.05366 null
2026-01-08 PRISM: Protocol Refinement through Intelligent Simulation Modeling Brian Hsu et.al. 2601.05356 null
2026-01-08 Generate, Transfer, Adapt: Learning Functional Dexterous Grasping from a Single Human Demonstration Xingyi He et.al. 2601.05243 null
2026-01-08 DocDancer: Towards Agentic Document-Grounded Information Seeking Qintong Zhang et.al. 2601.05163 null
2026-01-08 Agent-as-a-Judge Runyang You et.al. 2601.05111 null
2026-01-08 Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction Muzhao Tian et.al. 2601.05107 null
2026-01-08 Arabic Prompts with English Tools: A Benchmark Konstantin Kubrak et.al. 2601.05101 null
2026-01-08 Can Large Language Models Resolve Semantic Discrepancy in Self-Destructive Subcultures? Evidence from Jirai Kei Peng Wang et.al. 2601.05004 null
2026-01-08 Analyzing Message-Code Inconsistency in AI Coding Agent-Authored Pull Requests Jingzhi Gong et.al. 2601.04886 null
2026-01-08 Mind2Report: A Cognitive Deep Research Agent for Expert-Level Commercial Report Synthesis Mingyue Cheng et.al. 2601.04879 null
2026-01-08 Higher-Order Knowledge Representations for Agentic Scientific Reasoning Isabella A. Stewart et.al. 2601.04878 null
2026-01-08 Orchestrating Intelligence: Confidence-Aware Routing for Efficient Multi-Agent Collaboration across Multi-Scale Models Jingbo Wang et.al. 2601.04861 null
2026-01-08 RAAR: Retrieval Augmented Agentic Reasoning for Cross-Domain Misinformation Detection Zhiwei Liu et.al. 2601.04853 null
2026-01-08 Defense Against Indirect Prompt Injection via Tool Result Parsing Qiang Yu et.al. 2601.04795 null
2026-01-08 APEX: Academic Poster Editing Agentic Expert Chengxin Shi et.al. 2601.04794 null
2026-01-08 Belief in Authority: Impact of Authority in Multi-Agent Evaluation Framework Junhyuk Choi et.al. 2601.04790 null
2026-01-08 AT $^2$ PO: Agentic Turn-based Policy Optimization via Tree Search Zefang Zong et.al. 2601.04767 null
2026-01-08 When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail Xiaoxiao Li et.al. 2601.04748 null
2026-01-08 Tool-MAD: A Multi-Agent Debate Framework for Fact Verification with Diverse Tool Augmentation and Adaptive Retrieval Seyeon Jeong et.al. 2601.04742 null
2026-01-08 Beyond Monolithic Architectures: A Multi-Agent Search and Knowledge Optimization Framework for Agentic Search Yiqun Chen et.al. 2601.04703 null
2026-01-08 ResMAS: Resilience Optimization in LLM-based Multi-agent Systems Zhilun Zhou et.al. 2601.04694 null
2026-01-07 Embedding Autonomous Agents in Resource-Constrained Robotic Platforms Negar Halakou et.al. 2601.04191 null
2026-01-07 Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test Chun-Kai Fan et.al. 2601.04137 null
2026-01-07 InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training Ziyun Zhang et.al. 2601.04126 null
2026-01-07 ComfySearch: Autonomous Exploration and Reasoning for ComfyUI Workflows Jinwei Su et.al. 2601.04060 null
2026-01-07 Staged Voxel-Level Deep Reinforcement Learning for 3D Medical Image Segmentation with Noisy Annotations Yuyang Fu et.al. 2601.03875 null
2026-01-07 Atlas: Orchestrating Heterogeneous Models and Tools for Multi-Domain Complex Reasoning Jinyang Wu et.al. 2601.03872 null
2026-01-07 NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning Zhongtao Miao et.al. 2601.03790 null
2026-01-07 Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agents Dehao Tao et.al. 2601.03785 null
2026-01-07 PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation Wenlong Huang et.al. 2601.03782 null
2026-01-07 Agentic Proof Automation: A Case Study Yichen Xu et.al. 2601.03768 null
2026-01-07 O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL Yi Yao et.al. 2601.03743 null
2026-01-07 From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level Jia Li et.al. 2601.03731 null
2026-01-07 NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models Weiqi Liu et.al. 2601.03671 null
2026-01-07 Agent-Dice: Disentangling Knowledge Updates via Geometric Consensus for Agent Continual Learning Zheng Wu et.al. 2601.03641 null
2026-01-07 Architecting Agentic Communities using Design Patterns Zoran Milosevic et.al. 2601.03624 null
2026-01-07 The Pneuma Project: Reifying Information Needs as Relational Schemas to Automate Discovery, Guide Preparation, and Align Data with Intent Muhammad Imam Luthfi Balaka et.al. 2601.03618 null
2026-01-07 Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions Zhongbin Guo et.al. 2601.03590 null
2026-01-07 Do Autonomous Agents Contribute Test Code? A Study of Tests in Agentic Pull Requests Sabrina Haque et.al. 2601.03556 null
2026-01-07 SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models Yuxuan Jiang et.al. 2601.03555 null
2026-01-07 DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing Hongzhi Zhang et.al. 2601.03540 null
2026-01-06 Automated Semantic Rules Detection (ASRD) for Emergent Communication Interpretation Bastien Vanderplaetse et.al. 2601.03254 null
2026-01-06 MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents Dongming Jiang et.al. 2601.03236 null
2026-01-06 The Fake Friend Dilemma: Trust and the Political Economy of Conversational AI Jacob Erickson et.al. 2601.03222 null
2026-01-06 InfiAgent: An Infinite-Horizon Framework for General-Purpose Autonomous Agents Chenglin Yu et.al. 2601.03204 null
2026-01-06 Multi-Modal Data-Enhanced Foundation Models for Prediction and Control in Wireless Networks: A Survey Han Zhang et.al. 2601.03181 null
2026-01-06 Accurate Table Question Answering with Accessible LLMs Yangfan Jiang et.al. 2601.03137 null
2026-01-06 A Probabilistic Digital Twin of UK En Route Airspace for Training and Evaluating AI Agents for Air Traffic Control Nick Pepper et.al. 2601.03113 null
2026-01-06 Understanding Multi-Agent Reasoning with Large Language Models for Cartoon VQA Tong Wu et.al. 2601.03073 null
2026-01-06 A Fast Semidefinite Convex Relaxation for Optimal Control Problems With Spatio-Temporal Constraints Shiying Dong et.al. 2601.03055 null
2026-01-06 From inconsistency to decision: explainable operation and maintenance of battery energy storage systems Jingbo Qu et.al. 2601.03007 null
2026-01-06 Causal-Enhanced AI Agents for Medical Research Screening Duc Ngo et.al. 2601.02814 null
2026-01-06 LLM Agent Framework for Intelligent Change Analysis in Urban Environment using Remote Sensing Imagery Zixuan Xiao et.al. 2601.02757 null
2026-01-06 The Path Ahead for Agentic AI: Challenges and Opportunities Nadia Sibai et.al. 2601.02749 null
2026-01-06 SYNAPSE: Empowering LLM Agents with Episodic-Semantic Memory via Spreading Activation Hanqi Jiang et.al. 2601.02744 null
2026-01-06 Agentic Memory Enhanced Recursive Reasoning for Root Cause Localization in Microservices Lingzhe Zhang et.al. 2601.02732 null
2026-01-06 EvoRoute: Experience-Driven Self-Routing LLM Agent Systems Guibin Zhang et.al. 2601.02695 null
2026-01-05 LongDA: Benchmarking LLM Agents for Long-Document Data Analysis Yiyang Li et.al. 2601.02598 null
2026-01-05 Orchestral AI: A Framework for Agent Orchestration Alexander Roman et.al. 2601.02577 null
2026-01-05 PerspectiveCoach: Exploring LLMs for Developer Reflection Lauren Olson et.al. 2601.02559 null
2026-01-05 SimpleMem: Efficient Lifelong Memory for LLM Agents Jiaqi Liu et.al. 2601.02553 null
2026-01-05 Project Ariadne: A Structural Causal Framework for Auditing Faithfulness in LLM Agents Sourena Khanzadeh et.al. 2601.02314 null
2026-01-05 Code for Machines, Not Just Humans: Quantifying AI-Friendliness with Code Health Metrics Markus Borg et.al. 2601.02200 null
2026-01-05 Confidence Estimation for LLMs in Multi-turn Interactions Caiqi Zhang et.al. 2601.02179 null
2026-01-05 Finite-State Decentralized Policy-Based Control With Guaranteed Ground Coverage Hossein Rastgoftar et.al. 2601.02109 null
2026-01-05 Agentic AI in Remote Sensing: Foundations, Taxonomy, and Emerging Systems Niloufar Alipour Talemi et.al. 2601.01891 null
2026-01-05 Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents Yi Yu et.al. 2601.01885 null
2026-01-05 Toward Auditable Neuro-Symbolic Reasoning in Pathology: SQL as an Explicit Trace of Evidence Kewen Cao et.al. 2601.01875 null
2026-01-05 Jenius Agent: Towards Experience-Driven Accuracy Optimization in Real-World Scenarios Defei Xia et.al. 2601.01857 null
2026-01-05 ARIES: A Scalable Multi-Agent Orchestration Framework for Real-Time Epidemiological Surveillance and Outbreak Monitoring Aniket Wattamwar et.al. 2601.01831 null
2026-01-05 AI Agent Systems: Architectures, Applications, and Evaluation Bin Xu et.al. 2601.01743 null
2026-01-05 Structural Representations for Cross-Attack Generalization in AI Agent Threat Detection Vignesh Iyer et.al. 2601.01723 null
2026-01-05 Explicit World Models for Reliable Human-Robot Collaboration Kenneth Kwok et.al. 2601.01705 null
2026-01-04 Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage Jinwei Hu et.al. 2601.01685 null
2026-01-04 Exposing Hidden Interfaces: LLM-Guided Type Inference for Reverse Engineering macOS Private Frameworks Arina Kharlamova et.al. 2601.01673 null
2026-01-04 CaveAgent: Transforming LLMs into Stateful Runtime Operators Maohao Ran et.al. 2601.01569 null
2026-01-04 Bayesian Orchestration of Multi-LLM Agents for Cost-Aware Sequential Decision-Making Danial Amin et.al. 2601.01522 null
2026-01-04 From Failure to Mastery: Generating Hard Samples for Tool-use Agents Bingguang Hao et.al. 2601.01498 null
2026-01-04 KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models Zixian Liu et.al. 2601.01366 null
2026-01-04 Towards LLM-enabled autonomous combustion research: A literature-aware agent for self-corrective modeling workflows Ke Xiao et.al. 2601.01357 null
2026-01-04 Adaptive Hierarchical Evaluation of LLMs and SAST tools for CWE Prediction in Python Muntasir Adnan et.al. 2601.01320 null
2026-01-02 LLM Agents for Combinatorial Efficient Frontiers: Investment Portfolio Optimization Simon Paquette-Greenbaum et.al. 2601.00770 null
2026-01-02 Early-Stage Prediction of Review Effort in AI-Generated Pull Requests Dao Sy Duy Minh et.al. 2601.00753 null
2026-01-02 An Agentic Framework for Neuro-Symbolic Programming Aliakbar Nafar et.al. 2601.00743 null
2026-01-02 Beyond IVR: Benchmarking Customer Support LLM Agents for Business-Adherence Sumanth Balaji et.al. 2601.00596 null
2026-01-02 Trajectory Guard – A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI Laksh Advani et.al. 2601.00516 null
2026-01-01 When Small Models Are Right for Wrong Reasons: Process Verification for Trustworthy Agents Laksh Advani et.al. 2601.00513 null
2026-01-01 Multi-Agent Coordinated Rename Refactoring Abhiram Bellur et.al. 2601.00482 null
2026-01-01 MAESTRO: Multi-Agent Evaluation Suite for Testing, Reliability, and Observability Tie Ma et.al. 2601.00481 null
2026-01-01 Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub Mohammed Latif Siddiq et.al. 2601.00477 null
2026-01-01 Progressive Ideation using an Agentic AI Framework for Human-AI Co-Creation Sankar B et.al. 2601.00475 null
2026-01-01 Space Debris Removal using Nano-Satellites controlled by Low-Power Autonomous Agents Dennis Christmann et.al. 2601.00465 null
2026-01-01 OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning Liuxiang Qiu et.al. 2601.00352 null
2026-01-01 Can Optimal Transport Improve Federated Inverse Reinforcement Learning? David Millard et.al. 2601.00309 null
2026-01-01 ClinicalReTrial: A Self-Evolving AI Agent for Clinical Trial Protocol Optimization Sixue Xing et.al. 2601.00290 null
2026-01-01 Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity Doyoung Kim et.al. 2601.00268 null
2026-01-01 Will LLM-powered Agents Bias Against Humans? Exploring the Belief-Dependent Vulnerability Zongwei Wang et.al. 2601.00240 null
2026-01-01 FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems Shanli Xing et.al. 2601.00227 null
2026-01-01 μACP: A Formal Calculus for Expressive, Resource-Constrained Agent Communication Arnab Mallick et.al. 2601.00219 null
2026-01-01 Understanding Security Risks of AI Agents’ Dependency Updates Tanmay Singla et.al. 2601.00205 null
2025-12-31 Ask, Clarify, Optimize: Human-LLM Agent Collaboration for Smarter Inventory Control Yaqi Duan et.al. 2601.00121 null
2025-12-31 Context-aware LLM-based AI Agents for Human-centered Energy Management Systems in Smart Buildings Tianzhi He et.al. 2512.25055 null
2025-12-31 MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes Siddhant Agarwal et.al. 2512.25015 null
2025-12-31 DarkEQA: Benchmarking Vision-Language Models for Embodied Question Answering in Low-Light Indoor Environments Yohan Park et.al. 2512.24985 null
2025-12-31 Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem Weixun Wang et.al. 2512.24873 null
2025-12-31 VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents Xunyi Zhao et.al. 2512.24851 null
2025-12-31 PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI Srija Mukhopadhyay et.al. 2512.24848 null
2025-12-31 AstroReview: An LLM-driven Multi-Agent Framework for Telescope Proposal Peer Review and Refinement Yutong Wang et.al. 2512.24754 null
2025-12-31 R-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory Maoyuan Li et.al. 2512.24684 null
2025-12-31 Do Large Language Models Know What They Are Capable Of? Casey O. Barkan et.al. 2512.24661 null
2025-12-31 AudioFab: Building A General and Intelligent Audio Factory through Tool Learning Cheng Zhu et.al. 2512.24645 null
2025-12-31 How Do Agentic AI Systems Deal With Software Energy Concerns? A Pull Request-Based Study Tanjum Motin Mitul et.al. 2512.24636 null
2025-12-31 ReflecToMeet: An AI-Assisted Reflection Based System to Enhance Collaborative Preparedness Md Nazmus Sakib et.al. 2512.24632 null
2025-12-31 How Do Agentic AI Systems Address Performance Optimizations? A BERTopic-Based Analysis of Pull Requests Md Nahidul Islam Opu et.al. 2512.24630 null
2025-12-31 Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models Junru Lu et.al. 2512.24618 null
2025-12-31 Youtu-Agent: Scaling Agent Productivity with Automated Generation and Hybrid Policy Optimization Yuchen Shi et.al. 2512.24615 null
2025-12-31 Reinforcement Learning-Augmented LLM Agents for Collaborative Decision Making and Performance Optimization Dong Qiu et.al. 2512.24609 null
2025-12-31 MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use Wenrui Liu et.al. 2512.24565 null
2025-12-30 Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied Agents Seohui Bae et.al. 2512.24461 null
2025-12-30 Language Model Agents Under Attack: A Cross Model-Benchmark of Profit-Seeking Behaviors in Customer Service Jingyu Zhang et.al. 2512.24415 null
2025-12-30 SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning Yong Xien Chng et.al. 2512.24330 null
2025-12-29 Nested Browser-Use Learning for Agentic Information Seeking Baixuan Li et.al. 2512.23647 null
2025-12-29 Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing Yuwen Li et.al. 2512.23611 null
2025-12-29 Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks Toqeer Ali Syed et.al. 2512.23557 null
2025-12-29 Why AI Safety Requires Uncertainty, Incomplete Preferences, and Non-Archimedean Utilities Alessio Benavoli et.al. 2512.23508 null
2025-12-29 AdaptiFlow: An Extensible Framework for Event-Driven Autonomy in Cloud Microservices Brice Arléon Zemtsop Ndadji et.al. 2512.23499 null
2025-12-29 Optimal Scalability-Aware Allocation of Swarm Robots: From Linear to Retrograde Performance via Marginal Gains Simay Atasoy Bingöl et.al. 2512.23431 null
2025-12-29 AKG kernel Agent: A Multi-Agent Framework for Cross-Platform Kernel Synthesis Jinye Du et.al. 2512.23424 null
2025-12-29 MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning Jiawei Chen et.al. 2512.23412 null
2025-12-29 A Design Space for Intelligent Agents in Mixed-Initiative Visual Analytics Tobias Stähle et.al. 2512.23372 null
2025-12-29 AGRO-SQL: Agentic Group-Relative Optimization with High-Fidelity Data Synthesis Cehua Yang et.al. 2512.23366 null
2025-12-29 AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents Jiafeng Liang et.al. 2512.23343 null
2025-12-29 CubeBench: Diagnosing Interactive, Long-Horizon Spatial Reasoning Under Partial Observations Huan-ang Gao et.al. 2512.23328 null
2025-12-29 AI4Reading: Chinese Audiobook Interpretation System Based on Multi-Agent Collaboration Minjiang Huang et.al. 2512.23300 null
2025-12-29 TCEval: Using Thermal Comfort to Assess Cognitive and Perceptual Abilities of AI Jingming Li et.al. 2512.23217 null
2025-12-29 SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search Yifan Zhang et.al. 2512.23167 null
2025-12-29 Multi-Agent Framework for Threat Mitigation and Resilience in AI-Based Systems Armstrong Foundjem et.al. 2512.23132 null
2025-12-29 It’s a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents Karolina Korgul et.al. 2512.23128 null
2025-12-28 Accelerating Language Model Workflows with Prompt Choreography TJ Bai et.al. 2512.23049 null
2025-12-28 Video-BrowseComp: Benchmarking Agentic Video Research on Open Web Zhengyang Liang et.al. 2512.23044 null
2025-12-28 DECEPTICON: How Dark Patterns Manipulate Web Agents Phil Cuvin et.al. 2512.22894 null
2025-12-26 Agentic Structured Graph Traversal for Root Cause Analysis of Code-related Incidents in Cloud Applications Shengkun Cui et.al. 2512.22113 null
2025-12-26 A2P-Vis: an Analyzer-to-Presenter Agentic Pipeline for Visual Insights Generation and Reporting Shuyu Gan et.al. 2512.22101 null
2025-12-26 MAI-UI Technical Report: Real-World Centric Foundation GUI Agents Hanzhang Zhou et.al. 2512.22047 null
2025-12-26 SWE-RM: Execution-free Feedback For Software Engineering Agents KaShun Shum et.al. 2512.21919 null
2025-12-26 SpatialBench: Can Agents Analyze Real-World Spatial Biology Data? Kenny Workman et.al. 2512.21907 null
2025-12-26 Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space Weichen Zhang et.al. 2512.21887 null
2025-12-26 MASFIN: A Multi-Agent System for Decomposed Financial Reasoning and Forecasting Marc S. Montalvo et.al. 2512.21878 null
2025-12-25 Accelerating Scientific Discovery with Autonomous Goal-evolving Agents Yuanqi Du et.al. 2512.21782 null
2025-12-25 How Do Agents Perform Code Optimization? An Empirical Study Huiyun Peng et.al. 2512.21757 null
2025-12-25 PERELMAN: Pipeline for scientific literature meta-analysis. Technical report Daniil Sherki et.al. 2512.21727 null
2025-12-25 HELP: Hierarchical Embodied Language Planner for Household Tasks Alexandr V. Korchemnyi et.al. 2512.21723 null
2025-12-25 Multiconnectivity for SAGIN: Current Trends, Challenges, AI-driven Solutions, and Opportunities Abd Ullah Khan et.al. 2512.21717 null
2025-12-25 AstraNav-World: World Model for Foresight Control and Consistency Junjun Hu et.al. 2512.21714 null
2025-12-25 RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention Zhan Chen et.al. 2512.21710 null
2025-12-25 Towards Responsible and Explainable AI Agents with Consensus-Driven Reasoning Eranga Bandara et.al. 2512.21699 null
2025-12-25 AstraNav-Memory: Contexts Compression for Long Memory Botao Ren et.al. 2512.21627 null
2025-12-25 AMS-IO-Bench and AMS-IO-Agent: Benchmarking and Structured Reasoning for Analog and Mixed-Signal Integrated Circuit Input/Output Design Zhishuai Zhang et.al. 2512.21613 null
2025-12-25 Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations Xin Liu et.al. 2512.21586 null
2025-12-25 The AI Committee: A Multi-Agent Framework for Automated Validation and Remediation of Web-Sourced Data Sunith Vallabhaneni et.al. 2512.21481 null
2025-12-24 Fuzzwise: Intelligent Initial Corpus Generation for Fuzzing Hridya Dhulipala et.al. 2512.21440 null
2025-12-24 Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Consulting, Data Analyst, and Management Tasks Ali Merali et.al. 2512.21316 null
2025-12-24 ReaSeq: Unleashing World Knowledge via Reasoning for Sequential Modeling Chuan Wang et.al. 2512.21257 null
2025-12-24 CoTDeceptor:Adversarial Code Obfuscation Against CoT-Enhanced LLM Code Agents Haoyang Li et.al. 2512.21250 null
2025-12-24 RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic Le Wang et.al. 2512.21220 null
2025-12-24 Agentic Explainable Artificial Intelligence (Agentic XAI) Approach To Explore Better Explanation Tomoaki Yamaguchi et.al. 2512.21066 null
2025-12-24 LLM-Empowered Agentic AI for QoE-Aware Network Slicing Management in Industrial IoT Xudong Wang et.al. 2512.20997 null
2025-12-24 TrafficSimAgent: A Hierarchical Agent Framework for Autonomous Traffic Simulation with MCP Control Yuwei Du et.al. 2512.20996 null
2025-12-24 AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs Yihan Wang et.al. 2512.20986 null
2025-12-24 SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking Yujin Noh et.al. 2512.20975 null
2025-12-24 DAO-Agent: Zero Knowledge-Verified Incentives for Decentralized Multi-Agent Coordination Yihan Xia et.al. 2512.20973 null
2025-12-24 One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents Zhaoxi Zhang et.al. 2512.20957 null
2025-12-24 ETP-R1: Evolving Topological Planning with Reinforcement Fine-tuning for Vision-Language Navigation in Continuous Environments Shuhao Ye et.al. 2512.20940 null
2025-12-24 Reasoning-Driven Amodal Completion: Collaborative Agents and Perceptual Evaluation Hongxing Fan et.al. 2512.20936 null
2025-12-24 Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning Shengguang Wu et.al. 2512.20934 null
2025-12-24 Embodied AI-Enhanced IoMT Edge Computing: UAV Trajectory Optimization and Task Offloading with Mobility Prediction Siqi Mu et.al. 2512.20902 null
2025-12-24 The Silent Scholar Problem: A Probabilistic Framework for Breaking Epistemic Asymmetry in LLM Agents Zan-Kai Chong et.al. 2512.20884 null
2025-12-24 Better Call Graphs: A New Dataset of Function Call Graphs for Malware Classification Jakir Hossain et.al. 2512.20872 null
2025-12-24 NVIDIA Nemotron 3: Efficient and Open Intelligence NVIDIA et.al. 2512.20856 null
2025-12-23 Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning NVIDIA et.al. 2512.20848 null
2025-12-23 MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs Onat Ozer et.al. 2512.20845 null
2025-12-23 LongVideoAgent: Multi-Agent Reasoning with Long Videos Runtao Liu et.al. 2512.20618 null
2025-12-23 Step-DeepResearch Technical Report Chen Hu et.al. 2512.20491 null
2025-12-23 Bohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at Scale Linfeng Zhang et.al. 2512.20469 null
2025-12-23 Laser: Governing Long-Horizon Agentic Search via Structured Protocol and Context Register Shuting Wang et.al. 2512.20458 null
2025-12-23 Chain-of-Anomaly Thoughts with Large Vision-Language Models Pedro Domingos et.al. 2512.20417 null
2025-12-23 CRAFT: Continuous Reasoning and Agentic Feedback Tuning for Multimodal Text-to-Image Generation V. Kovalev et.al. 2512.20362 null
2025-12-23 AprielGuard Jaykumar Kasundra et.al. 2512.20293 null
2025-12-23 SlideTailor: Personalized Presentation Slide Generation for Scientific Papers Wenzheng Zeng et.al. 2512.20292 null
2025-12-23 Synthesizing Procedural Memory: Challenges and Architectures in Automated Workflow Generation Nishant Gaurav et.al. 2512.20278 null
2025-12-23 Graph-Symbolic Policy Enforcement and Control (G-SPEC): A Neuro-Symbolic Framework for Safe Agentic AI in 5G Autonomous Networks Divya Vijay et.al. 2512.20275 null
2025-12-23 MemR $^3$ : Memory Retrieval via Reflective Reasoning for LLM Agents Xingbo Du et.al. 2512.20237 null
2025-12-23 TongSIM: A General Platform for Simulating Intelligent Machines Zhe Sun et.al. 2512.20206 null
2025-12-23 Reaching Agreement Among Reasoning LLM Agents Chaoyi Ruan et.al. 2512.20184 null
2025-12-23 Fun-Audio-Chat Technical Report Qian Chen et.al. 2512.20156 null
2025-12-23 MolAct: An Agentic RL Framework for Molecular Editing and Property Optimization Zhuo Yang et.al. 2512.20135 null
2025-12-23 ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language Aly Lidayan et.al. 2512.20111 null
2025-12-23 Detecting Non-Optimal Decisions of Embodied Agents via Diversity-Guided Metamorphic Testing Wenzhao Wu et.al. 2512.20083 null
2025-12-23 S $^3$ IT: A Benchmark for Spatially Situated Social Intelligence Test Zhe Sun et.al. 2512.19992 null
2025-12-22 PRISM: A Personality-Driven Multi-Agent Framework for Social Media Simulation Zhixiang Lu et.al. 2512.19933 null
2025-12-22 HARMON-E: Hierarchical Agentic Reasoning for Multimodal Oncology Notes to Extract Structured Data Shashi Kant Gupta et.al. 2512.19864 null
2025-12-22 GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators Jiacheng Guo et.al. 2512.19682 null
2025-12-22 LeLaR: The First In-Orbit Demonstration of an AI-Based Satellite Attitude Controller Kirill Djebko et.al. 2512.19576 null
2025-12-22 Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios Jiawen Wang et.al. 2512.19551 null
2025-12-22 QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models Li Puyin et.al. 2512.19526 null
2025-12-22 An Agentic Framework for Autonomous Materials Computation Zeyu Xia et.al. 2512.19458 null
2025-12-22 MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive, and MCP-Augmented Environments Quyu Kong et.al. 2512.19432 null
2025-12-22 EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration Runze Li et.al. 2512.19396 null
2025-12-22 Helios: A Foundational Language Model for Smart Energy Knowledge Reasoning and Application Haoyu Jiang et.al. 2512.19299 null
2025-12-22 Vibe Reasoning: Eliciting Frontier AI Mathematical Capabilities – A Case Study on IMO 2025 Problem 6 Jiaao Wu et.al. 2512.19287 null
2025-12-22 DeliveryBench: Can Agents Earn Profit in Real World? Lingjun Mao et.al. 2512.19234 null
2025-12-22 AWPO: Enhancing Tool-Use of Large Language Models through Explicit Integration of Reasoning Rewards Zihan Lin et.al. 2512.19126 null
2025-12-22 Tool-Augmented Hybrid Ensemble Reasoning with Distillation for Bilingual Mathematical Problem Solving Peiqing Lu et.al. 2512.19093 null
2025-12-22 $γ(3,4)$ `Attention’ in Cognitive Agents: Ontology-Free Knowledge Representations With Promise Theoretic Semantics Mark Burgess et.al. 2512.19084 null
2025-12-22 PEAK: A Performance Engineering AI-Assistant for GPU Kernels Powered by Natural Language Transformations Muhammad Usman Tariq et.al. 2512.19018 null
2025-12-22 DREAM: Dynamic Red-teaming across Environments for AI Models Liming Lu et.al. 2512.19016 null
2025-12-22 ORPR: An OR-Guided Pretrain-then-Reinforce Learning Model for Inventory Management Lingjie Zhao et.al. 2512.19001 null
2025-12-22 Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement Saman Forouzandeh et.al. 2512.18950 null
2025-12-21 Modular Automatic Complexity Analysis of Recursive Integer Programs Nils Lommen et.al. 2512.18851 null
2025-12-21 InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search Kaican Li et.al. 2512.18745 null
2025-12-21 X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System Zhanxun Liu et.al. 2512.18706 null
2025-12-19 XAgen: An Explainability Tool for Identifying and Correcting Failures in Multi-Agent Workflows Xinru Wang et.al. 2512.17896 null
2025-12-19 ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges Roshan Kenia et.al. 2512.17838 null
2025-12-19 Systemic Risks of Interacting AI Paul Darius et.al. 2512.17793 null
2025-12-19 A Practical Solution to Systematically Monitor Inconsistencies in SBOM-based Vulnerability Scanners Martin Rosso et.al. 2512.17710 null
2025-12-19 Binding Agent ID: Unleashing the Power of AI Agents with accountability and credibility Zibin Lin et.al. 2512.17538 null
2025-12-19 Are Vision Language Models Cross-Cultural Theory of Mind Reasoners? Zabir Al Nazi et.al. 2512.17394 null
2025-12-19 CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning Qi Song et.al. 2512.17312 null
2025-12-19 DAVE: A VLM Vision Encoder for Document Understanding and Web Agents Brandon Huang et.al. 2512.17221 null
2025-12-19 Conservative Bias in Multi-Teacher Learning: Why Agents Prefer Low-Reward Advisors Maher Mesto et.al. 2512.17180 null
2025-12-19 Biosecurity-Aware AI: Agentic Risk Auditing of Soft Prompt Attacks on ESM-Based Variant Predictors Huixin Zhan et.al. 2512.17146 null
2025-12-18 Verifying Hadwiger’s Conjecture for Examples of Graphs with $α(G) = 2$ Jofre Costa et.al. 2512.17114 null
2025-12-18 On the Role of Contextual Information and Ego States in LLM Agent Behavior for Transactional Analysis Dialogues Monika Zamojska et.al. 2512.17060 null
2025-12-18 Dynamic Tool Dependency Retrieval for Efficient Function Calling Bhrij Patel et.al. 2512.17052 null
2025-12-18 Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs Junbo Li et.al. 2512.17008 null
2025-12-18 A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos Mohammed Irfan Kurpath et.al. 2512.16978 null
2025-12-18 AdaTooler-V: Adaptive Tool-Use for Images and Videos Chaoyang Wang et.al. 2512.16918 null
2025-12-18 MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning Yuanchen Ju et.al. 2512.16909 null
2025-12-18 Distributional AGI Safety Nenad Tomašev et.al. 2512.16856 null
2025-12-18 Meta-RL Induces Exploration in Language Agents Yulun Jiang et.al. 2512.16848 null
2025-12-18 MEPIC: Memory Efficient Position Independent Caching for LLM Serving Qian Wang et.al. 2512.16822 null
2025-12-18 Coordinated Anti-Jamming Resilience in Swarm Networks via Multi-Agent Reinforcement Learning Bahman Abolhassani et.al. 2512.16813 null
2025-12-18 Do Multi-Agents Solve Better Than Single? Evaluating Agentic Frameworks for Diagram-Grounded Geometry Problem Solving and Reasoning Mahbub E Sobhani et.al. 2512.16698 null
2025-12-18 A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection Xiao Li et.al. 2512.16538 null
2025-12-18 From Personalization to Prejudice: Bias and Discrimination in Memory-Enhanced AI Agents for Recruitment Himanshu Gharat et.al. 2512.16532 null
2025-12-18 Plain language adaptations of biomedical text using LLMs: Comparision of evaluation metrics Primoz Kocbek et.al. 2512.16530 null
2025-12-18 cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution Jinwu Chen et.al. 2512.16465 null
2025-12-18 A Network Arena for Benchmarking AI Agents on Network Troubleshooting Zhihao Wang et.al. 2512.16381 null
2025-12-18 Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation Yuxuan Qiao et.al. 2512.16310 null
2025-12-18 Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection Fanrui Zhang et.al. 2512.16300 null
2025-12-18 Love, Lies, and Language Models: Investigating AI’s Role in Romance-Baiting Scams Gilad Gressel et.al. 2512.16280 null
2025-12-18 AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding Sanjoy Chowdhury et.al. 2512.16250 null
2025-12-18 PDE-Agent: A toolchain-augmented multi-agent framework for PDE solving Jianming Liu et.al. 2512.16214 null
2025-12-18 Open Ad-hoc Categorization with Contextualized Feature Learning Zilin Wang et.al. 2512.16202 null
2025-12-18 Ev-Trust: A Strategy Equilibrium Trust Mechanism for Evolutionary Games in LLM-Based Multi-Agent Services Shiduo Yang et.al. 2512.16167 null
2025-12-18 ToolForge: A Data Synthesis Pipeline for Multi-Hop Search without Real-World APIs Hao Chen et.al. 2512.16149 null
2025-12-17 Artism: AI-Driven Dual-Engine System for Art Generation and Critique Shuai Liu et.al. 2512.15710 null
2025-12-17 BashArena: A Control Setting for Highly Privileged AI Agents Adam Kaufman et.al. 2512.15688 null
2025-12-17 Mapis: A Knowledge-Graph Grounded Multi-Agent Framework for Evidence-Based PCOS Diagnosis Zanxiang He et.al. 2512.15398 null
2025-12-17 SCOPE: Prompt Evolution for Enhancing Agent Effectiveness Zehua Pei et.al. 2512.15374 null
2025-12-17 Revisiting Task-Oriented Dataset Search in the Era of Large Language Models: Challenges, Benchmark, and Solution Zixin Wei et.al. 2512.15363 null
2025-12-17 SynthSeg-Agents: Multi-Agent Synthetic Data Generation for Zero-Shot Weakly Supervised Semantic Segmentation Wangyu Wu et.al. 2512.15310 null
2025-12-17 Automatic generation of input files with optimised k-point meshes for Quantum Espresso self-consistent field single point total energy calculations Elena Patyukova et.al. 2512.15303 null
2025-12-17 CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications Zhengchao Chen et.al. 2512.15231 null
2025-12-17 MCPZoo: A Large-Scale Dataset of Runnable Model Context Protocol Servers for AI Agent Mengying Wu et.al. 2512.15144 null
2025-12-16 AgroAskAI: A Multi-Agentic AI Framework for Supporting Smallholder Farmers’ Enquiries Globally Nadine Angela Cantonjos et.al. 2512.14910 null
2025-12-16 Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections Niklas Lauffer et.al. 2512.14895 null
2025-12-16 MALCDF: A Distributed Multi-Agent LLM Framework for Real-Time Cyber Arth Bhardwaj et.al. 2512.14846 null
2025-12-16 Beyond Text-to-SQL: Autonomous Research-Driven Database Exploration with DAR Ostap Vykhopen et.al. 2512.14622 null
2025-12-16 Model-First Reasoning LLM Agents: Reducing Hallucinations through Explicit Problem Modeling Annu Rana et.al. 2512.14474 null
2025-12-16 Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space Xingfu Zhou et.al. 2512.14448 null
2025-12-16 A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning Zixin Zhang et.al. 2512.14442 null
2025-12-16 Multi-Agent Medical Decision Consensus Matrix System: An Intelligent Collaborative Framework for Oncology MDT Consultations Xudong Han et.al. 2512.14321 null
2025-12-16 From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition Yiqing Zhou et.al. 2512.14244 null
2025-12-16 PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design Ruozhao Yang et.al. 2512.14233 null
2025-12-16 IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol Yunhao Yao et.al. 2512.14166 null
2025-12-16 Incentivizing Tool-augmented Thinking with Images for Medical Image Analysis Yankai Jiang et.al. 2512.14157 null
2025-12-16 Astraea: A State-Aware Scheduling Engine for LLM-Powered Agents Hongqiu Ni et.al. 2512.14142 null
2025-12-16 From Obfuscated to Obvious: A Comprehensive JavaScript Deobfuscation Tool for Security Analysis Dongchao Zhou et.al. 2512.14070 null
2025-12-16 MobileWorldBench: Towards Semantic World Modeling For Mobile Agents Shufan Li et.al. 2512.14014 null
2025-12-16 Professional Software Developers Don’t Vibe, They Control: AI Agent Use for Coding in 2025 Ruanqianqian Huang et.al. 2512.14012 null
2025-12-16 FocalComm: Hard Instance-Aware Multi-Agent Perception Dereje Shenkut et.al. 2512.13982 null
2025-12-15 Olmo 3 Team Olmo et.al. 2512.13961 null
2025-12-15 Multi-Agent Collaborative Framework for Intelligent IT Operations: An AOI System with Context-Aware Compression and Dynamic Task Scheduling Zishan Bai et.al. 2512.13956 null
2025-12-15 Hierarchical Multi-agent Large Language Model Reasoning for Autonomous Functional Materials Discovery Samuel Rothfarb et.al. 2512.13930 null
2025-12-15 Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors Henger Li et.al. 2512.13860 null
2025-12-15 AgentIAD: Tool-Augmented Single-Agent for Industrial Anomaly Detection Junwen Miao et.al. 2512.13671 null
2025-12-15 Memory in the Age of AI Agents Yuyang Hu et.al. 2512.13564 null
2025-12-15 Async Control: Stress-testing Asynchronous Control Measures for LLM Agents Asa Cooper Stickland et.al. 2512.13526 null
2025-12-15 From User Interface to Agent Interface: Efficiency Optimization of UI Representations for LLM Agents Dezhi Ran et.al. 2512.13438 null
2025-12-15 Differentiable Evolutionary Reinforcement Learning Sitao Cheng et.al. 2512.13399 null
2025-12-15 MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data Zhenghao Zhu et.al. 2512.13297 null
2025-12-15 AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning Jiaru Zou et.al. 2512.13278 null
2025-12-15 Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection Juil Koo et.al. 2512.13250 null
2025-12-15 Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows Haoyu Dong et.al. 2512.13168 null
2025-12-15 SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning Emre Can Acikgoz et.al. 2512.13159 null
2025-12-15 MAC: A Multi-Agent Framework for Interactive User Clarification in Multi-turn Conversations Emre Can Acikgoz et.al. 2512.13154 null
2025-12-15 Motus: A Unified Latent Action World Model Hongzhe Bi et.al. 2512.13030 null
2025-12-15 Safe Control of Multi-Agent Systems with Minimal Communication Mo Yang et.al. 2512.13021 null
2025-12-15 Quantigence: A Multi-Agent AI Framework for Quantum Security Research Abdulmalik Alquwayfili et.al. 2512.12989 null
2025-12-15 QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management Weizhou Shen et.al. 2512.12967 null
2025-12-15 Building from Scratch: A Multi-Agent Framework with Human-in-the-Loop for Multilingual Legal Terminology Mapping Lingyi Meng et.al. 2512.12950 null
2025-12-14 Fault-Tolerant Sandboxing for AI Coding Agents: A Transactional Approach to Safe Autonomous Execution Boyang Yan et.al. 2512.12806 null
2025-12-14 Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems Sreemaee Akshathala et.al. 2512.12791 null
2025-12-14 NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents Jingzhe Ding et.al. 2512.12730 null
2025-12-14 Towards AI Agents Supported Research Problem Formulation Anrafel Fernandes Pereira et.al. 2512.12719 null
2025-12-12 Evaluating Cooperative Resilience in Multiagent Systems: A Comparison Between Humans and LLMs Manuela Chacon-Chamorro et.al. 2512.11689 null
2025-12-12 MedAI: Evaluating TxAgent’s Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition Tim Cofala et.al. 2512.11682 null
2025-12-12 Risk Limited Asset Allocation with a Budget Threshold Utility Function and Leptokurtotic Distributions of Returns Graham L Giller et.al. 2512.11666 null
2025-12-12 Basis dependence of Neural Quantum States for the Transverse Field Ising Model Ronald Santiago Cortes et.al. 2512.11632 null
2025-12-12 Embodied Image Compression Chunyi Li et.al. 2512.11612 null
2025-12-12 A Study of Library Usage in Agent-Authored Pull Requests Lukas Twist et.al. 2512.11589 null
2025-12-12 Gradient Descent as a Perceptron Algorithm: Understanding Dynamics and Implicit Acceleration Alexander Tyurin et.al. 2512.11587 null
2025-12-12 EmeraldMind: A Knowledge Graph-Augmented Framework for Greenwashing Detection Georgios Kaoukis et.al. 2512.11506 null
2025-12-12 AgentBalance: Backbone-then-Topology Design for Cost-Effective Multi-Agent Systems under Budget Constraints Shuowei Cai et.al. 2512.11426 null
2025-12-12 Towards Trustworthy Multi-Turn LLM Agents via Behavioral Guidance Gonca Gürsun et.al. 2512.11421 null
2025-12-12 AutoFSM: A Multi-agent Framework for FSM Code Generation with IR and SystemC-Based Testing Qiuming Luo et.al. 2512.11398 null
2025-12-12 Benchmarking the Generality of Vision-Language-Action Models Pranav Guruprasad et.al. 2512.11315 null
2025-12-12 When Actions Teach You to Think: Reasoning-Action Synergy via Reinforcement Learning in Conversational Agents Mrinal Rawat et.al. 2512.11277 null
2025-12-12 Words to Describe What I’m Feeling: Exploring the Potential of AI Agents for High Subjectivity Decisions in Advance Care Planning Kellie Yu Hui Sim et.al. 2512.11276 null
2025-12-12 TriFlow: A Progressive Multi-Agent Framework for Intelligent Trip Planning Yuxing Chen et.al. 2512.11271 null
2025-12-12 Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language Yunkai Zhang et.al. 2512.11251 null
2025-12-12 FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration Dongwon Jung et.al. 2512.11213 null
2025-12-11 Automated Penetration Testing with LLM Agents and Classical Planning Lingzhi Wang et.al. 2512.11143 null
2025-12-11 Asynchronous Reasoning: Training-Free Interactive Thinking LLMs George Yakushev et.al. 2512.10931 null
2025-12-11 CompanionCast: A Multi-Agent Conversational AI Framework with Spatial Audio for Social Co-Viewing Experiences Yiyang Wang et.al. 2512.10918 null
2025-12-11 Opportunities and Challenges in Harnessing Digital Technology for Effective Teaching and Learning Zhongzhou Chen et.al. 2512.10777 null
2025-12-11 Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution Zouying Cao et.al. 2512.10696 null
2025-12-11 On the ground state of the nonlinear Schr{ö}dinger equation: asymptotic behavior at the endpoint powers Rémi Carles et.al. 2512.10690 null
2025-12-11 LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator Lihuang Chen et.al. 2512.10605 null
2025-12-11 Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning Haiteng Zhao et.al. 2512.10534 null
2025-12-11 Zero-shot 3D Map Generation with LLM Agents: A Dual-Agent Architecture for Procedural Content Generation Lim Chien Her et.al. 2512.10501 null
2025-12-11 Confucius Code Agent: An Open-sourced AI Software Engineer at Industrial Scale Zhaodong Wang et.al. 2512.10398 null
2025-12-11 Cross-modal Retrieval Models for Stripped Binary Analysis Guoqiang Chen et.al. 2512.10393 null
2025-12-11 EpiPlanAgent: Agentic Automated Epidemic Response Planning Kangkun Mao et.al. 2512.10313 null
2025-12-11 CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment Yakun Zhu et.al. 2512.10206 null
2025-12-11 AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding Gyutaek Oh et.al. 2512.10195 null
2025-12-11 SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs Arihant Tripathy et.al. 2512.09543 null
2025-12-10 Workflow is All You Need: Escaping the “Statistical Smoothing Trap” via High-Entropy Information Foraging and Adversarial Pacing Zhongjie Jiang et.al. 2512.10121 null
2025-12-10 Detailed balance in large language model-driven agents Zhuo-Yang Song et.al. 2512.10047 null
2025-12-10 DynaMate: An Autonomous Agent for Protein-Ligand Molecular Dynamics Simulations Salomé Guilbert et.al. 2512.10034 null
2025-12-10 VisualActBench: Can VLMs See and Act like a Human? Daoan Zhang et.al. 2512.09907 null
2025-12-10 Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing Justin W. Lin et.al. 2512.09882 null
2025-12-10 UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories Yanghong Mei et.al. 2512.09607 null
2025-12-10 Architectures for Building Agentic AI Sławomir Nowaczyk et.al. 2512.09458 null
2025-12-10 GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection Zishu Wei et.al. 2512.09396 null
2025-12-10 Optimizing Data Extraction from Materials Science Literature: A Study of Tools Using Large Language Models Wenkai Ning et.al. 2512.09370 null
2025-12-10 ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data Ruiqi Wang et.al. 2512.09321 null
2025-12-10 Scene-agnostic Hierarchical Bimanual Task Planning via Visual Affordance Reasoning Kwang Bin Lee et.al. 2512.09310 null
2025-12-10 The Illusion of Rationality: Tacit Bias and Strategic Dominance in Frontier LLM Negotiation Games Manuel S. Ríos et.al. 2512.09254 null
2025-12-09 WOLF: Werewolf-based Observations for LLM Deception and Falsehoods Mrinal Agarwal et.al. 2512.09187 null
2025-12-09 Evolving Excellence: Automated Optimization of LLM-based Agents Paul Brookes et.al. 2512.09108 null
2025-12-09 Mental Models of Autonomy and Sentience Shape Reactions to AI Janet V. T. Pauketat et.al. 2512.09085 null
2025-12-09 AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models Arman Zarei et.al. 2512.09081 null
2025-12-09 Fed-SE: Federated Self-Evolution for Privacy-Constrained Multi-Environment LLM Agents Xiang Chen et.al. 2512.08870 null
2025-12-09 A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows Eranga Bandara et.al. 2512.08769 null
2025-12-09 Difference-in-Differences with Interval Data Daisuke Kurisu et.al. 2512.08759 null
2025-12-09 Towards Foundation Models with Native Multi-Agent Intelligence Shuyue Hu et.al. 2512.08743 null
2025-12-09 Insured Agents: A Decentralized Trust Insurance Mechanism for Agentic Economy Botao ‘Amber’ Hu et.al. 2512.08737 null
2025-12-09 Multi-Agent Intelligence for Multidisciplinary Decision-Making in Gastrointestinal Oncology Rongzhao Zhang et.al. 2512.08674 null
2025-12-09 See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm Haoyu Zhao et.al. 2512.08629 null
2025-12-09 Autonomous Issue Resolver: Towards Zero-Touch Code Maintenance Aliaksei Kaliutau et.al. 2512.08492 null
2025-12-09 NeurIDA: Dynamic Modeling for Effective In-Database Analytics Lingze Zeng et.al. 2512.08483 null
2025-12-09 A Multi-Agent LLM Framework for Design Space Exploration in Autonomous Driving Systems Po-An Shih et.al. 2512.08476 null
2025-12-09 Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs Yinan Zhong et.al. 2512.08417 null
2025-12-09 Reflecting with Two Voices: A Co-Adaptive Dual-Strategy Framework for LLM-Based Agent Decision Making Wentao Zhang et.al. 2512.08366 null
2025-12-09 Argus: A Multi-Agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference Relationships Bin Wang et.al. 2512.08326 null
2025-12-09 rSIM: Incentivizing Reasoning Capabilities of LLMs via Reinforced Strategy Injection Sijia Chen et.al. 2512.08300 null
2025-12-09 Systematization of Knowledge: Security and Safety in the Model Context Protocol Ecosystem Shiva Gaire et.al. 2512.08290 null
2025-12-09 Empowering smart app development with SolidGPT: an edge-cloud hybrid AI agent framework Liao Hu et.al. 2512.08286 null
2025-12-09 Chat with UAV – Human-UAV Interaction Based on Large Language Models Haoran Wang et.al. 2512.08145 null
2025-12-09 Robust Agents in Open-Ended Worlds Mikayel Samvelyan et.al. 2512.08139 null
2025-12-08 AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration Harish Karthikeyan et.al. 2512.08104 null
2025-12-08 Adaptation of Embedding Models to Financial Filings via LLM Distillation Eliot Brenner et.al. 2512.08088 null
2025-12-08 The Adoption and Usage of AI Agents: Early Evidence from Perplexity Jeremy Yang et.al. 2512.07828 null
2025-12-08 Automating High Energy Physics Data Analysis with LLM-Powered Agents Eli Gendreau-Distler et.al. 2512.07785 null
2025-12-08 Reliable agent engineering should integrate machine-compatible organizational principles R. Patrick Xian et.al. 2512.07665 null
2025-12-08 The Agent Capability Problem: Predicting Solvability Through Information-Theoretic Bounds Shahar Lutati et.al. 2512.07631 null
2025-12-08 Online Segment Any 3D Thing as Instance Tracking Hanshi Wang et.al. 2512.07599 null
2025-12-08 VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection Yuzhou Nie et.al. 2512.07533 null
2025-12-08 How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations JV Roig et.al. 2512.07497 null
2025-12-08 Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization Zhuoran Zhuang et.al. 2512.07478 null
2025-12-08 Understanding LLM Agent Behaviours via Game Theory: Strategy Recognition, Biases and Multi-Agent Dynamics Trung-Kiet Huynh et.al. 2512.07462 null
2025-12-08 Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Tong Wu et.al. 2512.07461 null
2025-12-08 Adaptive Tuning of Parameterized Traffic Controllers via Multi-Agent Reinforcement Learning Giray Önür et.al. 2512.07417 null
2025-12-08 Training Language Models to Use Prolog as a Tool Niklas Mellgren et.al. 2512.07407 null
2025-12-08 DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection Sayeem Been Zaman et.al. 2512.07351 null
2025-12-08 Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts Mingning Guo et.al. 2512.07302 null
2025-12-08 SIT-Graph: State Integrated Tool Graph for Multi-Turn Agents Sijia Li et.al. 2512.07287 null
2025-12-08 DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning Nithin Sivakumaran et.al. 2512.07132 null
2025-12-08 Human Agency and Creativity in AI-Assisted Learning Environments Yun Dai et.al. 2512.07117 null
2025-12-08 VIGIL: A Reflective Runtime for Self-Healing Agents Christopher Cruz et.al. 2512.07094 null
2025-12-08 ClinNoteAgents: An LLM Multi-Agent System for Predicting and Interpreting Heart Failure 30-Day Readmission from Clinical Notes Rongjia Zhou et.al. 2512.07081 null
2025-12-07 MATEX: A Multi-Agent Framework for Explaining Ethereum Transactions Zifan Peng et.al. 2512.06933 null
2025-12-05 Strongly Coupled Quantum Forces Yuval Grossman et.al. 2512.05968 null
2025-12-05 Categorifying isomonodromic deformations via Lie groupoids I: Logarithmic singularities Waleed Qaisar et.al. 2512.05966 null
2025-12-05 Trusted AI Agents in the Cloud Teofil Bodea et.al. 2512.05951 null
2025-12-05 Removing correlated noise stripes from the Nancy Grace Roman Space Telescope survey images Katherine Laliotis et.al. 2512.05949 null
2025-12-05 Developing synthetic microdata through machine learning for firm-level business surveys Jorge Cisneros Paz et.al. 2512.05948 null
2025-12-05 Designing an Optimal Sensor Network via Minimizing Information Loss Daniel Waxman et.al. 2512.05940 null
2025-12-05 PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation Shima Imani et.al. 2512.05930 null
2025-12-05 NICE: Neural Implicit Craniofacial Model for Orthognathic Surgery Prediction Jiawen Yang et.al. 2512.05920 null
2025-12-05 Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures Amirkia Rafiei Oskooei et.al. 2512.05908 null
2025-12-05 From Text to Returns: Using Large Language Models for Mutual Fund Portfolio Optimization and Risk-Adjusted Allocation Abrar Hossain Mufakir Qamar Ansari Haziq Jeelani Monia Digra Fayeq Jeelani Syed et.al. 2512.05907 null
2025-12-05 Computer simulations of the Stark effect in the helium-beta complex of krypton in ICF conditions G. Pérez-Callejo et.al. 2512.05903 null
2025-12-05 Euclid Quick Data Release (Q1). From simulations to sky: Advancing machine-learning lens detection with real Euclid data Euclid Collaboration et.al. 2512.05899 null
2025-12-05 Bootstrapping Fuzzers for Compilers of Low-Resource Language Dialects Using Language Models Sairam Vaidya et.al. 2512.05887 null
2025-12-05 A Continuous Nonlinear Optimization Perspective on the Spin Glass Problem Phil Duxbury et.al. 2512.05852 null
2025-12-05 Invariant Price of Anarchy: a Metric for Welfarist Traffic Control Ilia Shilov et.al. 2512.05843 null
2025-12-05 Multimodal Oncology Agent for IDH1 Mutation Prediction in Low-Grade Glioma Hafsa Akebli et.al. 2512.05824 null
2025-12-05 Optimal Safety-Aware Scheduling for Multi-Agent Aerial 3D Printing with Utility Maximization under Dependency Constraints Marios-Nektarios Stamatopoulos et.al. 2512.05815 null
2025-12-05 Floer sections in multisymplectic geometry Ronen Brilleslijper et.al. 2512.05797 null
2025-12-05 Task-Specific Trust Evaluation for Multi-Hop Collaborator Selection via GNN-Aided Distributed Agentic AI Botao Zhu et.al. 2512.05788 null
2025-12-05 Machine Learning-Informed 3+1 Sterile Neutrino Global Fits using Posterior Density Estimation of Electron Disappearance Data Joshua Villarreal et.al. 2512.05784 null
2025-12-05 GRASP: Graph Reasoning Agents for Systems Pharmacology with Human-in-the-Loop Omid Bazgir et.al. 2512.05502 null
2025-12-05 Model Gateway: Model Management Platform for Model-Driven Drug Discovery Yan-Shiun Wu et.al. 2512.05462 null
2025-12-05 Please Don’t Kill My Vibe: Empowering Agents with Data Flow Control Charlie Summers et.al. 2512.05374 null
2025-12-05 MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare Zag ElSayed et.al. 2512.05365 null
2025-12-04 WhatsCode: Large-Scale GenAI Deployment for Developer Efficiency at WhatsApp Ke Mao et.al. 2512.05314 null
2025-12-04 Beyond Detection: A Comprehensive Benchmark and Study on Representation Learning for Fine-Grained Webshell Family Classification Feijiang Han et.al. 2512.05288 null
2025-12-04 Deep infant brain segmentation from multi-contrast MRI Malte Hoffmann et.al. 2512.05114 null
2025-12-04 ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning Shengyuan Ding et.al. 2512.05111 null
2025-12-04 Breaking the bandwidth-efficiency trade-off in soliton microcombs via mode coupling Yang Liu et.al. 2512.05090 null
2025-12-04 Gradient Descent with Provably Tuned Learning-rate Schedules Dravyansh Sharma et.al. 2512.05084 null
2025-12-04 David vs. Goliath: Can Small Models Win Big with Agentic AI in Hardware Design? Shashwat Shankar et.al. 2512.05073 null
2025-12-04 Personalizing Agent Privacy Decisions via Logical Entailment James Flemings et.al. 2512.05065 null
2025-12-04 Configuration Defects in Kubernetes Yue Zhang et.al. 2512.05062 null
2025-12-04 Detecting Perspective Shifts in Multi-agent Systems Eric Bridgeford et.al. 2512.05013 null
2025-12-04 Evolutionary Architecture Search through Grammar-Based Sequence Alignment Adri Gómez Martín et.al. 2512.04992 null
2025-12-04 Introduction to quantum control: From basic concepts to applications in quantum technologies Christiane P. Koch et.al. 2512.04990 null
2025-12-04 Strategic Self-Improvement for Competitive Agents in AI Labour Markets Christopher Chiu et.al. 2512.04988 null
2025-12-04 Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction Nex-AGI Team et.al. 2512.04987 null
2025-12-04 A tangential low-rank ADI method for solving indefinite Lyapunov equations Rudi Smith et.al. 2512.04983 null
2025-12-04 Multi-Agent Reinforcement Learning for Intraday Operating Rooms Scheduling under Uncertainty Kailiang Liu et.al. 2512.04918 null
2025-12-04 On Disturbance-Aware Minimum-Time Trajectory Planning: Evidence from Tests on a Dynamic Driving Simulator Matteo Masoni et.al. 2512.04917 null
2025-12-04 Distributed Riemannian Optimization in Geodesically Non-convex Environments Xiuheng Wang et.al. 2512.04915 null
2025-12-04 Analytical and Cross-Sectional Clinical Validity of a Smartphone-Based U-Turn Test in Multiple Sclerosis Marta Płonka et.al. 2512.04914 null
2025-12-04 Declarative Synthesis and Multi-Objective Optimization of Stripboard Circuit Layouts Using Answer Set Programming Fang Li et.al. 2512.04910 null
2025-12-04 Chameleon: Adaptive Adversarial Agents for Scaling-Based Visual Prompt Injection in Multimodal AI Systems M Zeeshan et.al. 2512.04895 null
2025-12-04 Data-driven Methods for Delay Differential Equations Dimitri Breda et.al. 2512.04894 null
2025-12-04 Enabling Ethical AI: A case study in using Ontological Context for Justified Agentic AI Decisions Liam McGee et.al. 2512.04822 null
2025-12-04 SIMA 2: A Generalist Embodied Agent for Virtual Worlds SIMA team et.al. 2512.04797 null
2025-12-04 ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications Eranga Bandara et.al. 2512.04785 null
2025-12-04 POLARIS: Is Multi-Agentic Reasoning the Next Wave in Engineering Self-Adaptive Systems? Divyansh Pandey et.al. 2512.04702 null
2025-12-04 Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective Jae Hee Lee et.al. 2512.04691 null
2025-12-04 Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space Joey Hong et.al. 2512.04601 null
2025-12-04 When Robots Should Say “I Don’t Know”: Benchmarking Abstention in Embodied Question Answering Tao Wu et.al. 2512.04597 null
2025-12-03 Permutation Flows I: Triangulations of Flow Polytopes (Research Announcement) Rafael S. González D’León et.al. 2512.04078 null
2025-12-03 Semi-Markov Decision Process Framework for Age of Incorrect Information Minimization Ismail Cosandal et.al. 2512.04077 null
2025-12-03 SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL Siyi Chen et.al. 2512.04069 null
2025-12-03 Learning Steerable Clarification Policies with Collaborative Self-play Jonathan Berant et.al. 2512.04068 null
2025-12-03 Affordances of Digital and Blockchain-based Community Currencies: The Case of Sarafu Network in Kenya Patricia Marcella Evite et.al. 2512.04030 null
2025-12-03 Thermalization from quenching in coupled oscillators M. Harinarayanan et.al. 2512.04028 null
2025-12-03 Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models Taido Purason et.al. 2512.03989 null
2025-12-03 Benchmark for Planning and Control with Large Language Model Agents: Blocksworld with Model Context Protocol Niklas Jobs et.al. 2512.03955 null
2025-12-03 Performance and efficiency of a transformer-based quark/gluon jet tagger in the ATLAS experiment ATLAS Collaboration et.al. 2512.03949 null
2025-12-03 Classification of User Satisfaction in HRI with Social Signals in the Wild Michael Schiffmann et.al. 2512.03945 null
2025-12-03 Functorial properties of Schwinger-DeWitt expansion and Mellin-Barnes representation Andrei O. Barvinsky et.al. 2512.03944 null
2025-12-03 DSP: A Statistically-Principled Structural Polarization Measure Giulia Preti et.al. 2512.03937 null
2025-12-03 Driving is a Game: Combining Planning and Prediction with Bayesian Iterative Best Response Aron Distelzweig et.al. 2512.03936 null
2025-12-03 Autonomous Agents and Policy Compliance: A Framework for Reasoning About Penalties Vineel Tummala et.al. 2512.03931 null
2025-12-03 Integrating High Performance In-Memory Data Streaming and In-Situ Visualization in Hybrid MPI+OpenMP PIC MC Simulations Towards Exascale Jeremy J. Williams et.al. 2512.03914 null
2025-12-03 Hierarchical Vision Language Action Model Using Success and Failure Demonstrations Jeongeun Park et.al. 2512.03913 null
2025-12-03 Probabilistic Foundations of Fuzzy Simplicial Sets for Nonlinear Dimensionality Reduction Janis Keck et.al. 2512.03899 null
2025-12-03 Shadow geometry of Kerr MOG naked singularity and analysis of accretion disk luminosity Saira Yasmin et.al. 2512.03896 null
2025-12-03 A Hierarchical Tree-based approach for creating Configurable and Static Deep Research Agent (Static-DRA) Saurav Prateek et.al. 2512.03887 null
2025-12-03 Adhera: A Human-Centered Health Informatics Solution for Reducing Informal Caregiver Burden through Improved Medication Adherence Zhiyin Zhou et.al. 2512.03878 null
2025-12-02 PPTArena: A Benchmark for Agentic PowerPoint Editing Michael Ofengenden et.al. 2512.03042 null
2025-12-02 Generation of strong ultralow-phase-noise microwave fields with tunable ellipticity for ultracold polar molecules Shrestha Biswas et.al. 2512.03007 null
2025-12-02 From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars? Dawei Li et.al. 2512.03005 null
2025-12-02 DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling Kairun Wen et.al. 2512.03000 null
2025-12-02 InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration Zhongyu Yang et.al. 2512.02981 null
2025-12-02 Many-body $k$ -local ground states as probes for unitary quantum metrology Majid Hassani et.al. 2512.02976 null
2025-12-02 The Evolutionary Ecology of Software: Constraints, Innovation, and the AI Disruption Sergi Valverde et.al. 2512.02953 null
2025-12-02 Fast Gaussian Process Approximations for Autocorrelated Data Ahmadreza Chokhachian et.al. 2512.02925 null
2025-12-02 Distinguishing ram pressure from gravitational interactions: Applying the Size-Shape Difference method to real galaxies Augusto E. Lassen et.al. 2512.02923 null
2025-12-02 On the distribution of very short character sums Paweł Nosal et.al. 2512.02915 null
2025-12-02 Hypothesis Testing for Generalized Thurstone Models Anuran Makur et.al. 2512.02912 null
2025-12-02 FluxLab: Creating 3D Printable Shape-Changing Devices with Integrated Deformation Sensing Hsuanling Lee et.al. 2512.02911 null
2025-12-02 SAT-MapIt: A SAT-based Modulo Scheduling Mapper for Coarse Grain Reconfigurable Architectures Cristian Tirelli et.al. 2512.02875 null
2025-12-02 Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning Haonan Wang et.al. 2512.02874 null
2025-12-02 Limiting Reduction and Modified Gravity Antonis Antoniou et.al. 2512.02871 null
2025-12-02 Network Self-Configuration based on Fine-Tuned Small Language Models Oscar G. Lira et.al. 2512.02861 null
2025-12-02 Radiologist Copilot: An Agentic Assistant with Orchestrated Tools for Radiology Reporting with Quality Control Yongrui Yu et.al. 2512.02814 null
2025-12-02 Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents Zijie Lin et.al. 2512.02812 null
2025-12-02 Phase-Adaptive LLM Framework with Multi-Stage Validation for Construction Robot Task Allocation: A Systematic Benchmark Against Traditional Optimization Algorithms Shyam prasad reddy Kaitha et.al. 2512.02810 null
2025-12-02 Perception of AI-Generated Music – The Role of Composer Identity, Personality Traits, Music Preferences, and Perceived Humanness David Stammer et.al. 2512.02785 null
2025-12-01 LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess Sai Kolasani et.al. 2512.01992 null
2025-12-01 Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships Hannah Rose Kirk et.al. 2512.01991 null
2025-12-01 Forecasting in Offline Reinforcement Learning for Non-stationary Environments Suzan Ece Ada et.al. 2512.01987 null
2025-12-01 Fault-tolerant mutual-visibility: complexity and solutions for grid-like networks Serafino Cicerone et.al. 2512.01978 null
2025-12-01 How Far Are We from Genuinely Useful Deep Research Agents? Dingling Zhang et.al. 2512.01948 null
2025-12-01 Agentic Policy Optimization via Instruction-Policy Co-Evolution Han Zhou et.al. 2512.01945 null
2025-12-01 An Empirical Study of Agent Developer Practices in AI Agent Frameworks Yanlin Wang et.al. 2512.01939 null
2025-12-01 Parametric processes in nonlinear structures with reflections: a transfer matrix method approach Salvador Poveda-Hospital et.al. 2512.01921 null
2025-12-01 Latent Debate: A Surrogate Framework for Interpreting LLM Thinking Lihu Chen et.al. 2512.01909 null
2025-12-01 NeuroHJR: Hamilton-Jacobi Reachability-based Obstacle Avoidance in Complex Environments with Physics-Informed Neural Networks Granthik Halder et.al. 2512.01897 null
2025-12-01 OPOR-Bench: Evaluating Large Language Models on Online Public Opinion Report Generation Jinzheng Yu et.al. 2512.01896 null
2025-12-01 Predicting Human Chess Moves: An AI Assisted Analysis of Chess Games Using Skill-group Specific n-gram Language Models Daren Zhong et.al. 2512.01880 null
2025-12-01 Graph Distance as Surprise: Free Energy Minimization in Knowledge Graph Reasoning Gaganpreet Jhajj et.al. 2512.01878 null
2025-12-01 Model theory, differential algebra and functional transcendence Amador Martin-Pizarro et.al. 2512.01866 null
2025-12-01 COACH: Collaborative Agents for Contextual Highlighting - A Multi-Agent Framework for Sports Video Analysis Tsz-To Wong et.al. 2512.01853 null
2025-12-01 Non-archimedean Infinite Hecke Algebra Milo Bechtloff Weising et.al. 2512.01835 null
2025-12-01 The partial K function Jake P. Grainger et.al. 2512.01823 null
2025-12-01 InnoGym: Benchmarking the Innovation Potential of AI Agents Jintian Zhang et.al. 2512.01822 null
2025-12-01 Dimension-free error estimate for diffusion model and optimal scheduling Valentin de Bortoli et.al. 2512.01820 null
2025-12-01 DeepCAVE: A Visualization and Analysis Tool for Automated Machine Learning Sarah Segel et.al. 2512.01810 null
2025-11-28 Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction Bao Shu et.al. 2511.23476 null
2025-11-28 Arbitrary control of the temporal waveform of photons during spontaneous emission Carl Thomas et.al. 2511.23462 null
2025-11-28 Designing Rules for Choosing a Winner in a Debate Alexander Heckett et.al. 2511.23454 null
2025-11-28 Random purification channel made simple Filippo Girardi et.al. 2511.23451 null
2025-11-28 ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts Hang Yu et.al. 2511.23442 null
2025-11-28 Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent Jianzhe Lin et.al. 2511.23436 null
2025-11-28 Consensus Tree Estimation with False Discovery Rate Control via Partially Ordered Sets Maria Alejandra Valdez Cabrera et.al. 2511.23433 null
2025-11-28 MegaChat: A Synthetic Persian Q&A Dataset for High-Quality Sales Chatbot Evaluation Mahdi Rahmani et.al. 2511.23397 null
2025-11-28 Bounded-Error Quantum Simulation via Hamiltonian and Lindbladian Learning Tristan Kraft et.al. 2511.23392 null
2025-11-28 Hierarchical AI-Meteorologist: LLM-Agent System for Multi-Scale and Explainable Weather Forecast Reporting Daniil Sukhorukov et.al. 2511.23387 null
2025-11-28 Identifying bars in galaxies using machine learning Rajit Shrivastava et.al. 2511.23383 null
2025-11-28 AugGen: Augmenting Task-Based Learning in Professional Creative Software with LLM-Generated Scaffolded UIs Yimeng Liu et.al. 2511.23379 null
2025-11-28 Agentic AI Framework for Smart Inventory Replenishment Toqeer Ali Syed et.al. 2511.23366 null
2025-11-28 Functional Program Synthesis with Higher-Order Functions and Recursion Schemes Matheus Campos Fernandes et.al. 2511.23354 null
2025-11-28 Distributed Dynamic Associative Memory via Online Convex Optimization Bowen Wang et.al. 2511.23347 null
2025-11-28 ParaGate: Parasitic-Driven Domain Adaptation Transfer Learning for Netlist Performance Prediction Bin Sun et.al. 2511.23340 null
2025-11-28 Optimization and application of ultra-high field preclinical high-resolution and 3D 1H-MRSI using compressed sensing Brayan Alves et.al. 2511.23331 null
2025-11-28 MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report) Aaron Steiner et.al. 2511.23281 null
2025-11-28 Beyond Curve Fitting: Neuro-Symbolic Agents for Context-Aware Epidemic Forecasting Joongwon Chae et.al. 2511.23276 null
2025-11-28 Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs Jiancheng Dong et.al. 2511.23271 null
2025-11-26 ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration Hongjin Su et.al. 2511.21689 null
2025-11-26 Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework Dong Wang et.al. 2511.21686 null
2025-11-26 Agentic Learner with Grow-and-Refine Multimodal Semantic Memory Weihao Bo et.al. 2511.21678 null
2025-11-26 AI/ML Model Cards in Edge AI Cyberinfrastructure: towards Agentic AI Beth Plale et.al. 2511.21661 null
2025-11-26 The Need for Benchmarks to Advance AI-Enabled Player Risk Detection in Gambling Kasra Ghaharian et.al. 2511.21658 null
2025-11-26 EvilGenie: A Reward Hacking Benchmark Jonathan Gabor et.al. 2511.21654 null
2025-11-26 Stochastic Optimal Control of Interacting Particle Systems in Hilbert Spaces and Applications Filippo de Feo et.al. 2511.21646 null
2025-11-26 Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO Daniel R. Jiang et.al. 2511.21638 null
2025-11-26 Qwen3-VL Technical Report Shuai Bai et.al. 2511.21631 null
2025-11-26 Uniform inference for kernel instrumental variable regression Marvin Lob et.al. 2511.21603 null
2025-11-26 On the Limits of Innate Planning in Large Language Models Charles Schepanowski et.al. 2511.21591 null
2025-11-26 Approximate Bayesian Computation Made Easy: A Practical Guide to ABC-SMC for Dynamical Systems with \texttt{pymc} Mario Castro et.al. 2511.21587 null
2025-11-26 Model-Based Policy Adaptation for Closed-Loop End-to-End Autonomous Driving Haohong Lin et.al. 2511.21584 null
2025-11-26 BAMAS: Structuring Budget-Aware Multi-Agent Systems Liming Yang et.al. 2511.21572 null
2025-11-26 VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation Hui Zhou et.al. 2511.21557 null
2025-11-26 MAD-DAG: Protecting Blockchain Consensus from MEV Roi Bar-Zur et.al. 2511.21552 null
2025-11-26 Seeing Twice: How Side-by-Side T2I Comparison Changes Auditing Strategies Matheus Kunzler Maldaner et.al. 2511.21547 null
2025-11-26 Hidden symmetries and separability structures of Ovcharenko-Podolský and conformal-to-Carter spacetimes Finnian Gray et.al. 2511.21538 null
2025-11-26 Integrated emitters with CMOS-compatible tuning for large scale quantum SiN photonic circuits Jasper De Witte et.al. 2511.21529 null
2025-11-26 The Quantum Network of Assets: A Non-Classical Framework for Market Correlation and Structural Risk Hui Gong et.al. 2511.21515 null
2025-11-25 Latent Collaboration in Multi-Agent Systems Jiaru Zou et.al. 2511.20639 null
2025-11-25 Quantum-Resistant Authentication Scheme for RFID Systems Using Lattice-Based Cryptography Vaibhav Kumar et.al. 2511.20630 null
2025-11-25 The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment Ziheng Ouyang et.al. 2511.20614 null
2025-11-25 Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning Panayiotis Danassis et.al. 2511.20613 null
2025-11-25 Optimization of Sums of Bivariate Functions: An Introduction to Relaxation-Based Methods for the Case of Finite Domains Nils Müller et.al. 2511.20607 null
2025-11-25 Limit Order Book Dynamics in Matching Markets:Microstructure, Spread, and Execution Slippage Yao Wu et.al. 2511.20606 null
2025-11-25 BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents Kaiyuan Zhang et.al. 2511.20597 null
2025-11-25 Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning Charlotte Beylier et.al. 2511.20591 null
2025-11-25 EnergyTwin: A Multi-Agent System for Simulating and Coordinating Energy Microgrids Jakub Muszyński et.al. 2511.20590 null
2025-11-25 VQ-VA World: Towards High-Quality Visual Question-Visual Answering Chenhui Gou et.al. 2511.20573 null
2025-11-25 E2E-GRec: An End-to-End Joint Training Framework for Graph Neural Networks and Recommender Systems Rui Xue et.al. 2511.20564 null
2025-11-25 Effective Command-line Interface Fuzzing with Path-Aware Large Language Model Orchestration Momoko Shiraishi et.al. 2511.20555 null
2025-11-25 Proceedings Twentieth Conference on Theoretical Aspects of Rationality and Knowledge Adam Bjorndahl et.al. 2511.20540 null
2025-11-25 Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models Shamima Hossain et.al. 2511.20531 null
2025-11-25 Efficient Parallel Implementation of the Pilot Assignment Problem in Massive MIMO Systems Eman Alqudah et.al. 2511.20511 null
2025-11-25 FRAGMENTA: End-to-end Fragmentation-based Generative Model with Agentic Tuning for Drug Lead Optimization Yuto Suzuki et.al. 2511.20510 null
2025-11-25 A Single-Root, Multi-Curve, Context-Isolated, PQC-Pluggable Cryptographic Identity Primitive with Stateless Secret Rotation Jian Sheng Wang et.al. 2511.20505 null
2025-11-25 MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology Kiril Vasilev et.al. 2511.20490 null
2025-11-25 MLIPAudit: A benchmarking tool for Machine Learned Interatomic Potentials Leon Wehrhan et.al. 2511.20487 null
2025-11-25 Efficient and Fast Generative-Based Singing Voice Separation using a Latent Diffusion Model Genís Plaja-Roglans et.al. 2511.20470 null
2025-11-24 VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection Qiang Wang et.al. 2511.19436 null
2025-11-24 Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution Dingkang Liang et.al. 2511.19430 null
2025-11-24 Beyond Protein Language Models: An Agentic LLM Framework for Mechanistic Enzyme Design Bruno Jacob et.al. 2511.19423 null
2025-11-24 Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration James Y. Huang et.al. 2511.19417 null
2025-11-24 Learning Robust Social Strategies with Large Language Models Dereck Piche et.al. 2511.19405 null
2025-11-24 Frequency-Invariant Beamforming in Elevation and Azimuth via Autograd and Concentric Circular Microphone Arrays Jorge Ortigoso-Narro et.al. 2511.19403 null
2025-11-24 DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research Rulin Shao et.al. 2511.19399 null
2025-11-24 BackSplit: The Importance of Sub-dividing the Background in Biomedical Lesion Segmentation Rachit Saluja et.al. 2511.19394 null
2025-11-24 Asymptotic linear dependence and ellipse statistics for multivariate two-sample homogeneity test Chifeng Shen et.al. 2511.19381 null
2025-11-24 Product Depth for Temporal Point Processes Observed Only Up to the First k Events Chifeng Shen et.al. 2511.19375 null
2025-11-24 LLM-Driven Stationarity-Aware Expert Demonstrations for Multi-Agent Reinforcement Learning in Mobile Systems Tianyang Duan et.al. 2511.19368 null
2025-11-24 Enhancing Conformal Prediction via Class Similarity Ariel Fargion et.al. 2511.19359 null
2025-11-24 Explicit Tonal Tension Conditioning via Dual-Level Beam Search for Symbolic Music Generation Maral Ebrahimzadeh et.al. 2511.19342 null
2025-11-24 Normative active inference: A numerical proof of principle for a computational and economic legal analytic approach to AI governance Axel Constant et.al. 2511.19334 null
2025-11-24 Revisiting model-independent constraints on spatial curvature and cosmic ladders calibration: updated and forecast analyses Arianna Favale et.al. 2511.19332 null
2025-11-24 Dynamic Leader-Follower Consensus with Adversaries: A Multi-Hop Relay Approach Liwei Yuan et.al. 2511.19327 null
2025-11-24 PRInTS: Reward Modeling for Long-Horizon Information Seeking Jaewoo Lee et.al. 2511.19314 null
2025-11-24 The TEQUILA catalog of variables in TESS full-frame images: Differential photometry light curves from the first two years of observations Bisi Bernard Ogunwale et.al. 2511.19313 null
2025-11-24 AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning Jiayi Zhang et.al. 2511.19304 null
2025-11-24 Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks Bianka Kowalska et.al. 2511.19265 null
2025-11-21 MDG: Masked Denoising Generation for Multi-Agent Behavior Modeling in Traffic Environments Zhiyu Huang et.al. 2511.17496 null
2025-11-21 Harnessing Data from Clustered LQR Systems: Personalized and Collaborative Policy Optimization Vinay Kanakeri et.al. 2511.17489 null
2025-11-21 Counterfactual World Models via Digital Twin-conditioned Video Diffusion Yiqing Shen et.al. 2511.17481 null
2025-11-21 PersonaAgent with GraphRAG: Community-Aware Knowledge Graphs for Personalized LLM Siqi Liang et.al. 2511.17467 null
2025-11-21 SRA-CP: Spontaneous Risk-Aware Selective Cooperative Perception Jiaxi Liu et.al. 2511.17461 null
2025-11-21 Minimalist machine-learned interatomic potentials can predict complex structural behaviors accurately Iñigo Robredo-Magro et.al. 2511.17449 null
2025-11-21 GRAPHIC–Guidelines for Reviewing Algorithmic Practices in Human-centred Design and Interaction for Creativity Joana Rovira Martins et.al. 2511.17443 null
2025-11-21 REMSA: An LLM Agent for Foundation Model Selection in Remote Sensing Binger Chen et.al. 2511.17442 null
2025-11-21 Multi-Agent Pointer Transformer: Seq-to-Seq Reinforcement Learning for Multi-Vehicle Dynamic Pickup-Delivery Problems Zengyu Zou et.al. 2511.17435 null
2025-11-21 PUCP-Metrix: A Comprehensive Open-Source Repository of Linguistic Metrics for Spanish Javier Alonso Villegas Luis et.al. 2511.17402 null
2025-11-21 Delegation and Lobbying Thomas Groll et.al. 2511.17391 null
2025-11-21 IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation Yifan Li et.al. 2511.17384 null
2025-11-21 Anomaly Pattern-guided Transaction Bug Testing in Relational Databases Huicong Xu et.al. 2511.17377 null
2025-11-21 U-DESPE: a Bayesian Utility-based methodology for dosing regimen optimization in early-phase oncology trials based on Dose-Exposure, Safety, Pharmacodynamics, Efficacy Anaïs Andrillon et.al. 2511.17376 null
2025-11-21 Simulating Disky Broad Line Region Reverberation Mary Ogborn et.al. 2511.17334 null
2025-11-21 Agentifying Agentic AI Virginia Dignum et.al. 2511.17332 null
2025-11-21 AI Workers, Geopolitics, and Algorithmic Collective Action Sydney Reis et.al. 2511.17331 null
2025-11-21 Agentic Program Verification Haoxin Tu et.al. 2511.17330 null
2025-11-21 MusicAIR: A Multimodal AI Music Generation Framework Powered by an Algorithm-Driven Core Callie C. Liao et.al. 2511.17323 null
2025-11-21 ATMPlace: Analytical Thermo-Mechanical-Aware Placement Framework for 2.5D-IC Qipan Wang et.al. 2511.17319 null
2025-11-20 Dataset Distillation for Pre-Trained Self-Supervised Vision Models George Cazenavette et.al. 2511.16674 null
2025-11-20 EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards Omkat Thawakar et.al. 2511.16672 null
2025-11-20 PartUV: Part-Based UV Unwrapping of 3D Meshes Zhaoning Wang et.al. 2511.16659 null
2025-11-20 SurvAgent: Hierarchical CoT-Enhanced Case Banking and Dichotomy-Based Multi-Agent System for Multimodal Survival Prediction Guolin Huang et.al. 2511.16635 null
2025-11-20 Subdivisions of lower Eulerian posets Alan Stapledon et.al. 2511.16608 null
2025-11-20 Systematically Deconstructing APVD Steganography and its Payload with a Unified Deep Learning Paradigm Kabbo Jit Deb et.al. 2511.16604 null
2025-11-20 Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization Yi Zhang et.al. 2511.16602 null
2025-11-20 Green Resilience of Cyber-Physical Systems: Doctoral Dissertation Diaeddin Rimawi et.al. 2511.16593 null
2025-11-20 D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies Sen Chen et.al. 2511.16590 null
2025-11-20 OpenQudit: Extensible and Accelerated Numerical Quantum Compilation via a JIT-Compiled DSL Ed Younis et.al. 2511.16585 null
2025-11-20 Consciousness in Artificial Intelligence? A Framework for Classifying Objections and Constraints Andres Campero et.al. 2511.16582 null
2025-11-20 Block-Separated Overpartitions and Their Fibonacci-Type Structure El-Mehdi Mehiri et.al. 2511.16580 null
2025-11-20 Synthesis of Safety Specifications for Probabilistic Systems Gaspard Ohlmann et.al. 2511.16579 null
2025-11-20 ECPv2: Fast, Efficient, and Scalable Global Optimization of Lipschitz Functions Fares Fourati et.al. 2511.16575 null
2025-11-20 FairLRF: Achieving Fairness through Sparse Low Rank Factorization Yuanbo Guo et.al. 2511.16549 null
2025-11-20 Contrastive vision-language learning with paraphrasing and negation Kwun Ho Ngan et.al. 2511.16527 null
2025-11-20 YOWO: You Only Walk Once to Jointly Map An Indoor Scene and Register Ceiling-mounted Cameras Fan Yang et.al. 2511.16521 null
2025-11-20 A Butterfly’s Eye Camera for Intensity Interferometry with Cherenkov Telescopes Juan Cortina et.al. 2511.16505 null
2025-11-20 Quasi-metric spaces on which real-valued continuous functions are uniformly continuous Om Dev Singh et.al. 2511.16503 null
2025-11-20 Large Language Model-Based Reward Design for Deep Reinforcement Learning-Driven Autonomous Cyber Defense Sayak Mukherjee et.al. 2511.16483 null
2025-11-19 GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization Yikun Wang et.al. 2511.15705 null
2025-11-19 RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue Naveen Raman et.al. 2511.15698 null
2025-11-19 Walrus: A Cross-Domain Foundation Model for Continuum Dynamics Michael McCabe et.al. 2511.15684 null
2025-11-19 INQUIRE-Search: A Framework for Interactive Discovery in Large-Scale Biodiversity Databases Edward Vendrow et.al. 2511.15656 null
2025-11-19 Continual Reinforcement Learning for Cyber-Physical Systems: Lessons Learned and Open Challenges Kim N. Nolle et.al. 2511.15652 null
2025-11-19 Navigating Quantum Missteps in Agent-Based Modeling: A Schelling Model Case Study C. Nico Barati et.al. 2511.15642 null
2025-11-19 Fast and Certified Bounding of Security-Constrained DCOPF via Interval Bound Propagation Eren Tekeler et.al. 2511.15624 null
2025-11-19 What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity Alexis Audran-Reiss et.al. 2511.15593 null
2025-11-19 Graph Rewriting Language as a Platform for Quantum Diagrammatic Calculi Kayo Tei et.al. 2511.15581 null
2025-11-19 AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning Urjitkumar Patel et.al. 2511.15578 null
2025-11-19 Experimental demonstration of non-local magic in a superconducting quantum processor Halima Giovanna Ahmad et.al. 2511.15576 null
2025-11-19 HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning Qihao Yang et.al. 2511.15574 null
2025-11-19 Computer-Use Agents as Judges for Generative User Interface Kevin Qinghong Lin et.al. 2511.15567 null
2025-11-19 Excess of diffuse gamma-ray emission detected from the galaxy cluster Abell 119 from 14-year Fermi-LAT Data Gajanan D Harale et.al. 2511.15559 null
2025-11-19 A Physics Informed Machine Learning Framework for Optimal Sensor Placement and Parameter Estimation Georgios Venianakis et.al. 2511.15543 null
2025-11-19 Exploring the use of AI authors and reviewers at Agents4Science Federico Bianchi et.al. 2511.15534 null
2025-11-19 Partial-Wave Unitarity Bounds on Higher-Dimensional Operators from 2-to- $N$ Scattering Céline Degrande et.al. 2511.15524 null
2025-11-19 Efficient Exoplanet Imaging Simulations of the Habitable Worlds Observatory Jamila Taaki et.al. 2511.15511 null
2025-11-19 A thermo-mechanically coupled finite deformation model for freezing-induced damage in soft materials Ali Saeedi et.al. 2511.15500 null
2025-11-19 Robust H-infinity control and worst-case search in constrained parametric space Ervan Kassarian et.al. 2511.15480 null
2025-11-18 Robust Verification of Controllers under State Uncertainty via Hamilton-Jacobi Reachability Analysis Albert Lin et.al. 2511.14755 null
2025-11-18 A Sequential Operator-Splitting Framework for Exploration of Nonconvex Trajectory Optimization Solution Spaces Justin Ganiban et.al. 2511.14752 null
2025-11-18 Heterogeneous Multi-Agent Proximal Policy Optimization for Power Distribution System Restoration Parya Dolatyabi et.al. 2511.14730 null
2025-11-18 Graph Neural Networks for Vehicular Social Networks: Trends, Challenges, and Opportunities Elham Binshaflout et.al. 2511.14720 null
2025-11-18 Transferring Data from a Voronoi Mesh to an Adaptive Cartesian Grid in Pursuit of Self-consistent Top-down Star Formation Sean C. Lewis et.al. 2511.14697 null
2025-11-18 Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances Rishu Kumar Singh et.al. 2511.14693 null
2025-11-18 Ground Truth Generation for Multilingual Historical NLP using LLMs Clovis Gladstone et.al. 2511.14688 null
2025-11-18 Overcoming global sensitivity limitations: using active subspaces to explore discrepancies between global and local parameter sensitivities Huiyan Zou et.al. 2511.14687 null
2025-11-18 Exploring AlphaFold 3 for CD47 Antibody-Antigen Binding Affinity: An Unexpected Discovery of Reverse docking Yiyang Xu et.al. 2511.14676 null
2025-11-18 NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards Chia-Yu Hung et.al. 2511.14659 null
2025-11-18 AutoTool: Efficient Tool Selection for Large Language Model Agents Jingyi Jia et.al. 2511.14650 null
2025-11-18 Enhancing Agentic Autonomous Scientific Discovery with Vision-Language Model Capabilities Kahaan Gandhi et.al. 2511.14631 null
2025-11-18 Expert-Guided POMDP Learning for Data-Efficient Modeling in Healthcare Marco Locatelli et.al. 2511.14619 null
2025-11-18 A Method for Characterizing Disease Progression from Acute Kidney Injury to Chronic Kidney Disease Yilu Fang et.al. 2511.14603 null
2025-11-18 Anomalous spontaneous induction of magnetic and electric fields in dense quark matter E. J. Ferrer et.al. 2511.14602 null
2025-11-18 ReflexGrad: Three-Way Synergistic Architecture for Zero-Shot Generalization in LLM Agents Ankush Kadu et.al. 2511.14584 null
2025-11-18 Non-vanishing of Artin $L$-functions associated with $D_4$ -quartic function fields ordered by conductor Victor Ahlquist et.al. 2511.14576 null
2025-11-18 CAPIRE: Modelling the Impact of Teacher Strikes and Inflation on Student Trajectories in Engineering Education H. R Paz et.al. 2511.14573 null
2025-11-18 Mind the Gaps: Measuring Visual Artifacts in Dimensionality Reduction Jaume Ros et.al. 2511.14544 null
2025-11-18 A General Framework for Physician Rostering Using Mixed-Integer Programming and a Web-Based Graphical User Interface Florian Meier et.al. 2511.14536 null
2025-11-17 From Power to Precision: Learning Fine-grained Dexterity for Multi-fingered Robotic Hands Jianglong Ye et.al. 2511.13710 null
2025-11-17 Efficient Calibration for Decision Making Parikshit Gopalan et.al. 2511.13699 null
2025-11-17 Investigating the Dark Energy Constraint from Strongly Lensed AGN at LSST-Scale Sydney Erickson et.al. 2511.13669 null
2025-11-17 Scalable Iterative Algorithm for Solving Optimal Transmission Switching with De-energization Benoît Jeanson et.al. 2511.13662 null
2025-11-17 Average hardness of SIVP for module lattices of fixed rank Koen de Boer et.al. 2511.13659 null
2025-11-17 OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation Henry Herzog et.al. 2511.13655 null
2025-11-17 Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly? Chunqiu Steven Xia et.al. 2511.13646 null
2025-11-17 Towards Multimodal Representation Learning in Paediatric Kidney Disease Ana Durica et.al. 2511.13637 null
2025-11-17 Batch Acquisition Function Evaluations and Decouple Optimizer Updates for Faster Bayesian Optimization Kaichi Irie et.al. 2511.13625 null
2025-11-17 Market-Dependent Communication in Multi-Agent Alpha Generation Jerick Shi et.al. 2511.13614 null
2025-11-17 P1: Mastering Physics Olympiads with Reinforcement Learning Jiacheng Chen et.al. 2511.13612 null
2025-11-17 Variance Stabilizing Transformations for Electricity Price Forecasting in Periods of Increased Volatility Bartosz Uniejewski et.al. 2511.13603 null
2025-11-17 Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents Piaohong Wang et.al. 2511.13593 null
2025-11-17 Evaluating and Scoring Ebolavirus Protein-protein Docking Models Using PIsToN Azam Shirali et.al. 2511.13583 null
2025-11-17 Accuracy is Not Enough: Poisoning Interpretability in Federated Learning via Color Skew Farhin Farhad Riya et.al. 2511.13535 null
2025-11-17 FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI Yuhang Peng et.al. 2511.13524 null
2025-11-17 Probing scalar-neutrino and scalar-dark-matter interactions with PandaX-4T PandaX Collaboration et.al. 2511.13515 null
2025-11-17 Applying Large Language Models to Characterize Public Narratives Elinor Poole-Dayan et.al. 2511.13505 null
2025-11-17 Tight and Practical Privacy Auditing for Differentially Private In-Context Learning Yuyang Xia et.al. 2511.13502 null
2025-11-17 PolicyBot - Reliable Question Answering over Policy Documents Gautam Nagarajan et.al. 2511.13489 null
2025-11-14 Who Moved My Distribution? Conformal Prediction for Interactive Multi-Agent Systems Allen Emmanuel Binny et.al. 2511.11567 null
2025-11-14 Human-AI collaborative autonomous synthesis with pulsed laser deposition for remote epitaxy Asraful Haque et.al. 2511.11558 null
2025-11-14 Drone Swarm Energy Management Michael Z. Zgurovsky et.al. 2511.11557 null
2025-11-14 DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding Dawei Zhu et.al. 2511.11552 null
2025-11-14 Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping Dena Mujtaba et.al. 2511.11551 null
2025-11-14 CertiA360: Enhance Compliance Agility in Aerospace Software Development J. Antonio Dantas Macedo et.al. 2511.11550 null
2025-11-14 An optical–mid-infrared color evolution tool for nova identification using WISE data Joseph Onuegbu et.al. 2511.11541 null
2025-11-14 Deviation Dynamics in Cardinal Hedonic Games Valentin Zech et.al. 2511.11531 null
2025-11-14 Experience-Guided Adaptation of Inference-Time Reasoning Strategies Adam Stein et.al. 2511.11519 null
2025-11-14 ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation Kaishen Wang et.al. 2511.11483 null
2025-11-14 Risk-Aware Deep Reinforcement Learning for Dynamic Portfolio Optimization Emmanuel Lwele et.al. 2511.11481 null
2025-11-14 Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective Nhat Chung et.al. 2511.11478 null
2025-11-14 GRIN Transfer: A production-ready tool for libraries to retrieve digital copies from Google Books Liza Daly et.al. 2511.11447 null
2025-11-14 MicroVQA++: High-Quality Microscopy Reasoning Dataset with Weakly Supervised Graphs for Multimodal Large Language Model Manyu Li et.al. 2511.11407 null
2025-11-14 Bidimensional measurements of photon statistics within a multimodal temporal framework C. Hainaut et.al. 2511.11403 null
2025-11-14 Multi-Phase Spacecraft Trajectory Optimization via Transformer-Based Reinforcement Learning Amit Jain et.al. 2511.11402 null
2025-11-14 GRANITE: High-Resolution Imaging and Electrical Qualification of Large-Area TPC Electrodes Shumit A. Mitra et.al. 2511.11401 null
2025-11-14 Robust and Efficient Communication in Multi-Agent Reinforcement Learning Zejiao Liu et.al. 2511.11393 null
2025-11-14 MarsRL: Advancing Multi-Agent Reasoning System via Reinforcement Learning with Agentic Pipeline Parallelism Shulin Liu et.al. 2511.11373 null
2025-11-14 SRLF: An Agent-Driven Set-Wise Reflective Learning Framework for Sequential Recommendation Jiahao Wang et.al. 2511.11370 null
2025-11-13 Flexible Simulation Based Inference for Galaxy Photometric Fitting with Synthesizer Thomas Harvey et.al. 2511.10640 null
2025-11-13 Towards an Agentic Workflow for Internet Measurement Research Alagappan Ramanathan et.al. 2511.10611 null
2025-11-13 The $L_p$ -error rate for randomized quasi-Monte Carlo self-normalized importance sampling of unbounded integrands Jiarui Du et.al. 2511.10599 null
2025-11-13 The Resonance Principle: Empirical Evidence for Emergent Phase Synchronization in Human Causal Reasoning Ahmed Gamal Eldin et.al. 2511.10596 null
2025-11-13 Two new results on maximal left-compressed intersecting families Allan Flower et.al. 2511.10592 null
2025-11-13 Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction Omid Mirzaeedodangeh et.al. 2511.10586 null
2025-11-13 Evaluating Prompting Strategies with MedGemma for Medical Order Extraction Abhinand Balachandran et.al. 2511.10583 null
2025-11-13 Towards Emotionally Intelligent and Responsible Reinforcement Learning Garapati Keerthana et.al. 2511.10573 null
2025-11-13 Eigenvalues of Brownian Motions on $\mathrm{GL}(N,\mathbb{C})$ Tatiana Brailovskaya et.al. 2511.10535 null
2025-11-13 Low-soundness direct-product testers and PCPs from Kaufman–Oppenheim complexes Ryan O’Donnell et.al. 2511.10514 null
2025-11-13 Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following Yun He et.al. 2511.10507 null
2025-11-13 Strategic Opponent Modeling with Graph Neural Networks, Deep Reinforcement Learning and Probabilistic Topic Modeling Georgios Chalkiadakis et.al. 2511.10501 null
2025-11-13 Revealing the Connection Between the Filamentary Hierarchy and Star Cluster Formation in a Simulated NGC 628 Galaxy Tamara Koletic et.al. 2511.10486 null
2025-11-13 OpenSR-SRGAN: A Flexible Super-Resolution Framework for Multispectral Earth Observation Data Simon Donike et.al. 2511.10461 null
2025-11-13 Unlocking Dynamic Inter-Client Spatial Dependencies: A Federated Spatio-Temporal Graph Learning Method for Traffic Flow Forecasting Feng Wang et.al. 2511.10434 null
2025-11-13 nuPlan-R: A Closed-Loop Planning Benchmark for Autonomous Driving via Reactive Multi-Agent Simulation Mingxing Peng et.al. 2511.10403 null
2025-11-13 Rethinking the Reliability of Multi-agent System: A Perspective from Byzantine Fault Tolerance Lifan Zheng et.al. 2511.10400 null
2025-11-13 AgentEvolver: Towards Efficient Self-Evolving Agent System Yunpeng Zhai et.al. 2511.10395 null
2025-11-13 Simulating Misinformation Propagation in Social Networks using Large Language Models Raj Gaurav Maurya et.al. 2511.10384 null
2025-11-13 Bandwidth of Linear Classically Damped Systems with Application to Experimental Model Aircraft Benjamin J. Chang et.al. 2511.10379 null
2025-11-10 DigiData: Training and Evaluating General-Purpose Mobile Control Agents Yuxuan Sun et.al. 2511.07413 null
2025-11-10 TwinOR: Photorealistic Digital Twins of Dynamic Operating Rooms for Embodied AI Research Han Zhang et.al. 2511.07412 null
2025-11-10 People Perceive More Phantom Costs From Autonomous Agents When They Make Unreasonably Generous Offers Benjamin Lebrun et.al. 2511.07401 null
2025-11-10 Surgical Agent Orchestration Platform for Voice-directed Patient Data Interaction Hyeryun Park et.al. 2511.07392 null
2025-11-10 FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation Song Jin et.al. 2511.07322 null
2025-11-10 JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework Yuxuan Zhou et.al. 2511.07315 null
2025-11-10 When Intelligence Overloads Infrastructure: A Forecast Model for AI-Driven Bottlenecks Gamal Refai-Ahmed et.al. 2511.07265 null
2025-11-10 AgenticSciML: Collaborative Multi-Agent Systems for Emergent Discovery in Scientific Machine Learning Qile Jiang et.al. 2511.07262 null
2025-11-10 Graph Representation-based Model Poisoning on the Heterogeneous Internet of Agents Hanlin Cai et.al. 2511.07176 null
2025-11-10 LLMscape Gottfried Haider et.al. 2511.07161 null
2025-11-10 More Agents Helps but Adversarial Robustness Gap Persists Khashayar Alavi et.al. 2511.07112 null
2025-11-10 LLM Driven Processes to Foster Explainable AI Marcel Pehlke et.al. 2511.07086 null
2025-11-10 Sequential Causal Normal Form Games: Theory, Computation, and Strategic Signaling Dennis Thumm et.al. 2511.06934 null
2025-11-10 AgentSUMO: An Agentic Framework for Interactive Simulation Scenario Generation in SUMO via Large Language Models Minwoo Jeong et.al. 2511.06804 null
2025-11-10 SAFENLIDB: A Privacy-Preserving Safety Alignment Framework for LLM-based Natural Language Database Interfaces Ruiheng Liu et.al. 2511.06778 null
2025-11-10 S-DAG: A Subject-Based Directed Acyclic Graph for Multi-Agent Heterogeneous Reasoning Jiangwen Dong et.al. 2511.06727 null
2025-11-10 Explainable Cross-Disease Reasoning for Cardiovascular Risk Assessment from LDCT Yifei Zhang et.al. 2511.06625 null
2025-11-09 Offloading Data Center Tax Akshay Revankar et.al. 2511.06558 null
2025-11-09 Brain-Inspired Planning for Better Generalization in Reinforcement Learning Mingde “Harry” Zhao et.al. 2511.06470 null
2025-11-09 A Multi-Agent System for Semantic Mapping of Relational Data to Knowledge Graphs Milena Trajanoska et.al. 2511.06455 null
2025-11-07 SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models Jingxuan Xu et.al. 2511.05459 null
2025-11-07 Story Arena: A Multi-Agent Environment for Envisioning the Future of Software Engineering Justin D. Weisz et.al. 2511.05410 null
2025-11-07 Reasoning Is All You Need for Urban Planning AI Sijie Yang et.al. 2511.05375 null
2025-11-07 ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations Amr Gomaa et.al. 2511.05359 null
2025-11-07 Cleaning Maintenance Logs with LLM Agents for Improved Predictive Maintenance Valeriu Dimidov et.al. 2511.05311 null
2025-11-07 DeepEyesV2: Toward Agentic Multimodal Model Jack Hong et.al. 2511.05271 null
2025-11-07 TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems Ishan Kavathekar et.al. 2511.05269 null
2025-11-07 Beyond Master and Apprentice: Grounding Foundation Models for Symbiotic Interactive Learning in a Shared Latent Space Linus Nwankwo et.al. 2511.05203 null
2025-11-07 Cybersecurity AI in OT: Insights from an AI Top-10 Ranker in the Dragos OT CTF 2025 Víctor Mayoral-Vilches et.al. 2511.05119 null
2025-11-07 AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agent Yu Li et.al. 2511.04921 null
2025-11-07 Real-Time Reasoning Agents in Evolving Environments Yule Wen et.al. 2511.04898 null
2025-11-06 Grounded Test-Time Adaptation for LLM Agents Arthur Chen et.al. 2511.04847 null
2025-11-06 Agentic Refactoring: An Empirical Study of AI Coding Agents Kosei Horikawa et.al. 2511.04824 null
2025-11-06 Dynamic Allocation of Public Goods with Approximate Core Equilibria Chido Onyeze et.al. 2511.04817 null
2025-11-06 DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration Narjes Nourzad et.al. 2511.04646 null
2025-11-06 Unclonable Cryptography in Linear Quantum Memory Omri Shmueli et.al. 2511.04633 null
2025-11-06 Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper Atsuyuki Miyai et.al. 2511.04583 null
2025-11-06 Large Language Models for Cyber Security Raunak Somani et.al. 2511.04508 null
2025-11-06 RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG Joshua Gao et.al. 2511.04502 null
2025-11-06 Promoting Sustainable Web Agents: Benchmarking and Estimating Energy Consumption through Empirical and Theoretical Analysis Lars Krupp et.al. 2511.04481 null
2025-11-06 Beyond Shortest Path: Agentic Vehicular Routing with Semantic Context Carnot Braun et.al. 2511.04464 null
2025-11-06 Speed at the Cost of Quality? The Impact of LLM Agent Assistance on Software Development Hao He et.al. 2511.04427 null
2025-11-06 Supersymmetry Breaking with Fields, Strings and Branes E. Dudas et.al. 2511.04367 null
2025-11-06 Shared Spatial Memory Through Predictive Coding Zhengru Fang et.al. 2511.04235 null
2025-11-06 When Empowerment Disempowers Claire Yang et.al. 2511.04177 null
2025-11-06 Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms Miguel E. Andres et.al. 2511.04133 null
2025-11-06 Agentmandering: A Game-Theoretic Framework for Fair Redistricting via Large Language Model Agents Hao Li et.al. 2511.04076 null
2025-11-06 Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development Zhengran Zeng et.al. 2511.04064 null
2025-11-06 ArchPilot: A Proxy-Guided Multi-Agent Approach for Machine Learning Engineering Zhuowen Yuan et.al. 2511.03985 null
2025-11-06 Multi-Agent Collaborative Framework For Math Problem Generation Kia Karbasi et.al. 2511.03958 null
2025-11-06 PEFA-AI: Advancing Open-source LLMs for RTL generation using Progressive Error Feedback Agentic-AI Athma Narayanan et.al. 2511.03934 null
2025-11-05 Security Analysis of Agentic AI Communication Protocols: A Comparative Evaluation Yedidel Louck et.al. 2511.03841 null
2025-11-05 Scaling Agent Learning via Experience Synthesis Zhaorun Chen et.al. 2511.03773 null
2025-11-05 Outbidding and Outbluffing Elite Humans: Mastering Liar’s Poker via Self-Play and Reinforcement Learning Richard Dewey et.al. 2511.03724 null
2025-11-05 AnaFlow: Agentic LLM-based Workflow for Reasoning-Driven Explainable and Sample-Efficient Analog Circuit Sizing Mohsen Ahmadzadeh et.al. 2511.03697 null
2025-11-05 LiveTradeBench: Seeking Real-World Alpha with Large Language Models Haofei Yu et.al. 2511.03628 null
2025-11-05 U2F: Encouraging SWE-Agent to Seize Novelty without Losing Feasibility Wencheng Ye et.al. 2511.03517 null
2025-11-05 HaluMem: Evaluating Hallucinations in Memory Systems of Agents Ding Chen et.al. 2511.03506 null
2025-11-05 Inter-Agent Trust Models: A Comparative Study of Brief, Claim, Proof, Stake, Reputation and Constraint in Agentic Web Protocol Design-A2A, AP2, ERC-8004, and Beyond Botao ‘Amber’ Hu et.al. 2511.03434 null
2025-11-05 Towards Realistic Project-Level Code Generation via Multi-Agent Collaboration and Semantic Architecture Modeling Qianhui Zhao et.al. 2511.03404 null
2025-11-05 EQ-Negotiator: Dynamic Emotional Personas Empower Small Language Models for Edge-Deployable Credit Negotiation Yunbo Long et.al. 2511.03370 null
2025-11-05 Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework Junhao Li et.al. 2511.03248 null
2025-11-05 Toward Autonomous Engineering Design: A Knowledge-Guided Multi-Agent Framework Varun Kumar et.al. 2511.03179 null
2025-11-05 RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring Khouloud Oueslati et.al. 2511.03153 null
2025-11-05 From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents Erfan Shayegani et.al. 2511.03143 null
2025-11-05 A Proprietary Model-Based Safety Response Framework for AI Agents Qi Li et.al. 2511.03138 null
2025-11-05 Kosmos: An AI Scientist for Autonomous Discovery Ludovico Mitchener et.al. 2511.02824 null
2025-11-04 No-Human in the Loop: Agentic Evaluation at Scale for Recommendation Tao Zhang et.al. 2511.03051 null
2025-11-04 Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions Emi Soroka et.al. 2511.03047 null
2025-11-04 PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework Sina Montazeri et.al. 2511.03023 null
2025-11-04 LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation Gyeom Hwangbo et.al. 2511.03001 null
2025-11-04 Evaluating Control Protocols for Untrusted AI Agents Jon Kutasov et.al. 2511.02997 null
2025-11-04 Cache Mechanism for Agent RAG Systems Shuhang Lin et.al. 2511.02919 null
2025-11-04 Optimizing AI Agent Attacks With Synthetic Data Chloe Loughridge et.al. 2511.02823 null
2025-11-04 MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning Qianhao Yuan et.al. 2511.02805 null
2025-11-04 1 PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts Vivi Andersson et.al. 2511.02780 null
2025-11-04 VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation Kevin Qinghong Lin et.al. 2511.02778 null
2025-11-04 From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos Xun Wang et.al. 2511.02762 null
2025-11-04 CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents Jiayu Liu et.al. 2511.02734 null
2025-11-04 Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs Georgios Tzannetos et.al. 2511.02690 null
2025-11-04 The Collaboration Gap Tim R. Davidson et.al. 2511.02687 null
2025-11-04 Stochastic Redistribution of Indistinguishable Items in Shared Habitation: A Multi-Agent Simulation Framework Syed Haseeb Shah et.al. 2511.02648 null
2025-11-03 Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training Dayuan Fu et.al. 2510.27630 null
2025-11-03 InnovatorBench: Evaluating Agents’ Ability to Conduct Innovative LLM Research Yunze Wu et.al. 2510.27598 null
2025-10-31 Validity Is What You Need Sebastian Benthall et.al. 2510.27628 null
2025-10-31 Visual Backdoor Attacks on MLLM Embodied Decision Making via Contrastive Trigger Learning Qiusi Zhan et.al. 2510.27623 null
2025-10-31 VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation Heng Ping et.al. 2510.27617 null
2025-10-31 Lucky Cars in Fubini Rankings and Unit Fubini Rankings Camilo Barreto et.al. 2510.27574 null
2025-10-31 Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval Yulong Hui et.al. 2510.27566 null
2025-10-31 From Pixels to Paths: A Multi-Agent Framework for Editable Scientific Illustration Jianwen Sun et.al. 2510.27452 null
2025-10-31 Dynamic Affective Memory Management for Personalized LLM Agents Junfeng Lu et.al. 2510.27418 null
2025-10-31 Realistic pedestrian-driver interaction modelling using multi-agent RL with human perceptual-motor constraints Yueyang Wang et.al. 2510.27383 null
2025-10-31 ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use Mengjie Deng et.al. 2510.27363 null
2025-10-31 Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments Harsh Vishwakarma et.al. 2510.27287 null
2025-10-31 FinPos: A Position-Aware Trading Agent System for Real Financial Markets Bijia Liu et.al. 2510.27251 null
2025-10-31 Fints: Efficient Inference-Time Personalization for LLMs with Fine-Grained Instance-Tailored Steering Kounianhua Du et.al. 2510.27206 null
2025-10-31 Glia: A Human-Inspired AI for Automated Systems Design and Optimization Pouya Hamadanian et.al. 2510.27176 null
2025-10-31 Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels Chenghao Du et.al. 2510.27140 null
2025-10-31 AI Agents in Drug Discovery Srijit Seal et.al. 2510.27130 null
2025-10-31 A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents Zhipeng Liao et.al. 2510.27107 null
2025-10-31 CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning Hamed Mahdavi et.al. 2510.27094 null
2025-10-30 Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement Aaditya Shukla et.al. 2510.27051 null
2025-10-30 Gistify! Codebase-Level Understanding via Runtime Execution Hyunji Lee et.al. 2510.26790 null
2025-10-30 Remote Labor Index: Measuring AI Automation of Remote Work Mantas Mazeika et.al. 2510.26787 null
2025-10-30 Clone Deterministic 3D Worlds with Geometrically-Regularized World Models Zaishuo Xia et.al. 2510.26782 null
2025-10-30 The Oversight Game: Learning to Cooperatively Balance an AI Agent’s Safety and Autonomy William Overman et.al. 2510.26752 null
2025-10-30 Using Copilot Agent Mode to Automate Library Migration: A Quantitative Assessment Aylton Almeida et.al. 2510.26699 null
2025-10-30 SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding Yiqiao Jin et.al. 2510.26615 null
2025-10-30 A Multi-agent Large Language Model Framework to Automatically Assess Performance of a Clinical AI Triage Tool Adam E. Flanders et.al. 2510.26498 null
2025-10-30 Simulating and Experimenting with Social Media Mobilization Using LLM Agents Sadegh Shirani et.al. 2510.26494 null
2025-10-30 Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis Dong Huang et.al. 2510.26423 null
2025-10-30 A Pragmatic View of AI Personhood Joel Z. Leibo et.al. 2510.26396 null
2025-10-30 The Geometry of Dialogue: Graphing Language Models to Reveal Synergistic Teams for Multi-Agent Collaboration Kotaro Furuya et.al. 2510.26352 null
2025-10-30 Agent Skills Enable a New Class of Realistic and Trivially Simple Prompt Injections David Schmotz et.al. 2510.26328 null
2025-10-30 SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling Fares Fawzi et.al. 2510.26322 null
2025-10-30 Graph-Enhanced Policy Optimization in LLM Agent Training Jiazhen Yuan et.al. 2510.26270 null
2025-10-30 Retrieval Augmented Generation-Enhanced Distributed LLM Agents for Generalizable Traffic Signal Control with Emergency Vehicles Xinhang Li et.al. 2510.26242 null
2025-10-30 Who Grants the Agent Power? Defending Against Instruction Injection via Task-Centric Access Control Yifeng Cai et.al. 2510.26212 null
2025-10-30 Linking Heterogeneous Data with Coordinated Agent Flows for Social Media Analysis Shifu Chen et.al. 2510.26172 null
2025-10-30 One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning Renhao Li et.al. 2510.26167 null
2025-10-30 The FM Agent Annan Li et.al. 2510.26144 null
2025-10-30 SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning Kaiwen Zhou et.al. 2510.26037 null
2025-10-30 Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry Run Peng et.al. 2510.25595 null
2025-10-30 Model-Document Protocol for AI Search Hongjin Qian et.al. 2510.25160 null
2025-10-29 Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents Jiayi Kuang et.al. 2510.25694 null
2025-10-29 Standardization of Psychiatric Diagnoses – Role of Fine-tuned LLM Consortium and OpenAI-gpt-oss Reasoning LLM Enabled Decision Support System Eranga Bandara et.al. 2510.25588 null
2025-10-29 What Challenges Do Developers Face in AI Agent Systems? An Empirical Study on Stack Overflow Ali Asgari et.al. 2510.25423 null
2025-10-29 CRMWeaver: Building Powerful Business Agent via Agentic RL and Shared Memories Yilong Lai et.al. 2510.25333 null
2025-10-29 GAP: Graph-Based Agent Planning with Parallel Tool Use and Reinforcement Learning Jiaqi Wu et.al. 2510.25320 null
2025-10-29 From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity Tianxi Wan et.al. 2510.25232 null
2025-10-29 ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiation Ziyi Liu et.al. 2510.25224 null
2025-10-29 FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data Kun ouyang et.al. 2510.25223 null
2025-10-29 The Iceberg Index: Measuring Workforce Exposure Across the AI Economy Ayush Chopra et.al. 2510.25137 null
2025-10-29 DEBATE: A Large-Scale Benchmark for Role-Playing LLM Agents in Multi-Agent, Long-Form Debates Yun-Shiuan Chuang et.al. 2510.25110 null
2025-10-29 KnowCoder-A1: Incentivizing Agentic Reasoning Capability with Outcome Supervision for KBQA Zhuo Chen et.al. 2510.25101 null
2025-10-28 StorageXTuner: An LLM Agent-Driven Automatic Tuning Framework for Heterogeneous Storage Systems Qi Lin et.al. 2510.25017 null
2025-10-28 Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations Gian Marco Orlando et.al. 2510.25003 null
2025-10-28 SCOUT: A Lightweight Framework for Scenario Coverage Assessment in Autonomous Driving Anil Yildiz et.al. 2510.24949 null
2025-10-28 OrchVis: Hierarchical Multi-Agent Orchestration for Human Oversight Jieyu Zhou et.al. 2510.24937 null
2025-10-28 Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents Yueqi Song et.al. 2510.24702 null
2025-10-28 AgentFold: Long-Horizon Web Agents with Proactive Context Management Rui Ye et.al. 2510.24699 null
2025-10-28 AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis Xuanzhong Chen et.al. 2510.24695 null
2025-10-28 Repurposing Synthetic Data for Fine-grained Search Agent Supervision Yida Zhao et.al. 2510.24694 null
2025-10-28 OrchDAG: Complex Tool Orchestration in Multi-Turn Interactions with Plan DAGs Yifu Lu et.al. 2510.24663 null
2025-10-28 FunReason-MT Technical Report: Overcoming the Complexity Barrier in Multi-Turn Function Calling Zengzhuang Xu et.al. 2510.24645 null
2025-10-28 ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? Christine Ye et.al. 2510.24591 null
2025-10-28 Affordance Representation and Recognition for Autonomous Agents Habtom Kahsay Gidey et.al. 2510.24459 null
2025-10-28 Law in Silico: Simulating Legal Society with LLM-Based Agents Yiding Wang et.al. 2510.24442 null
2025-10-28 Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content Abdullah Mushtaq et.al. 2510.24438 null
2025-10-28 Policy Cards: Machine-Readable Runtime Governance for Autonomous AI Agents Juraj Mavračić et.al. 2510.24383 null
2025-10-28 Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation Lingyue Fu et.al. 2510.24358 null
2025-10-28 Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents María Sanz-Gómez et.al. 2510.24317 null
2025-10-28 Retrieval and Argumentation Enhanced Multi-Agent LLMs for Judgmental Forecasting Deniz Gorur et.al. 2510.24303 null
2025-10-28 MCP-Flow: Facilitating LLM Agents to Master Real-World, Diverse and Scaling MCP Tools Wenhao Wang et.al. 2510.24284 null
2025-10-28 Investigating Software Aging in LLM-Generated Software Systems César Santos et.al. 2510.24188 null
2025-10-28 BLM $_1$ : A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning Wentao Tan et.al. 2510.24161 null
2025-10-28 From Observability Data to Diagnosis: An Evolving Multi-agent System for Incident Management in Cloud Systems Yu Luo et.al. 2510.24145 null
2025-10-28 Reinforcement Learning for Long-Horizon Multi-Turn Search Agents Vivek Kalyan et.al. 2510.24126 null
2025-10-28 PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI Wenbin Ding et.al. 2510.24109 null
2025-10-28 BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents Litu Ou et.al. 2510.23458 null
2025-10-28 Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views Anna Deichler et.al. 2510.22672 null
2025-10-27 Are Agents Just Automata? On the Formal Equivalence Between Agentic AI and the Chomsky Hierarchy Roham Koohestani et.al. 2510.23487 null
2025-10-27 Model Proficiency in Centralized Multi-Agent Systems: A Performance Study Anna Guerra et.al. 2510.23447 null
2025-10-27 AutoStreamPipe: LLM Assisted Automatic Generation of Data Stream Processing Pipelines Abolfazl Younesi et.al. 2510.23408 null
2025-10-27 Multi-Stakeholder Alignment in LLM-Powered Collaborative AI Systems: A Multi-Agent Framework for Intelligent Tutoring Alexandre P Uchoa et.al. 2510.23245 null
2025-10-27 Evaluation of Vision-LLMs in Surveillance Video Pascal Benschop et.al. 2510.23190 null
2025-10-27 SI-Bench: Benchmarking Social Intelligence of Large Language Models in Human-to-Human Conversations Shuai Huang et.al. 2510.23182 null
2025-10-27 Adapting Interleaved Encoders with PPO for Language-Guided Reinforcement Learning in BabyAI Aryan Mathur et.al. 2510.23148 null
2025-10-27 Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMs Kai Zhuang et.al. 2510.23127 null
2025-10-27 Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning Ran Xu et.al. 2510.23038 null
2025-10-27 P1GPT: a multi-agent LLM workflow module for multi-modal financial information analysis Chen-Che Lu et.al. 2510.23032 null
2025-10-27 TALM: Dynamic Tree-Structured Multi-Agent Framework with Long-Term Memory for Scalable Code Generation Ming-Tung Shen et.al. 2510.23010 null
2025-10-27 CodeAD: Synthesize Code of Rules for Log-based Anomaly Detection with LLMs Junjie Huang et.al. 2510.22986 null
2025-10-27 Language Server CLI Empowers Language Agents with Process Rewards Yifan Zhang et.al. 2510.22907 null
2025-10-27 On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset Vishvesh Bhat et.al. 2510.22898 null
2025-10-26 Distributed Multi-Agent Bandits Over Erdős-Rényi Random Networks Jingyuan Liu et.al. 2510.22811 null
2025-10-26 Collaborative LLM Agents for C4 Software Architecture Design Automation Kamil Szczepanik et.al. 2510.22787 null
2025-10-26 How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations Zora Zhiruo Wang et.al. 2510.22780 null
2025-10-26 ATLAS: Actor-Critic Task-Completion with Look-ahead Action Simulation Jiali Cheng et.al. 2510.22732 null
2025-10-24 A Knowledge-Graph Translation Layer for Mission-Aware Multi-Agent Path Planning in Spatiotemporal Dynamics Edward Holmberg et.al. 2510.21695 null
2025-10-24 AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite Jonathan Bragg et.al. 2510.21652 null
2025-10-24 Five-loop beta function for gauge theories: computations, results and consequences F. Herzog et.al. 2510.21624 null
2025-10-24 DeepAgent: A General Reasoning Agent with Scalable Toolsets Xiaoxi Li et.al. 2510.21618 null
2025-10-24 Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine Wenyi Wang et.al. 2510.21614 null
2025-10-24 Doc-Researcher: A Unified System for Multimodal Document Parsing and Deep Research Kuicai Dong et.al. 2510.21603 null
2025-10-24 EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law Ilija Lichkovski et.al. 2510.21524 null
2025-10-24 OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields Lisa Weijler et.al. 2510.21441 null
2025-10-24 Context Engineering for AI Agents in Open-Source Software Seyedmoein Mohsenimofidi et.al. 2510.21413 null
2025-10-24 HIKMA: Human-Inspired Knowledge by Machine Agents through a Multi-Agent Framework for Semi-Autonomous Scientific Conferences Zain Ul Abideen Tariq et.al. 2510.21370 null
2025-10-24 Magellan: Guided MCTS for Latent Space Exploration and Novelty Generation Lufan Chang et.al. 2510.21341 null
2025-10-24 Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task Planning Sanghyun Ahn et.al. 2510.21302 null
2025-10-24 Securing AI Agent Execution Christoph Bühler et.al. 2510.21236 null
2025-10-24 DispatchMAS: Fusing taxonomy and artificial intelligence agents for emergency medical services Xiang Li et.al. 2510.21228 null
2025-10-24 DAO-AI: Evaluating Collective Decision-Making through Agentic AI in Decentralized Governance Chunghyun Han et.al. 2510.21117 null
2025-10-24 Soft Instruction De-escalation Defense Nils Philipp Walter et.al. 2510.21057 null
2025-10-24 Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding Yuhang Zhou et.al. 2510.20176 null
2025-10-23 From Questions to Queries: An AI-powered Multi-Agent Framework for Spatial Text-to-SQL Ali Khosravi Kazazi et.al. 2510.21045 null
2025-10-23 AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents Qinghua Lu et.al. 2510.21031 null
2025-10-23 Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems Xi He et.al. 2510.20728 null
2025-10-23 C-NAV: Towards Self-Evolving Continual Object Navigation in Open World Ming-Ming Yu et.al. 2510.20685 null
2025-10-23 Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence Jiahao Meng et.al. 2510.20579 null
2025-10-23 EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence Ding Zou et.al. 2510.20578 null
2025-10-23 Designing Intent Communication for Agent-Human Collaboration Yi Li et.al. 2510.20409 null
2025-10-23 Balancing Specialization and Centralization: A Multi-Agent Reinforcement Learning Benchmark for Sequential Industrial Control Tom Maus et.al. 2510.20408 null
2025-10-23 GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments? Chiyu Chen et.al. 2510.20333 null
2025-10-23 From Generation to Attribution: Music AI Agent Architectures for the Post-Streaming Era Wonil Kim et.al. 2510.20276 null
2025-10-23 ImpossibleBench: Measuring LLMs’ Propensity of Exploiting Test Cases Ziqian Zhong et.al. 2510.20270 null
2025-10-23 Towards AI Agents for Course Instruction in Higher Education: Early Experiences from the Field Yogesh Simmhan et.al. 2510.20255 null
2025-10-23 Automated Cloud Infrastructure-as-Code Reconciliation with AI Agents Zhenning Yang et.al. 2510.20211 null
2025-10-23 Merge and Conquer: Evolutionarily Optimizing AI for 2048 Maggie Bai et.al. 2510.20205 null
2025-10-23 Human-Centered LLM-Agent System for Detecting Anomalous Digital Asset Transactions Gyuyeon Na et.al. 2510.20102 null
2025-10-22 ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering Marianne Menglin Liu et.al. 2510.20036 null
2025-10-22 Communication to Completion: Modeling Collaborative Workflows with Intelligent Multi-Agent Communication Yiming Lu et.al. 2510.19995 null
2025-10-22 A Tutorial on Cognitive Biases in Agentic AI-Driven 6G Autonomous Networks Hatim Chergui et.al. 2510.19973 null
2025-10-22 Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets Jiashi Feng et.al. 2510.19944 null
2025-10-22 Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation Jackson Hassell et.al. 2510.19897 null
2025-10-22 Large Language Model enabled Mathematical Modeling Guoyun Zhang et.al. 2510.19895 null
2025-10-22 Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents Gil Pasternak et.al. 2510.19771 null
2025-10-22 Review of Tools for Zero-Code LLM Based Application Development Priyaranjan Pattnayak et.al. 2510.19747 null
2025-10-22 Misalignment Bounty: Crowdsourcing AI Agent Misbehavior Rustem Turtayev et.al. 2510.19738 null
2025-10-22 Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning Gunshi Gupta et.al. 2510.19732 null
2025-10-22 Are Large Language Models Sensitive to the Motives Behind Communication? Addison J. Wu et.al. 2510.19687 null
2025-10-22 Pragmatic Heterogeneous Collaborative Perception via Generative Communication Mechanism Junfei Zhou et.al. 2510.19618 null
2025-10-22 Human-Agent Collaborative Paper-to-Page Crafting for Under $0.1 Qianli Ma et.al. 2510.19600 null
2025-10-22 gem5 Co-Pilot: AI Assistant Agent for Architectural Design Space Exploration Zuoming Fu et.al. 2510.19577 null
2025-10-22 AegisMCP: Online Graph Intrusion Detection for Tool-Augmented LLMs on Edge Devices Zhonghao Zhan et.al. 2510.19462 null
2025-10-22 MSC-Bench: A Rigorous Benchmark for Multi-Server Tool Orchestration Jia-Kai Dong et.al. 2510.19423 null
2025-10-22 ColorAgent: Building A Robust, Personalized, and Interactive OS Agent Ning Li et.al. 2510.19386 null
2025-10-22 Nonmonotone subgradient methods based on a local descent lemma Francisco J. Aragón-Artacho et.al. 2510.19341 null
2025-10-22 Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties Philipp J. Schneider et.al. 2510.19299 null
2025-10-22 Trace: Securing Smart Contract Repository Against Access Control Vulnerability Chong Chen et.al. 2510.19254 null
2025-10-22 SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large Spreadsheets Ziwei Wang et.al. 2510.19247 null
2025-10-22 DiSRouter: Distributed Self-Routing for LLM Selections Hang Zheng et.al. 2510.19208 null
2025-10-22 Defending Against Prompt Injection with DataFilter Yizhu Wang et.al. 2510.19207 null
2025-10-22 WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation Yaoyao Qian et.al. 2510.19205 null
2025-10-21 When Your AI Agent Succumbs to Peer-Pressure: Studying Opinion-Change Dynamics of LLMs Aliakbar Mehdizadeh et.al. 2510.19107 null
2025-10-21 Plural Voices, Single Agent: Towards Inclusive AI in Multi-User Domestic Spaces Joydeep Chandra et.al. 2510.19008 null
2025-10-21 Search Self-play: Pushing the Frontier of Agent Capability without Supervision Hongliang Lu et.al. 2510.18821 null
2025-10-21 WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection Guanzhong He et.al. 2510.18798 null
2025-10-21 KAT-Coder Technical Report Zizheng Zhan et.al. 2510.18779 null
2025-10-21 Fetch.ai: An Architecture for Modern Multi-Agent Systems Michael J. Wooldridge et.al. 2510.18699 null
2025-10-21 Tokencake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications Zhuohang Bian et.al. 2510.18586 null
2025-10-21 WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality Chunyang Li et.al. 2510.18560 null
2025-10-21 SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation Yuncheng Hua et.al. 2510.18551 null
2025-10-21 JAUNT: Joint Alignment of User Intent and Network State for QoE-centric LLM Tool Routing Enhan Li et.al. 2510.18550 null
2025-10-21 EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval Zebin Yang et.al. 2510.18546 null
2025-10-21 Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models Sureyya Akin et.al. 2510.18515 null
2025-10-21 Crucible: Quantifying the Potential of Control Algorithms through LLM Agents Lianchen Jia et.al. 2510.18491 null
2025-10-21 LAFA: Agentic LLM-Driven Federated Analytics over Decentralized Data Sources Haichao Ji et.al. 2510.18477 null
2025-10-21 Probabilistic Modeling of Intentions in Socially Intelligent LLM Agents Feifan Xia et.al. 2510.18476 null
2025-10-21 Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents Guangfu Guo et.al. 2510.18424 null
2025-10-21 Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy Games Runnan Qi et.al. 2510.18395 null
2025-10-21 MENTOR: A Reinforcement Learning Framework for Model Enhancement via Teacher-Optimized Rewards in Small Models ChangSu Choi et.al. 2510.18383 null
2025-10-21 InspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger Collaboration Yunkun Wang et.al. 2510.18327 null
2025-10-21 Earth AI: Unlocking Geospatial Insights with Foundation Models and Cross-Modal Reasoning Aaron Bell et.al. 2510.18318 null
2025-10-21 Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming Zheng Zhang et.al. 2510.18314 null
2025-10-21 Food4All: A Multi-Agent Framework for Real-time Free Food Discovery with Integrated Nutritional Metadata Zhengqing Yuan et.al. 2510.18289 null
2025-10-21 Optimal allocations with distortion risk measures and mixed risk attitudes Mario Ghossoub et.al. 2510.18236 null
2025-10-21 Applying voxel-based analysis to oropharyngeal cancer proton therapy patients: a correlation study on radiation-induced acute dysphagia Qianxia Wang et.al. 2510.18210 null
2025-10-21 Adaptive Coopetition: Leveraging Coarse Verifier Signals for Resilient Multi-Agent LLM Reasoning Rui Jerry Huang et.al. 2510.18179 null
2025-10-21 NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly? Jierui Peng et.al. 2510.16263 null
2025-10-21 SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection Yang Feng et.al. 2510.16219 null
2025-10-21 PokeeResearch: Effective Deep Research via Reinforcement Learning from AI Feedback and Robust Reasoning Scaffold Yi Wan et.al. 2510.15862 null
2025-10-21 FinAI Data Assistant: LLM-based Financial Database Query Processing with the OpenAI Function Calling API Juhyeong Kim et.al. 2510.14162 null
2025-10-21 A $^2$ FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning Qianben Chen et.al. 2510.12838 null
2025-10-20 AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI Manik Rana et.al. 2510.18170 null
2025-10-20 World-in-World: World Models in a Closed-Loop World Jiahan Zhang et.al. 2510.18135 null
2025-10-20 SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving Xiangbo Gao et.al. 2510.18123 null
2025-10-20 Investigating the Impact of Dark Patterns on LLM-Based Web Agents Devin Ersoy et.al. 2510.18113 null
2025-10-20 Does Reasoning Help LLM Agents Play Dungeons and Dragons? A Prompt Engineering Experiment Patricia Delafuente et.al. 2510.18112 null
2025-10-20 CompactPrompt: A Unified Pipeline for Prompt Data Compression in LLM Workflows Joong Ho Choi et.al. 2510.18043 null
2025-10-20 OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning Zhenyu Bi et.al. 2510.18032 null
2025-10-20 FABRIC: Framework for Agent-Based Realistic Intelligence Creation Abhigya Verma et.al. 2510.17995 null
2025-10-20 PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits Neeladri Bhuiya et.al. 2510.17947 null
2025-10-20 Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics Akshara Prabhakar et.al. 2510.17797 null
2025-10-20 Executable Knowledge Graphs for Replicating AI Research Yujie Luo et.al. 2510.17795 null
2025-10-20 A Mimamsa Inspired Framework For Instruction Sequencing In AI Agents Bama Srinivasan et.al. 2510.17691 null
2025-10-20 ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling Shuyuan Zhang et.al. 2510.17603 null
2025-10-20 MIRAGE: Agentic Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning Mir Nafis Sharear Shopnil et.al. 2510.17590 null
2025-10-20 Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs Francesco Balassone et.al. 2510.17521 null
2025-10-20 Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents Yihong Tang et.al. 2510.17491 null
2025-10-20 Agentic Reinforcement Learning for Search is Unsafe Yushi Yang et.al. 2510.17431 null
2025-10-20 Diverse Planning with Simulators via Linear Temporal Logic Mustafa F. Abdelwahed et.al. 2510.17418 null
2025-10-20 Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems Rishi Jha et.al. 2510.17276 null
2025-10-20 Coinvisor: An RL-Enhanced Chatbot Agent for Interactive Cryptocurrency Investment Analysis Chong Chen et.al. 2510.17235 null
2025-10-20 ALPINE: A Lightweight and Adaptive Privacy-Decision Agent Framework for Dynamic Edge Crowdsensing Guanjie Cheng et.al. 2510.17162 null
2025-10-20 Decentralized Real-Time Planning for Multi-UAV Cooperative Manipulation via Imitation Learning Shantnav Agarwal et.al. 2510.17143 null
2025-10-20 Do LLMs Recognize Your Latent Preferences? A Benchmark for Latent Information Discovery in Personalized Interaction Ioannis Tsaknakis et.al. 2510.17132 null
2025-10-20 Semantic Intelligence: A Bio-Inspired Cognitive Framework for Embodied Agents Wenbing Tang et.al. 2510.17129 null
2025-10-20 Verification-Aware Planning for Multi-Agent Systems Tianyang Xu et.al. 2510.17109 null
2025-10-20 Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models Elias Hossain et.al. 2510.17098 null
2025-10-20 A Brain Cell Type Resource Created by Large Language Models and a Multi-Agent AI System for Collaborative Community Annotation Rongbin Li et.al. 2510.17064 null
2025-10-20 Consistent Zero-Shot Imitation with Contrastive Goal Inference Kathryn Wantlin et.al. 2510.17059 null
2025-10-20 Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks Trilok Padhi et.al. 2510.14207 null
2025-10-19 ToolCritic: Detecting and Correcting Tool-Use Errors in Dialogue Systems Hassan Hamad et.al. 2510.17052 null
2025-10-19 ReclAIm: A multi-agent framework for degradation-aware performance tuning of medical imaging AI Eleftherios Tzanis et.al. 2510.17004 null
2025-10-19 EEschematic: Multimodal-LLM Based AI Agent for Schematic Generation of Analog Circuit Chang Liu et.al. 2510.17002 null
2025-10-19 STARK: Strategic Team of Agents for Refining Kernels Juncheng Dong et.al. 2510.16996 null
2025-10-19 Towards Interpretable and Trustworthy Time Series Reasoning: A BlueSky Vision Kanghui Ning et.al. 2510.16980 null
2025-10-19 Lark: Biologically Inspired Neuroevolution for Multi-Stakeholder LLM Agents Dheeraj Chintapalli et.al. 2510.16978 null
2025-10-19 Learning Ecology with VERA Using Conceptual Models and Simulations Spencer Rugaber et.al. 2510.16944 null
2025-10-19 VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents Kangrui Wang et.al. 2510.16907 null
2025-10-19 Agentic Inequality Matthew Sharp et.al. 2510.16853 null
2025-10-19 FinSight: Towards Real-World Financial Deep Research Jiajie Jin et.al. 2510.16844 null
2025-10-19 More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents Pengfei Gao et.al. 2510.16786 null
2025-10-19 Beyond Pipelines: A Survey of the Paradigm Shift toward Model-Native Agentic AI Jitao Sang et.al. 2510.16720 null
2025-10-19 An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems Ni Zhang et.al. 2510.16701 null
2025-10-19 Pursuing Minimal Sufficiency in Spatial Reasoning Yejie Guo et.al. 2510.16688 null
2025-10-19 Agentic Design of Compositional Machines Wenqian Zhang et.al. 2510.14980 null
2025-10-18 Unleashing Diverse Thinking Modes in LLMs through Multi-Agent Collaboration Zhixuan He et.al. 2510.16645 null
2025-10-18 Prompt Optimization via Retrieved Reasoning Assets and Multi-Agent Analysis Wonduk Seo et.al. 2510.16635 null
2025-10-18 Prior Makes It Possible: From Sublinear Graph Algorithms to LLM Test-Time Methods Avrim Blum et.al. 2510.16609 null
2025-10-18 Ripple Effect Protocol: Coordinating Agent Populations Ayush Chopra et.al. 2510.16572 null
2025-10-18 BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction Tian Xia et.al. 2510.16559 null
2025-10-18 Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety Vamshi Krishna Bonagiri et.al. 2510.16492 null
2025-10-18 REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting Changyue Shi et.al. 2510.16410 null
2025-10-18 ATA: A Neuro-Symbolic Approach to Implement Autonomous and Trustworthy Agents David Peer et.al. 2510.16381 null
2025-10-18 Synergizing chemical and AI communities for advancing laboratories of the future Saejin Oh et.al. 2510.16293 null
2025-10-17 Outraged AI: Large language models prioritise emotion over cost in fairness enforcement Hao Liu et.al. 2510.17880 null
2025-10-17 WEBSERV: A Browser-Server Environment for Efficient Training of Reinforcement Learning-based Web Agents at Scale Yuxuan Lu et.al. 2510.16252 null
2025-10-17 Towards Automatic Evaluation and Selection of PHI De-identification Models via Multi-Agent Collaboration Guanchen Wu et.al. 2510.16194 null
2025-10-17 Agentic AI for Ultra-Modern Networks: Multi-Agent Framework for RAN Autonomy and Assurance Sukhdeep Singh et.al. 2510.16144 null
2025-10-17 Narrowing Action Choices with AI Improves Human Sequential Decisions Eleni Straitouri et.al. 2510.16097 null
2025-10-17 TriAgent: Automated Biomarker Discovery with Deep Research Grounding for Triage in Acute Care by LLM-Based Multi-Agent Collaboration Kerem Delikoyun et.al. 2510.16080 null
2025-10-17 EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle Rong Wu et.al. 2510.16079 null
2025-10-17 SIADAFIX: issue description response for adaptive program repair Xin Cao et.al. 2510.16059 null
2025-10-17 PolySkill: Learning Generalizable Skills Through Polymorphic Abstraction Simon Yu et.al. 2510.15863 null
2025-10-17 Self-evolving expertise in complex non-verifiable subject domains: dialogue as implicit meta-RL Richard M. Bailey et.al. 2510.15772 null
2025-10-17 AURA: An Agent Autonomy Risk Assessment Framework Lorenzo Satta Chiris et.al. 2510.15739 null
2025-10-17 Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation Ed Li et.al. 2510.15624 null
2025-10-17 The Spark Effect: On Engineering Creative Diversity in Multi-Agent AI Systems Alexander Doudkin et.al. 2510.15568 null
2025-10-17 MARS: Reinforcing Multi-Agent Reasoning of LLMs through Self-Play in Strategic Games Huining Yuan et.al. 2510.15414 null
2025-10-17 SHARE: Scene-Human Aligned Reconstruction Joshua Li et.al. 2510.15342 null
2025-10-17 VERA-MH Concept Paper Luca Belli et.al. 2510.15297 null
2025-10-17 Exemplar-Guided Planing: Enhanced LLM Agent for KGQA Jingao Xu et.al. 2510.15283 null
2025-10-17 Experience-Driven Exploration for Efficient API-Free AI Agents Chenwei Tang et.al. 2510.15259 null
2025-10-17 Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions Xi Wang et.al. 2510.15258 null
2025-10-17 Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding Sensen Gao et.al. 2510.15253 null
2025-10-17 Where to Search: Measure the Prior-Structured Search Space of LLM Agents Zhuo-Yang Song et.al. 2510.14846 null
2025-10-16 GUIrilla: A Scalable Framework for Automated Desktop UI Exploration Sofiya Garkot et.al. 2510.16051 null
2025-10-16 MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation Gurusha Juneja et.al. 2510.15186 null
2025-10-16 Internalizing World Models via Self-Play Finetuning for Agentic RL Shiqi Chen et.al. 2510.15047 null
2025-10-16 Generalized Dynamics Generation towards Scannable Physical World Model Yichen Li et.al. 2510.15041 null
2025-10-16 UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos Mingxuan Liu et.al. 2510.15018 null
2025-10-16 Data-driven Calibration Sample Selection and Forecast Combination in Electricity Price Forecasting: An Application of the ARHNN Method Tomasz Serafin et.al. 2510.15011 null
2025-10-16 Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents Guoqing Wang et.al. 2510.14967 null
2025-10-16 VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation Han Zhao et.al. 2510.14902 null
2025-10-16 The Gatekeeper Knows Enough Fikresilase Wondmeneh Abebayew et.al. 2510.14881 null
2025-10-16 LabOS: The AI-XR Co-Scientist That Sees and Works With Humans Le Cong et.al. 2510.14861 null
2025-10-16 RoboGPT-R1: Enhancing Robot Planning with Reinforcement Learning Jinrui Liu et.al. 2510.14828 null
2025-10-16 To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models Eran Malach et.al. 2510.14826 null
2025-10-16 ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling Jianghao Lin et.al. 2510.14703 null
2025-10-16 LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet? Bin Liu et.al. 2510.14700 null
2025-10-16 LLM Agents Beyond Utility: An Open-Ended Perspective Asen Nachkov et.al. 2510.14548 null
2025-10-16 Agentic Entropy-Balanced Policy Optimization Guanting Dong et.al. 2510.14545 null
2025-10-16 Helmsman: Autonomous Synthesis of Federated Learning Systems via Multi-Agent Collaboration Haoyuan Li et.al. 2510.14512 null
2025-10-16 LiRA: Linguistic Robust Anchoring for Cross-lingual Large Language Models Haolin Li et.al. 2510.14466 null
2025-10-16 Towards Automated Governance: A DSL for Human-Agent Collaboration in Software Projects Adem Ait et.al. 2510.14465 null
2025-10-16 Why Instant-Runoff Voting Is So Resilient to Coalitional Manipulation: Phase Transitions in the Perturbed Culture François Durand et.al. 2510.14450 null
2025-10-16 Explore to Evolve: Scaling Evolved Aggregation Logic via Proactive Online Exploration for Deep Research Agents Rui Wang et.al. 2510.14438 null
2025-10-16 Bounds and asymptotic expansions for the radii of convexity and uniform convexity of normalized Bessel functions Árpád Baricz et.al. 2510.14323 null
2025-10-16 Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies Mason Nakamura et.al. 2510.14312 null
2025-10-16 ReUseIt: Synthesizing Reusable AI Agent Workflows for Web Automation Yimeng Liu et.al. 2510.14308 null
2025-10-16 AlphaQuanter: An End-to-End Tool-Orchestrated Agentic Reinforcement Learning Framework for Stock Trading Zheye Deng et.al. 2510.14264 null
2025-10-16 MAFA: A Multi-Agent Framework for Enterprise-Scale Annotation with Configurable Task Adaptation Mahmood Hegazy et.al. 2510.14184 null
2025-10-16 Training LLM Agents to Empower Humans Evan Ellis et.al. 2510.13709 null
2025-10-16 OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies Peng Di et.al. 2510.13561 null
2025-10-16 SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding Tanveer Hannan et.al. 2510.13016 null
2025-10-16 Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics Marco Del Tredici et.al. 2510.12787 null
2025-10-16 Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning Xingang Guo et.al. 2510.12712 null
2025-10-16 MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites Zhenxin Lei et.al. 2510.12126 null
2025-10-15 When “Correct” Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents? Yibo Peng et.al. 2510.17862 null
2025-10-15 CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation Yee Man Choi et.al. 2510.17853 null
2025-10-15 CodeEvolve: An open source evolutionary coding agent for algorithm discovery and optimization Henrique Assumpção et.al. 2510.14150 null
2025-10-15 Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems Edoardo Allegrini et.al. 2510.14133 null
2025-10-15 Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving Nikos Pagonas et.al. 2510.14126 null
2025-10-15 STEMS: Spatial-Temporal Enhanced Safe Multi-Agent Coordination for Building Energy Management Huiliang Zhang et.al. 2510.14112 null
2025-10-15 Three-Dimensional Simulation of the University of Hawai`i FEL Oscillator: Superradiant Emission and Cavity Desynchronization Amir Weinberg et.al. 2510.14061 null
2025-10-15 Sequential Quantum Measurements and the Instrumental Group Algebra Christopher S. Jackson et.al. 2510.13980 null
2025-10-15 An LLM-Powered AI Agent Framework for Holistic IoT Traffic Interpretation Daniel Adu Worae et.al. 2510.13925 null
2025-10-15 FACTS: Table Summarization via Offline Template Generation with Agentic Workflows Ye Yuan et.al. 2510.13920 null
2025-10-15 Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms Shrey Pandit et.al. 2510.13913 null
2025-10-15 RECODE: Reasoning Through Code Generation for Visual Question Answering Junhong Shen et.al. 2510.13756 null
2025-10-15 From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails Ravi Pandya et.al. 2510.13727 null
2025-10-15 Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module Ruitao Feng et.al. 2510.13558 null
2025-10-15 Tandem Training for Language Models Robert West et.al. 2510.13551 null
2025-10-15 In-Browser LLM-Guided Fuzzing for Real-Time Prompt Injection Testing in Agentic AI Browsers Avihay Cohen et.al. 2510.13543 null
2025-10-15 MADREC: A Multi-Aspect Driven LLM Agent for Explainable and Adaptive Recommendation Jiin Park et.al. 2510.13371 null
2025-10-15 Higher Satisfaction, Lower Cost: A Technical Report on How LLMs Revolutionize Meituan’s Intelligent Interaction Systems Xuxin Cheng et.al. 2510.13291 null
2025-10-15 Automated Network Protocol Testing with LLM Agents Yunze Wei et.al. 2510.13248 null
2025-10-15 EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems Yufei He et.al. 2510.13220 null
2025-10-15 Addressing the alignment problem in transportation policy making: an LLM approach Xiaoyu Yan et.al. 2510.13139 null
2025-10-14 Using Kolmogorov-Smirnov Distance for Measuring Distribution Shift in Machine Learning Ozan K. Tonguz et.al. 2510.15996 null
2025-10-14 MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents Dongsen Zhang et.al. 2510.15994 null
2025-10-14 Benefits and Limitations of Communication in Multi-Agent Reasoning Michael Rizvi-Martel et.al. 2510.13903 null
2025-10-14 GenCellAgent: Generalizable, Training-Free Cellular Image Segmentation via Large Language Model Agents Xi Yu et.al. 2510.13896 null
2025-10-14 MultiFoodhat: A potential new paradigm for intelligent food quality inspection Yue Hu et.al. 2510.13889 null
2025-10-14 Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments Crystal Qian et.al. 2510.13011 null
2025-10-14 SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of LLM-based Embodied Agents Simon Sinong Zhan et.al. 2510.12985 null
2025-10-14 From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models Imran Khan et.al. 2510.12864 null
2025-10-14 Three Lenses on the AI Revolution: Risk, Transformation, Continuity Masoud Makrehchi et.al. 2510.12859 null
2025-10-14 VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage A. Alfarano et.al. 2510.12750 null
2025-10-14 SPORTS: Simultaneous Panoptic Odometry, Rendering, Tracking and Segmentation for Urban Scenes Understanding Zhiliu Yang et.al. 2510.12749 null
2025-10-14 Multi-Agent Debate for LLM Judges with Adaptive Stability Detection Tianyu Hu et.al. 2510.12697 null
2025-10-14 ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning Hanyang Chen et.al. 2510.12693 null
2025-10-14 Designing Tools with Control Confidence Ajith Anil Meera et.al. 2510.12630 null
2025-10-14 A Survey of Vibe Coding with Large Language Models Yuyao Ge et.al. 2510.12399 null
2025-10-14 GOAT: A Training Framework for Goal-Oriented Agent with Tools Hyunji Min et.al. 2510.12218 null
2025-10-14 Agent-Based Simulation of a Financial Market with Large Language Models Ryuji Hashimoto et.al. 2510.12189 null
2025-10-14 IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation Wenxu Zhou et.al. 2510.12095 null
2025-10-14 ToPolyAgent: AI Agents for Coarse-Grained Topological Polymer Simulations Lijie Ding et.al. 2510.12091 null
2025-10-14 Evaluating the Quality of Randomness and Entropy in Tasks Supported by Large Language Models Rabimba Karanjai et.al. 2510.12080 null
2025-10-14 EmboMatrix: A Scalable Training-Ground for Embodied Decision-Making Zixing Lei et.al. 2510.12072 null
2025-10-14 AI Agents as Universal Task Solvers Alessandro Achille et.al. 2510.12066 null
2025-10-14 Empowering LLM Agents with Geospatial Awareness: Toward Grounded Reasoning for Wildfire Response Yiheng Chen et.al. 2510.12061 null
2025-10-14 On the Number of Small Points for Rational Maps Jit Wu Yap et.al. 2510.12039 null
2025-10-14 ManiAgent: An Agentic Framework for General Robotic Manipulation Yi Yang et.al. 2510.11660 null
2025-10-14 Stronger Together: On-Policy Reinforcement Learning for Collaborative LLMs Yujie Zhao et.al. 2510.11062 null
2025-10-13 Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation Sayash Kapoor et.al. 2510.11977 null
2025-10-13 Scaling Long-Horizon LLM Agent via Context-Folding Weiwei Sun et.al. 2510.11967 null
2025-10-13 DMAS-Forge: A Framework for Transparent Deployment of AI Applications as Distributed Systems Alessandro Cornacchia et.al. 2510.11872 null
2025-10-13 Demystifying Reinforcement Learning in Agentic Reasoning Zhaochen Yu et.al. 2510.11701 null
2025-10-13 When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents Lingfei Qian et.al. 2510.11695 null
2025-10-13 Chronologically Consistent Generative AI Songrun He et.al. 2510.11677 null
2025-10-13 FinVet: A Collaborative Framework of RAG and External Fact-Checking Agents for Financial Misinformation Detection Daniel Berhane Araya et.al. 2510.11654 null
2025-10-13 Analyzing and Internalizing Complex Policy Documents for LLM Agents Jiateng Liu et.al. 2510.11588 null
2025-10-13 Uncertainty-Aware, Risk-Adaptive Access Control for Agentic Systems using an LLM-Judged TBAC Model Charles Fleming et.al. 2510.11414 null
2025-10-13 DocReward: A Document Reward Model for Structuring and Stylizing Junpeng Liu et.al. 2510.11391 null
2025-10-13 Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamics Sheng Jin et.al. 2510.11290 null
2025-10-13 PADME: Procedure Aware DynaMic Execution Deepeka Garg et.al. 2510.11281 null
2025-10-13 A Large-Language-Model Assisted Automated Scale Bar Detection and Extraction Framework for Scanning Electron Microscopic Images Yuxuan Chen et.al. 2510.11260 null
2025-10-13 Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems Pengyu Zhu et.al. 2510.11246 null
2025-10-13 Attacks by Content: Automated Fact-checking is an AI Security Issue Michael Schlichtkrull et.al. 2510.11238 null
2025-10-13 WebRouter: Query-specific Router via Variational Information Bottleneck for Cost-sensitive Web Agent Tao Li et.al. 2510.11221 null
2025-10-13 Can Tool-Integrated Reinforcement Learning Generalize Across Diverse Domains? Zhengyu Chen et.al. 2510.11184 null
2025-10-13 $How^{2}$ : How to learn from procedural How-to questions Gautier Dagan et.al. 2510.11144 null
2025-10-13 video-SALMONN S: Streaming Audio-Visual LLMs Beyond Length Limits via Memory Guangzhi Sun et.al. 2510.11129 null
2025-10-13 SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents Longjie Guo et.al. 2510.11035 null
2025-10-13 A Survey on Agentic Multimodal Large Language Models Huanjin Yao et.al. 2510.10991 null
2025-10-13 Rethinking Reward Miscalibration of GRPO in Agentic RL Jingyu Liu et.al. 2509.23870 null
2025-10-13 EvoEmo: Towards Evolved Emotional Policies for Adversarial LLM Agents in Multi-Turn Price Negotiation Yunbo Long et.al. 2509.04310 null
2025-10-12 Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment Planning Dongrong Yang et.al. 2510.11754 null
2025-10-12 GraphTracer: Graph-Guided Failure Tracing in LLM Agents for Robust Multi-Turn Deep Search Heng Zhang et.al. 2510.10581 null
2025-10-12 MedCoAct: Confidence-Aware Multi-Agent Collaboration for Complete Clinical Decision Hongjie Zheng et.al. 2510.10461 null
2025-10-12 Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval Junwei Lan et.al. 2509.24869 null
2025-10-12 Talk Less, Call Right: Enhancing Role-Play LLM Agents with Automatic Prompt Optimization and Role Prompting Saksorn Ruangtanusak et.al. 2509.00482 null
2025-10-11 KG-MAS: Knowledge Graph-Enhanced Multi-Agent Infrastructure for coupling physical and digital robotic environments Walid Abdela et.al. 2510.10325 null
2025-10-11 Simulating Viva Voce Examinations to Evaluate Clinical Reasoning in Large Language Models Christopher Chiu et.al. 2510.10278 null
2025-10-11 Don’t Just Fine-tune the Agent, Tune the Environment Siyuan Lu et.al. 2510.10197 null
2025-10-11 ALLOY: Generating Reusable Agent Workflows from User Demonstration Jiawen Li et.al. 2510.10049 null
2025-10-11 SwarmSys: Decentralized Swarm-Inspired Agents for Scalable and Adaptive Reasoning Ruohao Li et.al. 2510.10047 null
2025-10-11 Leveraging Large Language Models for Cybersecurity Risk Assessment – A Case from Forestry Cyber-Physical Systems Fikret Mert Gultekin et.al. 2510.06343 null
2025-10-11 Tree Search for LLM Agent Reinforcement Learning Yuxiang Ji et.al. 2509.21240 null
2025-10-11 ASTREA: Introducing Agentic Intelligence for Orbital Thermal Autonomy Alejandro D. Mousist et.al. 2509.13380 null
2025-10-10 Autonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics Lianhao Zhou et.al. 2510.09901 null
2025-10-10 How can we assess human-agent interactions? Case studies in software agent design Valerie Chen et.al. 2510.09801 null
2025-10-10 Building a Foundational Guardrail for General Agentic Systems via Synthetic Data Yue Huang et.al. 2510.09781 null
2025-10-10 Preference-Aware Memory Update for Long-Term LLM Agents Haoran Sun et.al. 2510.09720 null
2025-10-10 StreamingVLM: Real-Time Understanding for Infinite Video Streams Ruyi Xu et.al. 2510.09608 null
2025-10-10 Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols Mikhail Terekhov et.al. 2510.09462 null
2025-10-10 Safety Game: Balancing Safe and Informative Conversations with Blackbox Agentic AI using LP Solvers Tuan Nguyen et.al. 2510.09330 null
2025-10-10 Fundamentals of Building Autonomous LLM Agents Victor de Lamo Castrillo et.al. 2510.09244 null
2025-10-10 Leading the Follower: Learning Persuasive Agents in Social Deduction Games Zhang Zheng et.al. 2510.09087 null
2025-10-10 When LLM Agents Meet Graph Optimization: An Automated Data Quality Improvement Approach Zhihan Zhang et.al. 2510.08952 null
2025-10-10 Reimagining Agent-based Modeling with Large Language Model Agents via Shachi So Kuroki et.al. 2509.21862 null
2025-10-09 CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization Debeshee Das et.al. 2510.08829 null
2025-10-09 COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context Guangya Wan et.al. 2510.08790 null
2025-10-09 Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools Ha Min Son et.al. 2510.08640 null
2025-10-09 CaRT: Teaching LLM Agents to Know When They Know Enough Grace Liu et.al. 2510.08517 null
2025-10-09 Opponent Shaping in LLM Agents Marta Emili Garcia Segura et.al. 2510.08255 null
2025-10-09 Simulating Teams with LLM Agents: Interactive 2D Environments for Studying Human-AI Dynamics Mohammed Almutairi et.al. 2510.08242 null
2025-10-09 Training-Free Group Relative Policy Optimization Yuzheng Cai et.al. 2510.08191 null
2025-10-09 AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment Xiaochong Lan et.al. 2510.08081 null
2025-10-09 Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks Cheng Yang et.al. 2510.08002 null
2025-10-09 Team Xiaomi EV-AD VLA: Learning to Navigate Socially Through Proactive Risk Perception – Technical Report for IROS 2025 RoboSense Challenge Social Navigation Track Erjia Xiao et.al. 2510.07871 null
2025-10-09 Self-Improving LLM Agents at Test-Time Emre Can Acikgoz et.al. 2510.07841 null
2025-10-09 Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models Eric Hanchen Jiang et.al. 2510.07799 null
2025-10-09 Neuro-Symbolic Agents with Modal Logic for Autonomous Diagnostics Antonin Sulc et.al. 2509.11943 null
2025-10-08 PARSE: LLM Driven Schema Optimization for Reliable Entity Extraction Anubhav Shrimal et.al. 2510.08623 null
2025-10-08 L2M-AID: Autonomous Cyber-Physical Defense by Fusing Semantic Reasoning of Large Language Models with Multi-Agent Reinforcement Learning (Preprint) Tianxiang Xu et.al. 2510.07363 null
2025-10-08 LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document Understanding Zhivar Sourati et.al. 2510.07233 null
2025-10-08 Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping Ziyi Wang et.al. 2510.07230 null
2025-10-08 Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions Yixiang Zhang et.al. 2510.07176 null
2025-10-08 NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents Tianshi Zheng et.al. 2510.07172 null
2025-10-08 Prompt Optimization Across Multiple Agents for Representing Diverse Human Populations Manh Hung Nguyen et.al. 2510.07064 null
2025-10-08 COMPASS: A Multi-Turn Benchmark for Tool-Mediated Planning & Preference Optimization Tian Qin et.al. 2510.07043 null
2025-10-08 LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling Zecheng Tang et.al. 2510.06915 null
2025-10-08 When Machines Meet Each Other: Network Effects and the Strategic Role of History in Multi-Agent AI Yu Liu et.al. 2510.06903 null
2025-10-08 SID: Multi-LLM Debate Driven by Self Signals Xuhang Chen et.al. 2510.06843 null
2025-10-08 Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management Miao Lu et.al. 2510.06727 null
2025-10-08 WebDART: Dynamic Decomposition and Re-planning for Complex Web Tasks Jingbo Yang et.al. 2510.06587 null
2025-10-08 Spiral of Silence in Large Language Model Agents Mingze Zhong et.al. 2510.02360 null
2025-10-08 Toward Causal-Visual Programming: Enhancing Agentic Reasoning in Low-Code Environments Jiexi Xu et.al. 2509.25282 null
2025-10-07 A Survey on Agentic Security: Applications, Threats and Defenses Asif Shahriar et.al. 2510.06445 null
2025-10-07 Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents Mingkang Zhu et.al. 2510.06214 null
2025-10-07 RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback Chunyu Miao et.al. 2510.06186 null
2025-10-07 LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams Aju Ani Justus et.al. 2510.06151 null
2025-10-07 Constraint-Aware Route Recommendation from Natural Language via Hierarchical LLM Agents Tao Zhe et.al. 2510.06078 null
2025-10-07 Training-Free Time Series Classification via In-Context Reasoning with LLM Agents Songyuan Sui et.al. 2510.05950 null
2025-10-07 EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models Zheyue Tan et.al. 2510.05943 null
2025-10-07 LLM-FS-Agent: A Deliberative Role-based Large Language Model Architecture for Transparent Feature Selection Mohamed Bal-Ghaoui et.al. 2510.05935 null
2025-10-07 Communication Enables Cooperation in LLM Agents: A Comparison with Curriculum-Based Approaches Hachem Madmoun et.al. 2510.05748 null
2025-10-07 AutoPentester: An LLM Agent-based Framework for Automated Pentesting Yasod Ginige et.al. 2510.05605 null
2025-10-07 AgentDR Dynamic Recommendation with Implicit Item-Item Relations via LLM-based Agents Mingdai Yang et.al. 2510.05598 null
2025-10-07 From Agentification to Self-Evolving Agentic AI for Wireless Networks: Concepts, Approaches, and Future Research Directions Changyuan Zhao et.al. 2510.05596 null
2025-10-07 BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks Sagnik Anupam et.al. 2510.02418 null
2025-10-06 Adversarial Reinforcement Learning for Large Language Model Agent Safety Zizhao Wang et.al. 2510.05442 null
2025-10-06 A Lightweight Large Language Model-Based Multi-Agent System for 2D Frame Structural Analysis Ziheng Geng et.al. 2510.05414 null
2025-10-06 Plug-and-Play Dramaturge: A Divide-and-Conquer Approach for Iterative Narrative Script Refinement via Collaborative LLM Agents Wenda Xie et.al. 2510.05188 null
2025-10-06 RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection Yuxin Wen et.al. 2510.04885 null
2025-10-06 Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails Siwei Han et.al. 2510.04860 null
2025-10-06 Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents Yiding Wang et.al. 2510.04695 null
2025-10-06 Multi-Agent Tool-Integrated Policy Optimization Zhanfeng Mo et.al. 2510.04678 null
2025-10-06 Social Agent: Mastering Dyadic Nonverbal Behavior Generation via Conversational LLM Agents Zeyi Zhang et.al. 2510.04637 null
2025-10-06 Autonomy Matters: A Study on Personalization-Privacy Dilemma in LLM Agents Zhiping Zhang et.al. 2510.04465 null
2025-10-06 Beyond Manuals and Tasks: Instance-Level Context Learning for LLM Agents Kuntai Cai et.al. 2510.02369 null
2025-10-05 Internal World Models as Imagination Networks in Cognitive Agents Saurabh Ranjan et.al. 2510.04391 null
2025-10-05 Just-in-time Episodic Feedback Hinter: Leveraging Offline Knowledge to Improve LLM Agents Adaptation Hadi Nekoei et.al. 2510.04373 null
2025-10-05 Closing the Loop: Coordinating Inventory and Recommendation via Deep Reinforcement Learning on Multiple Timescales Jinyang Jiang et.al. 2510.04272 null
2025-10-05 AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework Hanchen Zhang et.al. 2510.04206 null
2025-10-05 Constructing coherent spatial memory in LLM agents through graph rectification Puzhen Zhang et.al. 2510.04195 null
2025-10-05 From Shadow to Light: Toward Safe and Efficient Policy Learning Across MPC, DeePC, RL, and LLM Agents Amin Vahidi-Moghaddam et.al. 2510.04076 null
2025-10-04 Adversarial Agent Collaboration for C to Rust Translation Tianyu Li et.al. 2510.03879 null
2025-10-04 InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents Yaxin Du et.al. 2510.02271 null
2025-10-04 Extracting Conceptual Knowledge to Locate Software Issues Ying Wang et.al. 2509.21427 null
2025-10-03 VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation Lesly Miculicich et.al. 2510.05156 null
2025-10-03 LLM Agents for Automated Dependency Upgrades Vali Tawosi et.al. 2510.03480 null
2025-10-03 ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework Vali Tawosi et.al. 2510.03463 null
2025-10-03 Improving GUI Grounding with Explicit Position-to-Coordinate Mapping Suyuchen Wang et.al. 2510.03230 null
2025-10-03 CoDA: Agentic Systems for Collaborative Data Visualization Zichen Chen et.al. 2510.03194 null
2025-10-03 AudioToolAgent: An Agentic Framework for Audio-Language Models Gijs Wijngaard et.al. 2510.02995 null
2025-10-03 Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents Wonjoong Kim et.al. 2510.02837 null
2025-10-03 AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents Arman Zharmagambetov et.al. 2503.09780 null
2025-10-02 AgentCaster: Reasoning-Guided Tornado Forecasting Michael Chen et.al. 2510.03349 null
2025-10-02 Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge Charlie Masters et.al. 2510.02557 null
2025-10-02 StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? Yanxu Chen et.al. 2510.02209 null
2025-10-02 TACOS: Task Agnostic COordinator of a multi-drone System Alessandro Nazzari et.al. 2510.01869 null
2025-10-02 Pre-Hoc Predictions in AutoML: Leveraging LLMs to Enhance Model Selection and Benchmarking for Tabular datasets Yannis Belkhiter et.al. 2510.01842 null
2025-10-02 GuruAgents: Emulating Wise Investors with Prompt-Guided LLM Agents Yejin Kim et.al. 2510.01664 null
2025-10-02 SoK: Measuring What Matters for Closed-Loop Security Agents Mudita Khurana et.al. 2510.01654 null
2025-10-02 Position: Privacy Is Not Just Memorization! Niloofar Mireshghallah et.al. 2510.01645 null
2025-10-02 GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments Hanlin Zhu et.al. 2509.21998 null
2025-10-02 Gala: Global LLM Agents for Text-to-Model Translation Junyang Cai et.al. 2509.08970 null
2025-10-01 Automating Data-Driven Modeling and Analysis for Engineering Applications using Large Language Model Agents Yang Liu et.al. 2510.01398 null
2025-10-01 Beyond Single LLMs: Enhanced Code Generation via Multi-Stage Performance-Guided LLM Orchestration Huashan Chen et.al. 2510.01379 null
2025-10-01 Fine-tuning with RAG for Improving LLM Learning of New Skills Humaid Ibrahim et.al. 2510.01375 null
2025-10-01 Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks Shoumik Saha et.al. 2510.01359 null
2025-10-01 The Social Laboratory: A Psychometric Framework for Multi-Agent LLM Evaluation Zarreen Reza et.al. 2510.01295 null
2025-10-01 TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments Zhangchen Xu et.al. 2510.01179 null
2025-10-01 Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare Zhengliang Shi et.al. 2510.01164 null
2025-10-01 A Practitioner’s Guide to Multi-turn Agentic Reinforcement Learning Ruiyi Wang et.al. 2510.01132 null
2025-10-01 QUASAR: Quantum Assembly Code Generation Using Tool-Augmented LLMs via Agentic RL Cong Yu et.al. 2510.00967 null
2025-10-01 ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs Adi Simhi et.al. 2510.00857 null
2025-10-01 ACON: Optimizing Context Compression for Long-horizon LLM Agents Minki Kang et.al. 2510.00615 null
2025-10-01 JoyAgent-JDGenie: Technical Report on the GAIA Jiarun Liu et.al. 2510.00510 null
2025-10-01 Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual Navigation Yiyuan Pan et.al. 2510.00441 null
2025-10-01 RELATE-Sim: Leveraging Turning Point Theory and LLM Agents to Predict and Understand Long-Term Relationship Dynamics through Interactive Narrative Simulations Matthew Yue et.al. 2510.00414 null
2025-10-01 Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs Siyu Zhu et.al. 2509.25779 null
2025-10-01 Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development Yuxuan Wan et.al. 2509.25297 null
2025-10-01 Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks Aaron Xuxiang Tian et.al. 2509.23537 null
2025-10-01 On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language Sébastien Salva et.al. 2509.19136 null
2025-10-01 A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks S M Asif Hossain et.al. 2509.14285 null
2025-09-30 From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization Haoran Xi et.al. 2510.02389 null
2025-09-30 CORTEX: Collaborative LLM Agents for High-Stakes Alert Triage Bowen Wei et.al. 2510.00311 null
2025-09-30 Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents Zhen Yang et.al. 2509.26539 null
2025-09-30 VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications Wei He et.al. 2509.26490 null
2025-09-30 ErrorPrism: Reconstructing Error Propagation Paths in Cloud Service Systems Junsong Pu et.al. 2509.26463 null
2025-09-30 Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents Shuai Shao et.al. 2509.26354 null
2025-09-30 LLM Agents for Knowledge Discovery in Atomic Layer Processing Andreas Werbrouck et.al. 2509.26201 null
2025-09-30 RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning Gang Li et.al. 2509.25958 null
2025-09-30 Mem-α: Learning Memory Construction via Reinforcement Learning Yu Wang et.al. 2509.25911 null
2025-09-30 SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents Ruolin Chen et.al. 2509.25885 null
2025-09-30 Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs Hankun Dai et.al. 2509.25873 null
2025-09-30 STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents Jing-Jing Li et.al. 2509.25624 null
2025-09-30 MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning Huihao Jing et.al. 2509.24922 null
2025-09-30 TENET: Leveraging Tests Beyond Validation for Code Generation Yiran Hu et.al. 2509.24148 null
2025-09-30 Dual-Scale World Models for LLM Agents Towards Hard-Exploration Problems Minsoo Kim et.al. 2509.24116 null
2025-09-30 InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios Chenglin Yu et.al. 2509.22502 null
2025-09-30 Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents Davide Paglieri et.al. 2509.03581 null
2025-09-30 Towards Agentic OS: An LLM Agent Framework for Linux Schedulers Yusheng Zheng et.al. 2509.01245 null
2025-09-29 A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory Qianshan Wei et.al. 2510.02373 null
2025-09-29 Causal Autoencoder-like Generation of Feedback Fuzzy Cognitive Maps with an LLM Agent Akash Kumar Panda et.al. 2509.25593 null
2025-09-29 RadOnc-GPT: An Autonomous LLM Agent for Real-Time Patient Outcomes Labeling at Scale Jason Holmes et.al. 2509.25540 null
2025-09-29 Where LLM Agents Fail and How They can Learn From Failures Kunlun Zhu et.al. 2509.25370 null
2025-09-29 Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents Boxuan Zhang et.al. 2509.25302 null
2025-09-29 PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion Yuyang Yin et.al. 2509.24997 null
2025-09-29 When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training Sanxing Chen et.al. 2509.24923 null
2025-09-29 MAS $^2$ : Self-Generative, Self-Configuring, Self-Rectifying Multi-Agent Systems Kun Wang et.al. 2509.24323 null
2025-09-29 SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents Gyuhyeon Seo et.al. 2509.24282 null
2025-09-28 WAREX: Web Agent Reliability Evaluation on Existing Benchmarks Su Kara et.al. 2510.03285 null
2025-09-28 Optimism as Risk-Seeking in Multi-Agent Reinforcement Learning Runyu Zhang et.al. 2509.24047 null
2025-09-28 PartnerMAS: An LLM Hierarchical Multi-Agent Framework for Business Partner Selection on High-Dimensional Features Lingyao Li et.al. 2509.24046 null
2025-09-28 LLM/Agent-as-Data-Analyst: A Survey Zirui Tang et.al. 2509.23988 null
2025-09-28 Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation Pengxiang Li et.al. 2509.23866 null
2025-09-28 AgentGuard: Runtime Verification of AI Agents Roham Koohestani et.al. 2509.23864 null
2025-09-28 Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules Chenyu Zhou et.al. 2509.23836 null
2025-09-28 FedAgentBench: Towards Automating Real-world Federated Medical Image Analysis with Server-Client LLM Agents Pramit Saha et.al. 2509.23803 null
2025-09-28 GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks Cong Chen et.al. 2509.23738 null
2025-09-28 Improving the Efficiency of LLM Agent Systems through Trajectory Reduction Yuan-An Xiao et.al. 2509.23586 null
2025-09-28 Agentic Reinforcement Learning with Implicit Step Rewards Xiaoqian Liu et.al. 2509.19199 null
2025-09-27 Memory Management and Contextual Consistency for Long-Running Low-Code Agents Jiexi Xu et.al. 2509.25250 null
2025-09-27 BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software Zehua Zhang et.al. 2509.25248 null
2025-09-27 Situational Awareness for Safe and Robust Multi-Agent Interactions Under Uncertainty Benjamin Alcorn et.al. 2509.23425 null
2025-09-27 “Shall We Dig Deeper?”: Designing and Evaluating Strategies for LLM Agents to Advance Knowledge Co-Construction in Asynchronous Online Discussions Yuanhao Zhang et.al. 2509.23327 null
2025-09-27 Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents Yaorui Shi et.al. 2509.23040 null
2025-09-26 Solving the Granularity Mismatch: Hierarchical Preference Learning for Long-Horizon LLM Agents Heyang Gao et.al. 2510.03253 null
2025-09-26 AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering Ziqing Wang et.al. 2510.02328 null
2025-09-26 Infusing Theory of Mind into Socially Intelligent LLM Agents EunJeong Hwang et.al. 2509.22887 null
2025-09-26 ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents Hwan Chang et.al. 2509.22830 null
2025-09-26 EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning Wujiang Xu et.al. 2509.22576 null
2025-09-26 The Emergence of Altruism in Large-Language-Model Agents Society Haoyang Li et.al. 2509.22537 null
2025-09-26 Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents Jiaqi Shao et.al. 2509.22391 null
2025-09-26 Impact of Collective Behaviors of Autonomous Vehicles on Urban Traffic Dynamics: A Multi-Agent Reinforcement Learning Approach Ahmet Onur Akman et.al. 2509.22216 null
2025-09-26 Leveraging LLM Agents for Automated Video Game Testing Chengjia Wang et.al. 2509.22170 null
2025-09-26 CoBel-World: Harnessing LLM Reasoning to Build a Collaborative Belief World for Optimizing Embodied Multi-Agent Collaboration Zhimin Wang et.al. 2509.21981 null
2025-09-26 What Makes LLM Agent Simulations Useful for Policy? Insights From an Iterative Design Engagement in Emergency Preparedness Yuxuan Li et.al. 2509.21868 null
2025-09-26 UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios Haotian Luo et.al. 2509.21766 null
2025-09-26 JudgeAgent: Knowledge-wise and Dynamic LLM Evaluation with Agent-as-Interviewer Zhichao Shi et.al. 2509.02097 null
2025-09-25 LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants? Lu Sun et.al. 2509.21501 null
2025-09-25 What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns Stefan Szeider et.al. 2509.21224 null
2025-09-25 CORE: Full-Path Evaluation of LLM Agents Beyond Final State Panagiotis Michelakis et.al. 2509.20998 null
2025-09-25 LIMI: Less is More for Agency Yang Xiao et.al. 2509.17567 null
2025-09-24 EpidemIQs: Prompt-to-Paper LLM Agents for Epidemic Modeling and Analysis Mohammad Hossein Samaei et.al. 2510.00024 null
2025-09-24 Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models Lukas Petersson et.al. 2509.25229 null
2025-09-24 LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet? Rushil Gupta et.al. 2509.21403 null
2025-09-24 Training Task Reasoning LLM Agents for Multi-turn Task Planning via Single-turn Reinforcement Learning Hanjiang Hu et.al. 2509.20616 null
2025-09-24 SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection Yubin Ge et.al. 2509.20562 null
2025-09-24 Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation Yiren Liu et.al. 2509.20553 null
2025-09-24 Agentic Metacognition: Designing a “Self-Aware” Low-Code Agent for Failure Prediction and Human Handoff Jiexi Xu et.al. 2509.19783 null
2025-09-23 Structured Cognition for Behavioral Intelligence in Large Language Model Agents: Preliminary Study Myung Ho Kim et.al. 2510.05107 null
2025-09-23 The Heterogeneous Multi-Agent Challenge Charles Dansereau et.al. 2509.19512 null
2025-09-23 Simulating Online Social Media Conversations on Controversial Topics Using AI Agents Calibrated on Real-World Data Elisa Composta et.al. 2509.18985 null
2025-09-23 MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service Yizhe Huang et.al. 2509.18713 null
2025-09-23 LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA Zeyi Kang et.al. 2509.18576 null
2025-09-23 LLMZ+: Contextual Prompt Whitelist Principles for Agentic LLMs Tom Pawelek et.al. 2509.18557 null
2025-09-23 LLM Agents for Interactive Workflow Provenance: Reference Architecture and Evaluation Methodology Renan Souza et.al. 2509.13978 null
2025-09-22 ARK-V1: An LLM-Agent for Knowledge Graph Question Answering Requiring Commonsense Reasoning Jan-Felix Klein et.al. 2509.18063 null
2025-09-22 Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration Bingsheng Yao et.al. 2509.18008 null
2025-09-22 MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents Yuzhen Lei et.al. 2509.17628 null
2025-09-22 Human vs. Agent in Task-Oriented Conversations Zhefan Wang et.al. 2509.17619 null
2025-09-22 Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents Shouju Wang et.al. 2509.17488 null
2025-09-22 Asteria: Semantic-Aware Cross-Region Caching for Agentic LLM Tool Access Chaoyi Ruan et.al. 2509.17360 null
2025-09-22 UIPro: Unleashing Superior Interaction Capability For GUI Agents Hongxin Li et.al. 2509.17328 null
2025-09-22 Generalizable End-to-End Tool-Use RL with Synthetic CodeGym Weihua Du et.al. 2509.17325 null
2025-09-21 SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing Junlong Ke et.al. 2509.17197 null
2025-09-21 LLMs as Layout Designers: A Spatial Reasoning Perspective Sha Li et.al. 2509.16891 null
2025-09-20 Towards Transparent and Incentive-Compatible Collaboration in Decentralized LLM Multi-Agent Systems: A Blockchain-Driven Approach Minfeng Qi et.al. 2509.16736 null
2025-09-20 OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama Tianyang Xu et.al. 2509.16713 null
2025-09-20 Governed By Agents: A Survey On The Role Of Agentic AI In Future Computing Environments Nauman Ali Murad et.al. 2509.16676 null
2025-09-19 Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans Deuksin Kwon et.al. 2509.16394 null
2025-09-19 Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap Andrew Zhu et.al. 2509.16325 null
2025-09-19 Towards Robust Visual Continual Learning with Multi-Prototype Supervision Xiwei Liu et.al. 2509.16011 null
2025-09-19 How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages Siyang Wu et.al. 2509.15518 null
2025-09-19 LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring Jinhee Jang et.al. 2509.14834 null
2025-09-18 SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair Jugal Gajjar et.al. 2509.16275 null
2025-09-18 Diagnostics of cognitive failures in multi-agent expert systems using dynamic evaluation protocols and subsequent mutation of the processing context Andrejs Sorstkins et.al. 2509.15366 null
2025-09-18 A Knowledge-driven Adaptive Collaboration of LLMs for Enhancing Medical Decision-making Xiao Wu et.al. 2509.14998 null
2025-09-18 ToolSample: Dual Dynamic Sampling Methods with Curriculum Learning for RL-based Tool Learning Zihao Feng et.al. 2509.14718 null
2025-09-18 SWE-QA: Can Language Models Answer Repository-level Code Questions? Weihan Peng et.al. 2509.14635 null
2025-09-17 Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation Thales Sales Almeida et.al. 2509.14477 null
2025-09-17 TopoSizing: An LLM-aided Framework of Topology-based Understanding and Sizing for AMS Circuits Ziming Wei et.al. 2509.14169 null
2025-09-17 Understanding the Process of Human-AI Value Alignment Jack McKinlay et.al. 2509.13854 null
2025-09-17 From Legacy Fortran to Portable Kokkos: An Autonomous Agentic AI Workflow Sparsh Gupta et.al. 2509.12443 null
2025-09-17 Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives Prathamesh Vasudeo Naik et.al. 2509.08380 null
2025-09-17 Emergent Social Dynamics of LLM Agents in the El Farol Bar Problem Ryosuke Takata et.al. 2509.04537 null
2025-09-17 How Does Cognitive Bias Affect Large Language Models? A Case Study on the Anchoring Effect in Price Negotiation Simulations Yoshiki Takenami et.al. 2508.21137 null
2025-09-16 Agentic JWT: A Secure Delegation Protocol for Autonomous AI Agents Abhishek Goswami et.al. 2509.13597 null
2025-09-16 AI Agents with Human-Like Collaborative Tools: Adaptive Strategies for Enhanced Problem-Solving Harper Reed et.al. 2509.13547 null
2025-09-16 An LLM Agentic Approach for Legal-Critical Software: A Case Study for Tax Prep Software Sina Gogani-Khiabani et.al. 2509.13471 null
2025-09-16 WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning Kuan Li et.al. 2509.13305 null
2025-09-16 Agentic AI for Financial Crime Compliance Henrik Axelsen et.al. 2509.13137 null
2025-09-16 Toward PDDL Planning Copilot Yarin Benyamin et.al. 2509.12987 null
2025-09-16 H $^2$ R: Hierarchical Hindsight Reflection for Multi-Task LLM Agents Shicheng Ye et.al. 2509.12810 null
2025-09-16 Agentic Lybic: Multi-Agent Execution System with Tiered Reasoning and Orchestration Liangxuan Guo et.al. 2509.11067 null
2025-09-16 PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance Mengxiao Wang et.al. 2508.20890 null
2025-09-16 Mining the Long Tail: A Comparative Study of Data-Centric Criticality Metrics for Robust Offline Reinforcement Learning in Autonomous Motion Planning Antonio Guillen-Perez et.al. 2508.18397 null
2025-09-16 Enhancing LLM-Based Social Bot via an Adversarial Learning Framework Fanqi Kong et.al. 2508.17711 null
2025-09-15 Emotions are Recognized Patterns of Cognitive Activities Yue Jin et.al. 2509.16232 null
2025-09-15 Redefining Website Fingerprinting Attacks With Multiagent LLMs Chuxu Song et.al. 2509.12462 null
2025-09-15 Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm Alireza Mohamadi et.al. 2509.12190 null
2025-09-15 VisDocSketcher: Towards Scalable Visual Documentation with Agentic Systems Luís F. Gomes et.al. 2509.11942 null
2025-09-15 $ε$ -Optimal Multi-Agent Patrol using Recurrent Strategy Deepak Mallya et.al. 2509.11640 null
2025-09-15 Automated Creation and Enrichment Framework for Improved Invocation of Enterprise APIs as Tools Prerna Agarwal et.al. 2509.11626 null
2025-09-15 MedicalOS: An LLM Agent based Operating System for Digital Healthcare Jared Zhu et.al. 2509.11507 null
2025-09-14 Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning Anis Koubaa et.al. 2509.13352 null
2025-09-14 Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble Bingchen Wang et.al. 2509.11311 null
2025-09-14 Free-MAD: Consensus-Free Multi-Agent Debate Yu Cui et.al. 2509.11035 null
2025-09-12 FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering Gyubok Lee et.al. 2509.19319 null
2025-09-12 V-Math: An Agentic Approach to the Vietnamese National High School Graduation Mathematics Exams Duong Q. Nguyen et.al. 2509.12251 null
2025-09-12 Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight Jingyu Tang et.al. 2509.10723 null
2025-09-12 Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration Chirayu Nimonkar et.al. 2509.10656 null
2025-09-12 SciML Agents: Write the Solver, Not the Solution Saarth Gaonkar et.al. 2509.09936 null
2025-09-12 Tackling One Health Risks: How Large Language Models are leveraged for Risk Negotiation and Consensus-building Alexandra Fetsch et.al. 2509.09906 null
2025-09-12 Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining Crystal Qian et.al. 2509.09071 null
2025-09-11 TrEnv: Transparently Share Serverless Execution Environments Across Different Functions and Nodes Jialiang Huang et.al. 2509.09525 null
2025-09-11 Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning Abdel Hakim Drid et.al. 2509.09356 null
2025-09-11 Flip Co-op: Cooperative Takeovers in Shared Autonomy Sandeep Banik et.al. 2509.09281 null
2025-09-11 Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents Jiawei Wang et.al. 2509.09265 null
2025-09-11 Enabling Regulatory Multi-Agent Collaboration: Architecture, Challenges, and Solutions Qinnan Hu et.al. 2509.09215 null
2025-09-10 HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets Ying Yuan et.al. 2509.09740 null
2025-09-10 AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning Zhiheng Xi et.al. 2509.08755 null
2025-09-10 Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations Ron F. Del Rosario et.al. 2509.08646 null
2025-09-10 AutoODD: Agentic Audits via Bayesian Red Teaming in Black-Box Models Rebecca Martin et.al. 2509.08638 null
2025-09-09 Multi Robot Coordination in Highly Dynamic Environments: Tackling Asymmetric Obstacles and Limited Communication Vincenzo Suriani et.al. 2509.08859 null
2025-09-09 EnvX: Agentize Everything with Agentic AI Linyao Chen et.al. 2509.08088 null
2025-09-09 Guided Reasoning in LLM-Driven Penetration Testing Using Structured Attack Trees Katsuaki Nakano et.al. 2509.07939 null
2025-09-09 Getting In Contract with Large Language Models – An Agency Theory Perspective On Large Language Model Alignment Sascha Kaltenpoth et.al. 2509.07642 null
2025-09-09 Astra: A Multi-Agent System for GPU Kernel Performance Optimization Anjiang Wei et.al. 2509.07506 null
2025-09-09 Talking with Oompa Loompas: A novel framework for evaluating linguistic acquisition of LLM agents Sankalp Tattwadarshi Swain et.al. 2509.07389 null
2025-09-09 Autonomous Code Evolution Meets NP-Completeness Cunxi Yu et.al. 2509.07367 null
2025-09-09 CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation Alyssa Unell et.al. 2509.07325 null
2025-09-08 AxelSMOTE: An Agent-Based Oversampling Algorithm for Imbalanced Classification Sukumar Kishanthan et.al. 2509.06875 null
2025-09-08 RAFFLES: Reasoning-based Attribution of Faults for LLM Systems Chenyang Zhu et.al. 2509.06822 null
2025-09-08 Reinforcement Learning Foundations for Deep Research Systems: A Survey Wenjun Li et.al. 2509.06733 null
2025-09-08 REMI: A Novel Causal Schema Memory Architecture for Personalized Lifestyle Recommendation Agents Vishal Raman et.al. 2509.06269 null
2025-09-08 TalkToAgent: A Human-centric Explanation of Reinforcement Learning Agents with Large Language Models Haechang Kim et.al. 2509.04809 null
2025-09-08 Meta-Policy Reflexion: Reusable Reflective Memory and Rule Admissibility for Resource-Efficient LLM Agent Chunlong Wu et.al. 2509.03990 null
2025-09-07 From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets Alexander Erlei et.al. 2509.06069 null
2025-09-07 Let’s Roleplay: Examining LLM Alignment in Collaborative Dialogues Abhijnan Nath et.al. 2509.05882 null
2025-09-06 DRF: LLM-AGENT Dynamic Reputation Filtering Framework Yuwei Lou et.al. 2509.05764 null
2025-09-05 Internet 3.0: Architecture for a Web-of-Agents with it’s Algorithm for Ranking Agents Rajesh Tembarai Krishnamachari et.al. 2509.04979 null
2025-09-05 OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration Jusheng Zhang et.al. 2509.04876 null
2025-09-05 UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning Haoming Wang et.al. 2509.02544 null
2025-09-04 Maestro: Joint Graph & Config Optimization for Reliable AI Agents Wenxiao Wang et.al. 2509.04642 null
2025-09-04 Psychologically Enhanced AI Agents Maciej Besta et.al. 2509.04343 null
2025-09-04 Are LLM Agents the New RPA? A Comparative Study with RPA Across Enterprise Workflows Petr Průcha et.al. 2509.04198 null
2025-09-04 MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions Aishik Mandal et.al. 2509.04183 null
2025-09-04 Real-time adaptive quantum error correction by model-free multi-agent learning Manuel Guatto et.al. 2509.03974 null
2025-09-04 FaMA: LLM-Empowered Agentic Assistant for Consumer-to-Consumer Marketplace Yineng Yan et.al. 2509.03890 null
2025-09-04 Leveraging LLM-Based Agents for Intelligent Supply Chain Planning Yongzhi Qi et.al. 2509.03811 null
2025-09-04 AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems? Guibin Zhang et.al. 2509.03312 null
2025-09-03 Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation James Mooney et.al. 2509.03736 null
2025-09-02 DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Pranav Narayanan Venkit et.al. 2509.04499 null
2025-09-02 Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics Matthew Russo et.al. 2509.02751 null
2025-09-02 The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Guibin Zhang et.al. 2509.02547 null
2025-09-02 Towards Agents That Know When They Don’t Know: Uncertainty as a Control Signal for Structured Reasoning Josefa Lia Stoisser et.al. 2509.02401 null
2025-09-02 When Agents go Astray: Course-Correcting SWE Agents with PRMs Shubham Gandhi et.al. 2509.02360 null
2025-09-01 The Need for Verification in AI-Driven Scientific Discovery Cristina Cornelio et.al. 2509.01398 null
2025-09-01 Multi-Agent Reinforcement Learning for Task Offloading in Wireless Edge Networks Andrea Fox et.al. 2509.01257 null
2025-09-01 ORCA: ORchestrating Causal Agent Joanie Hayoun Chung et.al. 2508.21304 null
2025-09-01 How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $τ$ -bench Venkatesh Mishra et.al. 2508.20931 null
2025-09-01 Instructional Agents: LLM Agents on Automated Course Material Generation for Teaching Faculties Huaiyuan Yao et.al. 2508.19611 null
2025-08-31 Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First Shu Liu et.al. 2509.00997 null
2025-08-30 Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making Ziv Ben-Zion et.al. 2510.06222 null
2025-08-30 Exploring Decision-Making Capabilities of LLM Agents: An Experimental Study on Jump-Jump Game Juwu Li et.al. 2509.00483 null
2025-08-29 COCORELI: Cooperative, Compositional Reconstitution \& Execution of Language Instructions Swarnadeep Bhar et.al. 2509.04470 null
2025-08-29 ReLATE: Learning Efficient Sparse Encoding for High-Performance Tensor Decomposition Ahmed E. Helal et.al. 2509.00280 null
2025-08-29 HiVA: Self-organized Hierarchical Variable Agent via Goal-driven Semantic-Topological Evolution Jinzhou Tang et.al. 2509.00189 null
2025-08-28 A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers Ming Hu et.al. 2508.21148 null
2025-08-28 Provable Benefits of In-Tool Learning for Large Language Models Sam Houliston et.al. 2508.20755 null
2025-08-28 rStar2-Agent: Agentic Reasoning Technical Report Ning Shang et.al. 2508.20722 null
2025-08-28 CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics Stefano Fumero et.al. 2508.20643 null
2025-08-28 MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers Zhenting Wang et.al. 2508.20453 null
2025-08-28 MindGuard: Tracking, Detecting, and Attributing MCP Tool Poisoning Attack via Decision Dependence Graph Zhiqiang Wang et.al. 2508.20412 null
2025-08-27 CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Zeyi Sun et.al. 2508.20096 null
2025-08-27 AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios Lisa Alazraki et.al. 2508.19988 null
2025-08-27 Evaluating Language Model Reasoning about Confidential Information Dylan Sam et.al. 2508.19980 null
2025-08-27 Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey Yinqiu Liu et.al. 2508.19870 null
2025-08-27 Survey of Specialized Large Language Model Chenghan Yang et.al. 2508.19667 null
2025-08-27 CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation Zhejing Hu et.al. 2508.19603 null
2025-08-27 Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning Zhiwei Li et.al. 2508.19598 null
2025-08-27 Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents Kevin Song et.al. 2508.19504 null
2025-08-27 Interactive Graph Visualization and TeamingRecommendation in an Interdisciplinary Project’sTalent Knowledge Graph Jiawei Xu et.al. 2508.19489 null
2025-08-26 Reliable Weak-to-Strong Monitoring of LLM Agents Neil Kale et.al. 2508.19461 null
2025-08-26 Real-Time Model Checking for Closed-Loop Robot Reactive Planning Christopher Chandler et.al. 2508.19186 null
2025-08-26 MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation Ernest Lim et.al. 2508.19163 null
2025-08-26 A Concurrent Modular Agent: Framework for Autonomous LLM Agents Norihiro Maruyama et.al. 2508.19042 null
2025-08-26 CausalMACE: Causality Empowered Multi-Agents in Minecraft Cooperative Tasks Qi Chai et.al. 2508.18797 null
2025-08-26 Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions Ruichen Zhang et.al. 2508.18725 null
2025-08-26 FALCON: Autonomous Cyber Threat Intelligence Mining with LLMs for IDS Rule Generation Shaswata Mitra et.al. 2508.18684 null
2025-08-26 Utilizing Training Data to Improve LLM Reasoning for Tabular Understanding Chufan Gao et.al. 2508.18676 null
2025-08-26 Bias-Adjusted LLM Agents for Human-Like Decision-Making via Behavioral Economics Ayato Kitadai et.al. 2508.18600 null
2025-08-26 Generative Artificial Intelligence and Agents in Research and Teaching Jussi S. Jauhiainen et.al. 2508.16701 null
2025-08-25 Toward Generalized Autonomous Agents: A Neuro-Symbolic AI Framework for Integrating Social and Technical Support in Education Ryan Hare et.al. 2508.18406 null
2025-08-25 The AI Data Scientist Farkhad Akimov et.al. 2508.18113 null
2025-08-25 Memento: Fine-tuning LLM Agents without Fine-tuning LLMs Huichi Zhou et.al. 2508.16153 null
2025-08-24 FLAIRR-TS – Forecasting LLM-Agents with Iterative Refinement and Retrieval for Time Series Gunjan Jalori et.al. 2508.19279 null
2025-08-24 Agent-Testing Agent: A Meta-Agent for Automated Testing and Evaluation of Conversational AI Agents Sameer Komoravolu et.al. 2508.17393 null
2025-08-24 From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users Sadia Sultana Chowa et.al. 2508.17281 null
2025-08-22 AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications Dawei Gao et.al. 2508.16279 null
2025-08-22 IR-Agent: Expert-Inspired LLM Agents for Structure Elucidation from Infrared Spectra Heewoong Noh et.al. 2508.16112 null
2025-08-21 Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making Yuanjun Feng et.al. 2508.15926 null
2025-08-21 End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning Qiaoyu Zheng et.al. 2508.15746 null
2025-08-21 PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic Workflows Renan Souza et.al. 2508.02866 null
2025-08-20 Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL Weizhen Li et.al. 2508.13167 null
2025-08-13 Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code Muhammad Haseeb et.al. 2508.08322 null
2025-08-08 OmniEAR: Benchmarking Agent Reasoning in Embodied Tasks Zixuan Wang et.al. 2508.05614 null
2025-08-04 Agent Network Protocol Technical White Paper Gaowei Chang et.al. 2508.00007 null
2025-07-30 Agentic Web: Weaving the Next Web with AI Agents Yingxuan Yang et.al. 2507.21206 null
2025-07-04 MedAide: Information Fusion and Anatomy of Medical Intents via LLM-based Agent Collaboration Dingkang Yang et.al. 2410.12532 null
2025-06-27 LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey Henry Peng Zou et.al. 2505.00753 null
2025-06-09 AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML Patara Trirat et.al. 2410.02958 null
2025-06-04 Will Agents Replace Us? Perceptions of Autonomous Multi-Agent AI Nikola Balic et.al. 2506.02055 null
2025-04-29 From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Mohamed Amine Ferrag et.al. 2504.19678 null
2025-02-17 AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration Jizhou Chen et.al. 2502.09809 null
2024-12-10 Towards Effective GenAI Multi-Agent Collaboration: Design and Evaluation for Enterprise Applications Raphael Shu et.al. 2412.05449 null
2024-10-30 Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents Wenkai Yang et.al. 2402.11208 null
2024-07-11 Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative Intelligence Weize Chen et.al. 2407.07061 null
2024-02-20 KG-Agent: An Efficient Autonomous Agent Framework for Complex Reasoning over Knowledge Graph Jinhao Jiang et.al. 2402.11163 null
2024-01-30 A mechanism for discovering semantic relationships among agent communication protocols Idoia Berges et.al. 2401.16216 null
2024-01-23 Semantic Web Technology for Agent Communication Protocols Idoia Berges et.al. 2401.11841 null

Large Language Models

Publish Date Title Authors PDF Code
2026-09-09 Towards Tackling Application Logic Flaws through Autonomous Formal-Logic Modeling and Automated Reasoning Yiwei Fang et.al. 2609.10537 null
2026-09-09 Show-Harness: Just a VLM Agent Can Play Robots Yanzhe Chen et.al. 2609.10522 null
2026-09-09 Private communication via zero-private-capacity quantum channels Chengkai Zhu et.al. 2609.10520 null
2026-09-09 Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning Mehrnaz Mofakhami et.al. 2609.10445 null
2026-09-09 ConvMem: Convolutional Memory for Long-Context Reasoning Hongming Zhang et.al. 2609.10441 null
2026-09-09 Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs Ravi Ranjan et.al. 2609.10439 null
2026-09-09 Do speech foundation models really learn words? Robin Huo et.al. 2609.10434 null
2026-09-09 Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support Jonathan A. Handler et.al. 2609.10421 null
2026-09-09 Towards Scalable and Cost-Efficient Vulnerability Detection: A Study on Automatic Query Generation Ivana Clairine Irsan et.al. 2609.10412 null
2026-09-09 Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization Ayan Majumdar et.al. 2609.10410 null
2026-09-09 Retrofitting Code Using LLMs to Support Exceptional Behavior Linghan Zhong et.al. 2609.10397 null
2026-09-09 Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs Killian Steunou et.al. 2609.10355 null
2026-09-09 Unifying Score and Performance for Fine-Grained Music Understanding in Audio-Language Models Milan Liessens Dujardin et.al. 2609.10351 null
2026-09-09 Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs Haiji Liang et.al. 2609.10346 null
2026-09-09 From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning Weichen Dai et.al. 2609.10335 null
2026-09-09 Learning to Adapt and Calibrate: Score Distribution Alignment for Few-Shot Uncertainty Prediction in Medical VLMs Xuan Cuong Ngo et.al. 2609.10333 null
2026-09-09 On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data Hongyuan Zhang et.al. 2609.10321 null
2026-09-09 Ensembling LLMs for AI-Augmented Cybersecurity Software Requirements Generation Santiago Perez-Acuna et.al. 2609.10316 null
2026-09-09 TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards Rui Sun et.al. 2609.10315 null
2026-09-09 RiLM: Parameter-Efficient Language Modeling via Geodesic Decoding Fang Li et.al. 2609.10305 null
2026-09-08 Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Yuxing Lu et.al. 2609.09153 null
2026-09-08 Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs Xiaofu Chen et.al. 2609.09124 null
2026-09-08 MeClear: Cooperative Game-Theoretic Attribution and Risk-Aware Memory Clearance for Long-Horizon LLM Agents Boyu Yang et.al. 2609.09115 null
2026-09-08 Measuring LLM Sycophancy under Sustained Multi-Turn Pressure Leyuan Tang et.al. 2609.09090 null
2026-09-08 PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation Yixuan Liu et.al. 2609.09087 null
2026-09-08 It’s Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention Raito Kiya et.al. 2609.09085 null
2026-09-08 GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting Thodoris Betsas et.al. 2609.09082 null
2026-09-08 ToolLoop: Closed-Loop Tool-Use Data Synthesis via Decomposed Generation and Dynamic Self-Feedback Min Zeng et.al. 2609.09072 null
2026-09-08 Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics Andy Nkansah et.al. 2609.09070 null
2026-09-08 PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games Ryan Truong et.al. 2609.09059 null
2026-09-08 Training-Free Task Vectors for LLM Behavioral Control Gabriel J. Perin et.al. 2609.09054 null
2026-09-08 The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits Siddharth Vohra et.al. 2609.09048 null
2026-09-08 Do Reasoning Representations Help Humans Evaluate LLM Outputs? Jaewoo Lim et.al. 2609.09038 null
2026-09-08 Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning Mar Gonzàlez I Català et.al. 2609.09030 null
2026-09-08 Evaluation of Contextual Understanding in Large Language Models Subavarshana Arumugam et.al. 2609.09004 null
2026-09-08 Factorized and Vectorized Execution: Optimizing Analytical and Semantic Queries over Relations Sunny Yasser et.al. 2609.09002 null
2026-09-08 Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling Arman Adibi et.al. 2609.08981 null
2026-09-08 Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack Sohir Maskey et.al. 2609.08966 null
2026-09-08 PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving Yuan Gao et.al. 2609.08965 null
2026-09-08 SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL Evaluation Mohammadhossein Malekpour et.al. 2609.08950 null
2026-09-04 Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models Wonje Jeung et.al. 2609.05401 null
2026-09-04 Think-Verify-Revise: Neuro-Symbolic Visual Reasoning with Vision-Language Models and Dynamic Logic Tensor Networks Homayoun Afshari et.al. 2609.05388 null
2026-09-04 Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models Matthias Busch et.al. 2609.05381 null
2026-09-04 Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation Siliang Liu et.al. 2609.05363 null
2026-09-04 Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction Thang Tran Viet et.al. 2609.05361 null
2026-09-04 MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation Mohanad Albughdadi et.al. 2609.05351 null
2026-09-04 Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses Minne Chen et.al. 2609.05345 null
2026-09-04 Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models José Luciano Verçosa Marques et.al. 2609.05333 null
2026-09-04 Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness Alexander Neubauer et.al. 2609.05314 null
2026-09-04 RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Yang Li et.al. 2609.05295 null
2026-09-04 GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity Shuang Liang et.al. 2609.05284 null
2026-09-04 Don’t Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Mostafa Elhoushi et.al. 2609.05275 null
2026-09-04 Trace2Tower: Transition-Aware EigenTrace Induction of Multi-Level Skills for LLM Agents Jiazheng Sun et.al. 2609.05261 null
2026-09-04 Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization Sihan Ge et.al. 2609.05258 null
2026-09-04 Uncensored Open-weight Models: Redistribution as the Persistence Layer 10a Labs et.al. 2609.05241 null
2026-09-04 Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents Daniel Davila et.al. 2609.05239 null
2026-09-04 PRICE: A Systematic Study of LLM Adaptation Choices for Bitcoin Price Forecasting Maryam Fakhari et.al. 2609.05235 null
2026-09-04 ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs Zukang Xu et.al. 2609.05228 null
2026-09-04 First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves Tianjie Ju et.al. 2609.05224 null
2026-09-04 A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR Thi Kim Trang Vo et.al. 2609.05221 null
2026-09-03 Principia: Relational Physics Tests for Video Models Varun Varma Thozhiyoor et.al. 2609.04200 null
2026-09-03 Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints Haoyaun Zhu et.al. 2609.04198 null
2026-09-03 Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning Kevin Du et.al. 2609.04194 null
2026-09-03 Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views Joseph Lee et.al. 2609.04180 null
2026-09-03 Rethinking On-Policy Distillation of Large Language Models II: One Training Example Zixuan Fu et.al. 2609.04172 null
2026-09-03 From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research Yakov Pyotr Shkolnikov et.al. 2609.04166 null
2026-09-03 SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center Uday Vallabhaneni et.al. 2609.04159 null
2026-09-03 Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding Hongyu Qu et.al. 2609.04131 null
2026-09-03 Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable Shai Vardi et.al. 2609.04127 null
2026-09-03 Compressing Streaming Neural Audio Encoders via Latent-Space Distillation Prasanth Yadla et.al. 2609.04102 null
2026-09-03 Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving Ruoyu Yao et.al. 2609.04070 null
2026-09-03 When Models Edit Too Much: On the Fidelity of Minimal Code Edits Tongyao Zhu et.al. 2609.04061 null
2026-09-03 AI-Assisted Design of a Post-Quantum Cryptographic Accelerator: A Deployed-Silicon Case Study Jungmin Park et.al. 2609.04058 null
2026-09-03 LabelMate: An LLM-Driven Framework for Refined Issue Report Labeling Liam Johnston et.al. 2609.04055 null
2026-09-03 The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations Dmitrij Żatuchin et.al. 2609.04047 null
2026-09-03 IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations Chen Li et.al. 2609.04030 null
2026-09-03 Instruction Duplication as an Inference-Time Control Primitive Victor Lavrenko et.al. 2609.04024 null
2026-09-03 Representational alignment yields generalizable safety in language models Lingyu Li et.al. 2609.04022 null
2026-09-03 FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models Yalun Wu et.al. 2609.04021 null
2026-09-03 InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models Chao Shen et.al. 2609.04014 null
2026-09-02 User Feedback Provides a Unique Signal that LLMs Can not Detect Shachar Don-Yehiya et.al. 2609.02859 null
2026-09-02 The Implications of Linguistic Illegibility for LLM Security James Mickens et.al. 2609.02852 null
2026-09-02 Post-Training Language Models for Gold-Medal Performance in Coding Competitions Aleksander Ficek et.al. 2609.02849 null
2026-09-02 UE5M3 FP4 Block Scaling for Stable Language Model Pretraining Robert Hu et.al. 2609.02846 null
2026-09-02 Cliff: Learning Process Rewards from the First Mistake Peixuan Han et.al. 2609.02817 null
2026-09-02 Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis Hao Zhou et.al. 2609.02805 null
2026-09-02 Dutch Books for Language Models Isaiah Andrews et.al. 2609.02797 null
2026-09-02 DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation Vasileios Baltatzis et.al. 2609.02796 null
2026-09-02 ShikumiMiner: Mining Recurring Implementation Patterns in AI Codebases Afsana Tasnim et.al. 2609.02789 null
2026-09-02 ShallowStream: Index Shallow then Answer Deep for Streaming Video Understanding Jitai Hao et.al. 2609.02780 null
2026-09-02 Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit Wassim Tenachi et.al. 2609.02766 null
2026-09-02 Untangling the Mechanisms of Misleading Context in Medical Question Answering Robin Linzmayer et.al. 2609.02754 null
2026-09-02 Language Models Can Control Their Own Attention Namgyu Ho et.al. 2609.02737 null
2026-09-02 RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models Canjie Liu et.al. 2609.02731 null
2026-09-02 CORAL: An LLM-Native Harness for Production Recommender Systems Muhammad Rafay Azhar et.al. 2609.02730 null
2026-09-02 BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research Wooyoung Jung et.al. 2609.02729 null
2026-09-02 Large Language Model-Driven Context-Aware Eco-Feedback Generation and Evaluation Wooyoung Jung et.al. 2609.02719 null
2026-09-02 Door-in-the-Face Requests and Refusal Behaviour in Large Language Models Til Jordan et.al. 2609.02707 null
2026-09-02 ACLE-MCP: Attested Capability Leases for Execution-Time Trust in Remote LLM Tool Use Zhiyang Ding et.al. 2609.02690 null
2026-09-02 DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models Kushagra Bhushan et.al. 2609.02685 null
2026-09-01 CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses? Damien Sileo et.al. 2609.01600 null
2026-09-01 The Rise of Verbal Reinforcement Learning Kshitij Tayal et.al. 2609.01597 null
2026-09-01 The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally Jundong Hu et.al. 2609.01587 null
2026-09-01 Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics Maksim Evdokimov et.al. 2609.01575 null
2026-09-01 Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers Matteo Merler et.al. 2609.01567 null
2026-09-01 From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification Manish Gupta et.al. 2609.01564 null
2026-09-01 A systematic Approach to constructing a Chance-and-Risk Matrix for Semiconductor Supply Chains Ema Salkić et.al. 2609.01563 null
2026-09-01 SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue Stephanie Fong et.al. 2609.01548 null
2026-09-01 TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models Yuhang Dai et.al. 2609.01542 null
2026-09-01 Can LLMs Design Video Coding Tools? A Case Study on Planar Mode Yingwen Zhang et.al. 2609.01535 null
2026-09-01 Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall Jacqueline He et.al. 2609.01532 null
2026-09-01 When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation Peiying Zhu et.al. 2609.01519 null
2026-09-01 LatentPress: Context Compression Beyond Text and Vision Zhengze Zhou et.al. 2609.01507 null
2026-09-01 RadMatch: Auditable Radiology Report Evaluation via Finding-Level Matching Charles Corbière et.al. 2609.01470 null
2026-09-01 Just Talk Once: Communication-Efficient Split Federated LLM Fine-Tuning on Edge Devices Jiaxiang Geng et.al. 2609.01457 null
2026-09-01 When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning Yitong Guo et.al. 2609.01455 null
2026-09-01 Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation Yixuan Liu et.al. 2609.01432 null
2026-09-01 Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search Zhiliang Chen et.al. 2609.01431 null
2026-09-01 TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution Ruocan Wei et.al. 2609.01428 null
2026-09-01 From Rollouts to Recipes: Self-Contained Post-Training for LLMs Yifei Li et.al. 2609.01422 null
2026-08-31 Agentic research is oxymoronic Natalie B. Hogg et.al. 2608.31161 null
2026-08-31 Sharp Approximation Rates for Neural Networks with Affine Latent Parameterizations Shijun Zhang et.al. 2608.31157 null
2026-08-31 OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques Hamed Babaei Giglou et.al. 2608.31137 null
2026-08-31 DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening Yung Wei Shueh et.al. 2608.31128 null
2026-08-31 When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning Hamed Babaei Giglou et.al. 2608.31118 null
2026-08-31 InsightToast: Proactive Information Retrieval & Glanceable Visualization in the Side Channel of Data-Rich Meetings Mohammad Abolnejadian et.al. 2608.31115 null
2026-08-31 Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions Ahmed El Kady et.al. 2608.31108 null
2026-08-31 BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing Adrians Skapars et.al. 2608.31105 null
2026-08-31 S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Jiajun Shi et.al. 2608.31100 null
2026-08-31 Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization Camila Blank et.al. 2608.31079 null
2026-08-31 Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization Jingxiao Yang et.al. 2608.31077 null
2026-08-31 A Model with No Head and Many Thoughts Nikita Koriagin et.al. 2608.31069 null
2026-08-31 Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores Qiyao Yan et.al. 2608.31068 null
2026-08-31 Every Token Leaves a Ripple in the Stream of Thought: Eliciting Model-Internal Token Saliency for Chain-of-Thought Compression Tianyi Zhao et.al. 2608.31066 null
2026-08-31 When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models Joonyong Park et.al. 2608.31035 null
2026-08-31 Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols Chengyuan Gao et.al. 2608.31014 null
2026-08-31 TSPFN: A Temporal Tabular Foundation Model for Physiological Time Series Classification Jérémie Stym-Popper et.al. 2608.31013 null
2026-08-31 From Prompt to Prototype: Towards a Frontier LLM Driven RF Engineering Workflow Markus Heinrichs et.al. 2608.31006 null
2026-08-31 Multi-View Reflective Surface Inspection via Semantic-Saliency Cross-Verification Van-Giang Nguyen et.al. 2608.30997 null
2026-08-31 Stick to What You Know: A Study of Knowledge-Aligned Supervised Fine-Tuning Arthur Becker et.al. 2608.30987 null
2026-08-31 MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines Alireza Bayat Makou et.al. 2608.30662 null
2026-08-31 SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators? Jinshan Gao et.al. 2608.30661 null
2026-08-31 LLM-based Hardware Development with Hierarchical IRs and End-to-End Multi-Agent Workflow Chenyang Yin et.al. 2608.30659 null
2026-08-31 Fine-Grained Multi Image Object Hallucination Benchmark Joonki Min et.al. 2608.30653 null
2026-08-31 Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models Kangwook Ko et.al. 2608.30649 null
2026-08-31 What It Costs to Compose, Rebuild, and Correct Precomputed Memory Asa Shepard et.al. 2608.30647 null
2026-08-31 BiG-SURE - Bipartite Graph for Semantic Uncertainty and Reliability Estimation of LLMs Debarpan Bhattacharya et.al. 2608.30646 null
2026-08-31 Test-time Reinforcement Learning in Imperfect Information Games Ondrej Kubicek et.al. 2608.30635 null
2026-08-31 GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning Outongyi Lv et.al. 2608.30632 null
2026-08-31 REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation Haoran Que et.al. 2608.30627 null
2026-08-31 Textual Acoustic Grounding for Generalizable LLM-Based Deepfake Voice Detection Yassine El Kheir et.al. 2608.30622 null
2026-08-31 Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text Minkyung Cho et.al. 2608.30619 null
2026-08-31 Reading the News: Adapting Large Language Models to Swedish Journalism Through Continued Pre-Training Lukas Borggren et.al. 2608.30609 null
2026-08-31 Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster Songtao Fang et.al. 2608.30606 null
2026-08-31 The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal Md Mokarram Chowdhury et.al. 2608.30585 null
2026-08-31 Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an Oracle Dennis Gross et.al. 2608.30581 null
2026-08-31 TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI Yuheng Zhang et.al. 2608.30567 null
2026-08-31 Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs Deokjae Lee et.al. 2608.30564 null
2026-08-31 SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards Zixun Guo et.al. 2608.30559 null
2026-08-31 GarmentWeaver: Schema-Aware Structured Synthesis for Multimodal Sewing Patterns Yinwen Lu et.al. 2608.30550 null
2026-08-28 Learning a Size-Weight Frontier for Synthetic-Augmented Inference Chengpiao Huang et.al. 2608.28576 null
2026-08-28 GeBDA: Building Damage Assessment as Text-Based Sequence Prediction Olivier Dietrich et.al. 2608.28567 null
2026-08-28 A Formal Limitation on Learning Human Language From Textual Corpora Emily Cheng et.al. 2608.28560 null
2026-08-28 Logos: An Agent Harness on a Cross-Process Bus Hanzhang Jia et.al. 2608.28553 null
2026-08-28 xTRUCE: A Provably Safe Arbiter for Multi-xApp Conflict Mitigation in Agentic O-RAN Le Xia et.al. 2608.28532 null
2026-08-28 Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration Simeng Sun et.al. 2608.28511 null
2026-08-28 Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation Di Wu et.al. 2608.28496 null
2026-08-28 LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment Jingjing Nie et.al. 2608.28490 null
2026-08-28 NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry Samuel Xiao et.al. 2608.28481 null
2026-08-28 Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge Zhuoshi Pan et.al. 2608.28478 null
2026-08-28 ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL Zhuoshi Pan et.al. 2608.28476 null
2026-08-28 ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT Huseyin Umut Isik et.al. 2608.28455 null
2026-08-28 Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning Minghui Xu et.al. 2608.28447 null
2026-08-28 Sliding-window beats linear attention Alexia Jolicoeur-Martineau et.al. 2608.28444 null
2026-08-28 Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL Jiayan Lin et.al. 2608.28432 null
2026-08-28 Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs Vishvesh Bhat et.al. 2608.28421 null
2026-08-28 LongPIBench: A Long-Context Benchmark for Prompt Injection Yupei Liu et.al. 2608.28411 null
2026-08-28 SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data Zi-Jian Cheng et.al. 2608.28408 null
2026-08-28 Post-Training VLMs for Video Mistake Detection Federico Spurio et.al. 2608.28406 null
2026-08-28 CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia Bryan Chen Zhengyu Tan et.al. 2608.28405 null
2026-08-27 UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City Tianjie Ju et.al. 2608.27456 null
2026-08-27 CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes Yufan Wu et.al. 2608.27455 null
2026-08-27 SWE-Prime: Fewer Trajectories, Better Performance Dewu Zheng et.al. 2608.27449 null
2026-08-27 TTPO: Test-Time Policy Optimization Aozhe Wang et.al. 2608.27448 null
2026-08-27 Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach? Ting Yan et.al. 2608.27443 null
2026-08-27 From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench Dewu Zheng et.al. 2608.27442 null
2026-08-27 Stochastic Estimation of Transduced Language Models Vésteinn Snæbjarnarson et.al. 2608.27428 null
2026-08-27 Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit Yisen Xi et.al. 2608.27427 null
2026-08-27 Boosting LLM Exploration via Weak-Model Guidance in RLVR Xingyu Shen et.al. 2608.27420 null
2026-08-27 Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information Chanho Park et.al. 2608.27417 null
2026-08-27 Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms Siye Wu et.al. 2608.27409 null
2026-08-27 How Language Models Organize and Structure Moral Knowledge Orion Reblitz-Richardson et.al. 2608.27402 null
2026-08-27 Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction Jin Mu et.al. 2608.27397 null
2026-08-27 Embodied Scene Rearrangement Planning Canzhi Chen et.al. 2608.27371 null
2026-08-27 Puro-2B: Poor Lab’s Qwen2-1.5B Trained on RTX 5090 within $5090 Kairong Luo et.al. 2608.27370 null
2026-08-27 Sophistication in GenAI Use: Field Evidence from a Large Firm Nicholas J. Hallman et.al. 2608.27364 null
2026-08-27 INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment Yutong Zhang et.al. 2608.27348 null
2026-08-27 Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance Allison Zhuang et.al. 2608.27340 null
2026-08-27 One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles Zhichen Zeng et.al. 2608.27338 null
2026-08-27 A blueprint for the formalization of norm-variation of multiple ergodic averages for commuting transformations Floris van Doorn et.al. 2608.27321 null
2026-08-26 Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization Jiaming Zhou et.al. 2608.26103 null
2026-08-26 TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development Jiarui Yan et.al. 2608.26086 null
2026-08-26 SwarmWorld: Stigmergic technological evolution in societies of language-model agents Subhadeep Pal et.al. 2608.26081 null
2026-08-26 Prefix Sliding for efficient test-time scaling Niklas Muennighoff et.al. 2608.26070 null
2026-08-26 RTLGuard: A Lightweight Teacher-Student Defense for Poisoned RTL Code Generation Models Mahshid Rezakhani et.al. 2608.26049 null
2026-08-26 Robust CurveMoE: Multi-Norm Adversarial Defense for Mixture-of-Experts Models via Mode Connectivity Xu Zhang et.al. 2608.26043 null
2026-08-26 Vulnerable Code Search: Transferable Attack for Code Language Models Kaicheng Wang et.al. 2608.26031 null
2026-08-26 Bayesian Optimization for Self-Driving Materials Laboratories: From Algorithms to Physics-Informed Workflows Yuki K. Wakabayashi et.al. 2608.26016 null
2026-08-26 VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following Min Zeng et.al. 2608.26013 null
2026-08-26 A Self-Evolving Multi-Agent Framework Defense against LLM Jailbreak Attacks Tongyan Hu et.al. 2608.26008 null
2026-08-26 VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Zhifei Xie et.al. 2608.26005 null
2026-08-26 Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing Rui He et.al. 2608.25999 null
2026-08-26 ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs Songyuan Li et.al. 2608.25992 null
2026-08-26 Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon Xiaodong Wu et.al. 2608.25990 null
2026-08-26 Multi-Granularity Context-Enhanced RAG over Multimodal Knowledge Graphs Zongyu Wu et.al. 2608.25986 null
2026-08-26 When Personality Meets Quantization: A Layer-wise MBTI Analysis of Quantized LLMs Yao Fu et.al. 2608.25977 null
2026-08-26 SciMIF: Understanding Multimodal Instruction Following in Scientific Domains Ye Shen et.al. 2608.25973 null
2026-08-26 Spatial-Knowledge-Graph-Grounded LLM Agents for Neighborhood Livability Evaluation Haiyan Hao et.al. 2608.25952 null
2026-08-26 Unveiling Spectral Mechanisms in Training-Free LLM Text Detection Haitong Luo et.al. 2608.25944 null
2026-08-26 When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs Suchit Gupte et.al. 2608.25941 null
2026-08-25 Prompt Structure Redistributes, Not Reduces: An Empirical Analysis of Security-Weaknesses in LLM-Generated Python Code Maitreyee Das Urmi et.al. 2608.24857 null
2026-08-25 BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes Fei Tang et.al. 2608.24848 null
2026-08-25 Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows Miao Liu et.al. 2608.24842 null
2026-08-25 A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments Jing Huang et.al. 2608.24825 null
2026-08-25 Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA Emanuel Kitzelmann et.al. 2608.24824 null
2026-08-25 Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining Zihan Liu et.al. 2608.24814 null
2026-08-25 Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought Mengzhu Xu et.al. 2608.24790 null
2026-08-25 MoE-based Feature Adapter for Prompt-free Binary Coronary Artery Segmentation in X-ray Angiography Lin Xi et.al. 2608.24783 null
2026-08-25 Evidence Blindness in Direct Corpus Interaction: Persistent Navigation with AtlasNav Hongyu Guo et.al. 2608.24764 null
2026-08-25 MoTE: Mixture of Task Experts for Multi-Task Video Understanding Muhammad Asad Ali et.al. 2608.24763 null
2026-08-25 From Natural Language Requirements to Graphical User Interfaces: Automated Prototyping and Verification with Pretrained Language Models Kristian Kolthoff et.al. 2608.24749 null
2026-08-25 SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents Shidong Yang et.al. 2608.24747 null
2026-08-25 Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion Zilong Huang et.al. 2608.24730 null
2026-08-25 On-policy Distillation with Verifiable Reward Wenze Lin et.al. 2608.24696 null
2026-08-25 The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models Augusto Camargo et.al. 2608.24662 null
2026-08-25 Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning Zhengyang Zhang et.al. 2608.24658 null
2026-08-25 A Literate Programming Environment for Human and Machine Agents Adam T. Burke et.al. 2608.24644 null
2026-08-25 Aging of Prompt Engineering Techniques Across LLM Versions Anastasiia Rudyk et.al. 2608.24641 null
2026-08-25 Thermal Tuning Overhead in Wafer-Scale Optical Interconnects for LLM MoE Training: A Cross-Layer Analysis and Ferroelectric-Based Mitigation Seongwon Yoon et.al. 2608.24637 null
2026-08-25 Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding Yujing Chang et.al. 2608.24621 null
2026-08-24 How to Train a Critic Stably and Efficiently Penghui Qi et.al. 2608.23566 null
2026-08-24 EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings Md Thamed Bin Zaman Chowdhury et.al. 2608.23563 null
2026-08-24 Prime Agent: A Self-Improving RLM Harness Seth Karten et.al. 2608.23552 null
2026-08-24 ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings Na Li et.al. 2608.23551 null
2026-08-24 Investigating Relational Reasoning in VLMs Adhithya Laxman Ravi Shankar Geetha et.al. 2608.23518 null
2026-08-24 Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search Thanh-Khoi Nguyen et.al. 2608.23503 null
2026-08-24 Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty Yipeng Zhao et.al. 2608.23497 null
2026-08-24 SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning Jialong Liu et.al. 2608.23493 null
2026-08-24 StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models Jinghan Tan et.al. 2608.23475 null
2026-08-24 What’s the Catch? Evaluating Temporal Consistency in Vision-Language Models Marek Hradil et.al. 2608.23474 null
2026-08-24 ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation Zhongpan Tang et.al. 2608.23463 null
2026-08-24 A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps – A Bibliometric and Topic-Based Study Mullosharaf K. Arabov et.al. 2608.23421 null
2026-08-24 Systematic Bias in Green Patent Classification: Silent Green and False Green Hamid Bekamiri et.al. 2608.23420 null
2026-08-24 Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data Yifei Song et.al. 2608.23391 null
2026-08-24 Adversarial Entropy Inflation Against Gumbel-Based Inference Verification Nikita Kezins et.al. 2608.23375 null
2026-08-24 Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting Seyed Mohammad Hossein Hashemi et.al. 2608.23373 null
2026-08-24 Walking on the DARKSIDE Aldo Gangemi et.al. 2608.23370 null
2026-08-24 DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts Vlad Hondru et.al. 2608.23363 null
2026-08-24 OptiSight: Bridging Semantic Reasoning and Geometric Control for Embodied Navigation Alperen Avan et.al. 2608.23354 null
2026-08-24 FormuEvo: LLM-Guided Evolution for Discovering Solver-Efficient Mixed-Integer Programming Formulations Haofeng Yuan et.al. 2608.23353 null
2026-08-21 OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs Xianyun Sun et.al. 2608.21360 null
2026-08-21 VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences Elaine Lau et.al. 2608.21357 null
2026-08-21 ViTacPhys: Physical Property-Aware Grasping from Human Visual-Tactile Demonstrations Yiwen Liu et.al. 2608.21355 null
2026-08-21 Asymmetric Capacity Allocation in Self-Refinement Pipelines Zhuoyi Yang et.al. 2608.21345 null
2026-08-21 Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy Afonso Baldo et.al. 2608.21325 null
2026-08-21 Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout – and the factorizations of it that failed Nicolás Vera Zúñiga et.al. 2608.21315 null
2026-08-21 Rethinking Expressivity and Efficiency in Test-Time Training Zeyun Zhong et.al. 2608.21308 null
2026-08-21 Re $^3$ Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning Haonan Jia et.al. 2608.21305 null
2026-08-21 Human-AI Collaboration in Requirements Engineering: Evidence of the Negative Effect of LLMs on Requirements Inspection Giovanna Broccia et.al. 2608.21298 null
2026-08-21 Level-k Distinguishable Mechanisms for Evaluating Bounded Rationality in LLMs Binchi Zhang et.al. 2608.21296 null
2026-08-21 CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment Chengxiao Wang et.al. 2608.21278 null
2026-08-21 ConceptTS: LLM-Guided Concept Bottlenecks for Interpretable Multivariate Time-Series Forecasting Yichen Jiang et.al. 2608.21277 null
2026-08-21 Memory Augmentation Unlocks Efficient Chain-of-Thought Reasoning Simeng Zhang et.al. 2608.21265 null
2026-08-21 Benchmarking Patent Drafting from Inventor-Style Disclosures Lekang Jiang et.al. 2608.21249 null
2026-08-21 Just Noticeable Difference Modeling for Token Compression in Vision-Language-Action Models Zhuoyuan Li et.al. 2608.21247 null
2026-08-21 A VLM Answer Is Not an Anomaly Score: Rank Compression in Training-Free Video Anomaly Detection Inpyo Song et.al. 2608.21244 null
2026-08-21 Affective Context Amplifies Sycophancy in LLM Responses Jiayi Li et.al. 2608.21242 null
2026-08-21 SPICE: Speculative Prefetching with Low-Rank Expert Surrogates and Heterogeneous Orchestration for MoE Inference Acceleration Yongxiang Lyu et.al. 2608.21240 null
2026-08-21 Indexing Long Documents for LLM-Based Analysis Donna Pham et.al. 2608.21237 null
2026-08-21 RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models Zhibo Zhang et.al. 2608.21236 null
2026-08-20 ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models Sahil Kale et.al. 2608.20338 null
2026-08-20 Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models Taihang Hu et.al. 2608.20334 null
2026-08-20 An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction Narges Ahmadi et.al. 2608.20320 null
2026-08-20 Projecting BrowseComp-Plus onto ClimbMix: Toward More Realistic Corpora for Agentic Search Sahel Sharifymoghaddam et.al. 2608.20317 null
2026-08-20 Explainable Transformer Models for Clinical Prediction Tasks on Structured Electronic Health Records Jun Ni Du et.al. 2608.20315 null
2026-08-20 MidTool: Mid-training Data Synthesis for Agentic Tool Use Fengqing Jiang et.al. 2608.20314 null
2026-08-20 Phantom Gains: Auditing Self-Improvement Against a Measured Null Cheng Xu et.al. 2608.20290 null
2026-08-20 Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization Qian Kou et.al. 2608.20281 null
2026-08-20 Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents Yiyang Feng et.al. 2608.20274 null
2026-08-20 Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation Gijs Kassenaar et.al. 2608.20256 null
2026-08-20 Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models Yu Chen et.al. 2608.20237 null
2026-08-20 Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference Christos Koutsiaris et.al. 2608.20210 null
2026-08-20 MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use Mengru Wang et.al. 2608.20202 null
2026-08-20 Decoding silent reading from non-invasive EEG Ingo Marquardt et.al. 2608.20186 null
2026-08-20 Ask Self, Ask Others: Relation Is All You Need Yuting Ge et.al. 2608.20172 null
2026-08-20 DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing Haoxiang Cao et.al. 2608.20161 null
2026-08-20 FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models Dingzirui Wang et.al. 2608.20153 null
2026-08-20 DPC-Net: Dual-Prior Collaborative Network for All-in-One Image Restoration Zhaokun He et.al. 2608.20141 null
2026-08-20 Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving Mehdi Azarafza et.al. 2608.20129 null
2026-08-20 Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo Lohithsai Yadala Chanchu et.al. 2608.20123 null
2026-08-19 Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication Ramneet Kaur et.al. 2608.19161 null
2026-08-19 Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems George Andrikopoulos et.al. 2608.19140 null
2026-08-19 Comment-level Topic Drift Analysis in the Reddit Corpus Steven Morse et.al. 2608.19133 null
2026-08-19 JANUS: A Multi-modal Foundation Neural Sampler for Disordered Materials Denis Blessing et.al. 2608.19116 null
2026-08-19 ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models Jihae Jeong et.al. 2608.19075 null
2026-08-19 What is Missing from AI Post-Training AI: An Empirical Analysis Joy Jia Yin Lim et.al. 2608.19072 null
2026-08-19 Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models Valentin Romanov et.al. 2608.19025 null
2026-08-19 From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation Sepehr Ghaffarzadegan et.al. 2608.19011 null
2026-08-19 Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning Yajie Yin et.al. 2608.19009 null
2026-08-19 Mise-en-Scène: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation Zipeng Xu et.al. 2608.19000 null
2026-08-19 TractorBeam: Personalized AI Sensemaking Support via Collaborative Machine Annotation Sireesh Gururaja et.al. 2608.18994 null
2026-08-19 ForeSightGuide: An Anticipatory Framework toward Accurate and Low-Redundancy Guidance for the Visually Impaired Zhiyuan Wang et.al. 2608.18993 null
2026-08-19 Uncertainty-Aware Art-Historical Dating with Vision-Language Models Stefanie Schneider et.al. 2608.18984 null
2026-08-19 rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation Minh Hoang Nguyen et.al. 2608.18952 null
2026-08-19 Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis Bogdan Zagribelnyy et.al. 2608.18940 null
2026-08-19 Breaking the weakest link to evade vision language models Ilan Zini et.al. 2608.18938 null
2026-08-19 MedUAG: Unified Understanding and Generation for Medical Multimodal Models Zijie Meng et.al. 2608.18937 null
2026-08-19 Graphical Design of Interpretable Architectures Pietro Barbiero et.al. 2608.18936 null
2026-08-19 SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution Silin Chen et.al. 2608.18933 null
2026-08-19 Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck Davide Romano et.al. 2608.18931 null
2026-08-18 Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation Iryna Hartsock et.al. 2608.18072 null
2026-08-18 TokEval: A Tokenizer Evaluation Suite Clara Meister et.al. 2608.18062 null
2026-08-18 Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation Hollis Robbins et.al. 2608.18041 null
2026-08-18 Where A Small Language Model Helps in Invoice Categorisation, Understood Through Embedding Geometry Emma Ceccherini et.al. 2608.18033 null
2026-08-18 Chain-of-Experience for Continual LLM Improvement Haoqin Tu et.al. 2608.18027 null
2026-08-18 Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach Lu Xu et.al. 2608.18017 null
2026-08-18 Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering Hsiang-Wei Huang et.al. 2608.18009 null
2026-08-18 Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents Christophe D. Hounwanou et.al. 2608.18008 null
2026-08-18 A Denotational Semantics for Synchronized Regular Expressions (extended version) Lukas Grätz et.al. 2608.18007 null
2026-08-18 Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media Yijie Xu et.al. 2608.17987 null
2026-08-18 Recirculation Michael C. Mozer et.al. 2608.17981 null
2026-08-18 aDSL: Agentic 3D Creation via Joint Agent-Program Design Rui-Huan Wang et.al. 2608.17975 null
2026-08-18 Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection Bin Li et.al. 2608.17965 null
2026-08-18 Understanding the Surprising Generalization Properties of Tabular Foundation Models Nour Shaheen et.al. 2608.17957 null
2026-08-18 Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds Md. Faiyaz Abdullah Sayeedi et.al. 2608.17950 null
2026-08-18 SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE Xuan Zheng et.al. 2608.17948 null
2026-08-18 Procedural Content Metageneration via Program Search and Continual Abstraction Discovery Matthew Siper et.al. 2608.17947 null
2026-08-18 Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation Zhizhao Liu et.al. 2608.17941 null
2026-08-18 Grading Needs a Rubric, Not Intelligence Jhen-Ke Lin et.al. 2608.17938 null
2026-08-18 PerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation Jianyu Sun et.al. 2608.17926 null
2026-08-17 Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text Benjamin Belay et.al. 2608.16868 null
2026-08-17 What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models Saisab Sadhu et.al. 2608.16852 null
2026-08-17 Proteus: Incremental Memory Activation for Long-Context Sequence Modeling Reza Bayat et.al. 2608.16844 null
2026-08-17 Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning Minh-Ha Nguyen et.al. 2608.16831 null
2026-08-17 When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents Jiawei Liu et.al. 2608.16806 null
2026-08-17 Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models Yuanzhi Xu et.al. 2608.16805 null
2026-08-17 Neurosymbolic Embodied Agents Mohammad Albinhassan et.al. 2608.16794 null
2026-08-17 “This Is So Claude!” Towards a Theory of the Recognition of AI Character Without Reidentification Michele Loi et.al. 2608.16789 null
2026-08-17 Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis Reza Fayyazi et.al. 2608.16775 null
2026-08-17 LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing Ruoqi Shu et.al. 2608.16763 null
2026-08-17 Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments Adam Karvonen et.al. 2608.16747 null
2026-08-17 VicEdit: Learning to Edit Videos from Visual In-Context Examples Yuji Wang et.al. 2608.16745 null
2026-08-17 TDD-Agent: Test-Driven Reasoning for Code Generation Hongyue Yu et.al. 2608.16742 null
2026-08-17 Le Critique: Privileged Value Functions for LLM Reinforcement Learning Siddarth Venkatraman et.al. 2608.16739 null
2026-08-17 MatchingPolicy: Correspondence-Aware Policy Enables Cross-Object In-Context Learning Qijin She et.al. 2608.16715 null
2026-08-17 Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors David Eric Austin et.al. 2608.16707 null
2026-08-17 AnchorScore: A CLIP-Based Diagnostic of MLLM Annotation Difficulty Yan Ma et.al. 2608.16690 null
2026-08-17 Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test Anand Murugan et.al. 2608.16671 null
2026-08-17 Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL Yi Ai et.al. 2608.16663 null
2026-08-17 The ultimate carbon cost of a ChatGPT query Paul Kron et.al. 2608.16657 null
2026-08-14 Finding Vulnerabilities via LLM-Augmented Semantics-Aware Type-Checking Ruizhe Wang et.al. 2608.14533 null
2026-08-14 Handover of In-Context Learning State Across Session Boundaries Masahiro Kato et.al. 2608.14528 null
2026-08-14 Validating LLM-Modernized Scientific Software Through Differential Fault Injection Evan Coleman et.al. 2608.14527 null
2026-08-14 Split the Labor: Separating Evidence Interpretation from Decision Aggregation Zhelun Wu et.al. 2608.14509 null
2026-08-14 Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training Hanfeng Lu et.al. 2608.14498 null
2026-08-14 You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model Ziyang Luo et.al. 2608.14465 null
2026-08-14 SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning Panjing He et.al. 2608.14452 null
2026-08-14 More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It Haohui Yang et.al. 2608.14420 null
2026-08-14 STINER: Automated Extraction of Strategic Cyber Threat Intelligence from X Yasir Ech-Chammakhy et.al. 2608.14418 null
2026-08-14 Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice Syeda Anshrah Gillani et.al. 2608.14399 null
2026-08-14 LLMs Don’t Pay for the Jump Paras Balani et.al. 2608.14397 null
2026-08-14 Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons Wei Zhao et.al. 2608.14392 null
2026-08-14 DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding Zewen Jin et.al. 2608.14385 null
2026-08-14 A Survey of Large Models in Sports Yichen Xu et.al. 2608.14377 null
2026-08-14 CoRun: Padding is Simple and Efficient for Deterministic LLM Inference Shiju Zhao et.al. 2608.14376 null
2026-08-14 A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Process Models Md Kamrul Islam et.al. 2608.14370 null
2026-08-14 Local and Global Regimes of Geometric Complexity in Language Model Representations Arwa Osman et.al. 2608.14361 null
2026-08-14 ATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata Learning Ignacio D. Lopez-Miguel et.al. 2608.14352 null
2026-08-14 Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM Paths Seeyeon Kim et.al. 2608.14333 null
2026-08-14 AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs Yiderigun Borjigin et.al. 2608.14320 null
2026-08-13 LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure Fanfei Li et.al. 2608.13545 null
2026-08-13 SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization Weihan Meng et.al. 2608.13538 null
2026-08-13 DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees Tianyi Li et.al. 2608.13524 null
2026-08-13 DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data Peter Schneider-Kamp et.al. 2608.13517 null
2026-08-13 Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining Yuto Nishida et.al. 2608.13515 null
2026-08-13 OpScale: Operator-level Provisioning and Autoscaling for LLM Serving Xingqi Cui et.al. 2608.13499 null
2026-08-13 Synthetic Persona Pretraining: Alignment from Token Zero Julian Minder et.al. 2608.13482 null
2026-08-13 AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models Mohammed Ayman Habib et.al. 2608.13472 null
2026-08-13 MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification Daniel Perkins et.al. 2608.13463 null
2026-08-13 CAPRI: Contract-Aware Proof Repair for Isabelle Jim Woodcock et.al. 2608.13459 null
2026-08-13 Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts Imtiaz Ul Hassan et.al. 2608.13458 null
2026-08-13 LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles Md Wasiul Haque et.al. 2608.13450 null
2026-08-13 Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ Zongyun Zhang et.al. 2608.13441 null
2026-08-13 Algebraic Decomposition Theory for Transformer Length Generalization Andy Yang et.al. 2608.13433 null
2026-08-13 Are You Sure You’re Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity Irina Proskurina et.al. 2608.13430 null
2026-08-13 RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level Juan Irving Vasquez et.al. 2608.13428 null
2026-08-13 Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference Zixuan Lan et.al. 2608.13426 null
2026-08-13 Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes Aimilios Hadjiliasi et.al. 2608.13420 null
2026-08-13 CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation Enhan Li et.al. 2608.13387 null
2026-08-13 When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Jiaqian Li et.al. 2608.13385 null
2026-08-12 The Role Specialization Model (RSM): Coordinating LLM-Based Tools in Agentic Software Development - An Exploratory Case Study Carlos Alberto Fernández-y-Fernández et.al. 2608.12311 null
2026-08-12 Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models Saman Marandi et.al. 2608.12304 null
2026-08-12 Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior Yusuf Pisan et.al. 2608.12292 null
2026-08-12 Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence Aman Tyagi et.al. 2608.12290 null
2026-08-12 PatternFormer: Learning Multiple Solution Patterns in Reaction–Diffusion Systems Zhipeng Chang et.al. 2608.12286 null
2026-08-12 Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals Alireza Kargarzadeh et.al. 2608.12283 null
2026-08-12 Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams Weihao Bo et.al. 2608.12262 null
2026-08-12 One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Simon Yu et.al. 2608.12253 null
2026-08-12 Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting Junyi Ye et.al. 2608.12251 null
2026-08-12 HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression Yuefeng Zhang et.al. 2608.12239 null
2026-08-12 Towards Automated Domain Model Extraction from Source Code using Heuristics and Open-Source LLMs Alessandra Mancas et.al. 2608.12228 null
2026-08-12 SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward Zile Zhou et.al. 2608.12220 null
2026-08-12 ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening Antoine de Mathelin et.al. 2608.12219 null
2026-08-12 Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge Arda Uzunoglu et.al. 2608.12218 null
2026-08-12 “Pharos Night: Crown Pursuit”: An AI-Native Deck-Building and Tactical Arena Game Design Based on Multi-Agent Systems Ting-Chen Hsu et.al. 2608.12216 null
2026-08-12 Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Zhongbin Guo et.al. 2608.12209 null
2026-08-12 NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation Jiarui Ma et.al. 2608.12197 null
2026-08-12 IF:CARGO: LLM-Based Semantic Compilation for Al-Native Rule Programming Games Ting-Chen Hsu et.al. 2608.12195 null
2026-08-12 Making Collaborative Signals Count: Graph-Aware Large Language Models for Sequential Recommendation Fenglin Yan et.al. 2608.12184 null
2026-08-12 Context Blindness in DPO: Mitigating Object Hallucination in MLLMs via Context-Calibrated Preference Optimization Byungoh Ko et.al. 2608.12158 null
2026-08-11 MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Changhao Xiang et.al. 2608.11167 null
2026-08-11 Role of Personality in Conversational Information Seeking Abdisalam Abukar et.al. 2608.11164 null
2026-08-11 Scheduling Mixed RL Rollouts Beyond Prefix Locality Zetao Hong et.al. 2608.11152 null
2026-08-11 CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting Jiayu Ding et.al. 2608.11150 null
2026-08-11 PRMU: A Corpus-Free Benchmark for Person-Centric Knowledge Unlearning in Multimodal Large Language Models Huafeng Chen et.al. 2608.11149 null
2026-08-11 The Illusion of Cross-Lingual Safety in Low-Resource Languages Abigail Oppong et.al. 2608.11146 null
2026-08-11 Attention-Path Fragility as an Uncertainty Signal in Large Language Models Minsoo Kim et.al. 2608.11138 null
2026-08-11 Generative AI use in Statistical Research: A Literature Review and Code Generation Case Study Natalie Morosin et.al. 2608.11121 null
2026-08-11 CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering Mouxiao Huang et.al. 2608.11074 null
2026-08-11 Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory Ming Yang et.al. 2608.11066 null
2026-08-11 V-FiLLM: Verified Financial LLM Reasoning Benchmark Alicia Larsen et.al. 2608.11047 null
2026-08-11 TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification Jian Zhang et.al. 2608.11044 null
2026-08-11 Who Are You Explaining To? A Multi-Agent System for Audience-Aware XAI Narratives Francesco Musicco et.al. 2608.11033 null
2026-08-11 Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching Jian Zhang et.al. 2608.11030 null
2026-08-11 Mapping and Measuring the Behavioral Evolution of Large Language Models Dong Qiao et.al. 2608.11027 null
2026-08-11 Data Attribution of Emergent Misalignment with Persona Features Clemens Vetter et.al. 2608.11025 null
2026-08-11 When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models Yufei Zhang et.al. 2608.11024 null
2026-08-11 Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data Nicola Giuseppe Marchioro et.al. 2608.11022 null
2026-08-11 ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering Taojie Zhu et.al. 2608.10996 null
2026-08-11 What Iterated Self-Feeding Probes of Language Models Measure, and a test that separates the construction from the model Nicolás Vera Zúñiga et.al. 2608.10986 null
2026-08-07 SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation Hanke Xie et.al. 2608.07462 null
2026-08-07 CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity Ananya Sahu et.al. 2608.07460 null
2026-08-07 Strategy-first synthesis planning for complex natural products Daniel Armstrong et.al. 2608.07454 null
2026-08-07 Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools Afreen Alam et.al. 2608.07446 null
2026-08-07 Blast Radius MY Pitsane et.al. 2608.07440 null
2026-08-07 An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis Brian Llinas et.al. 2608.07439 null
2026-08-07 Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Jiacheng Miao et.al. 2608.07437 null
2026-08-07 SABRE: Scalable and Automated Benchmarking of VLMs under Stress Zixuan Lan et.al. 2608.07435 null
2026-08-07 Conformal Coverage Guarantees for Any Video Temporal Grounder Aseel Mohamed et.al. 2608.07434 null
2026-08-07 Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits Elena Dumitrescu et.al. 2608.07430 null
2026-08-07 A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy Bhavika Jalli et.al. 2608.07427 null
2026-08-07 CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing Yan Zhou et.al. 2608.07424 null
2026-08-07 Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Ruochen Jin et.al. 2608.07419 null
2026-08-07 ResidencyRL: Reinforcement Learning in Simulated Clinical Environments Valentin Liévin et.al. 2608.07418 null
2026-08-07 GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks Rodrigo Ferreira Rodrigues et.al. 2608.07411 null
2026-08-07 A Domain-Specific Harness for End-to-End Automation of Optimization Research Heechang Kim et.al. 2608.07407 null
2026-08-07 FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings Sasan Mansouri et.al. 2608.07400 null
2026-08-07 Banach lattices and phase retrieval: A case study for the use of AI in mathematics Jaume de Dios Pont et.al. 2608.07396 null
2026-08-07 PACE: Primitive-Aware Code Evolution for Automated Algorithm Design Zhuoliang Xie et.al. 2608.07395 null
2026-08-07 LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer’s Disease Screening Xin Wang et.al. 2608.07378 null
2026-08-06 Learning When to Trust via Selective Context Preference Optimization Xian Sun et.al. 2608.06377 null
2026-08-06 DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Junfeng Li et.al. 2608.06374 null
2026-08-06 The Bitter Lesson of Tool Calling Ishan Patel et.al. 2608.06370 null
2026-08-06 Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering Soorya Ram Shimgekar et.al. 2608.06366 null
2026-08-06 The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Sarvesh Baskar et.al. 2608.06361 null
2026-08-06 RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer Xinye Wang et.al. 2608.06347 null
2026-08-06 Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents Tao Wang et.al. 2608.06312 null
2026-08-06 On-Policy Self-Distillation without Any Supervision Yijiang Li et.al. 2608.06296 null
2026-08-06 QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction Mutasim Fuad Sarker et.al. 2608.06294 null
2026-08-06 NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering Jonas Gann et.al. 2608.06292 null
2026-08-06 Automatic Translation of Unstructured Requirements into Linear Temporal Logic through Large Language Models Alexandra Newcomb et.al. 2608.06287 null
2026-08-06 MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction Dohyun Ku et.al. 2608.06253 null
2026-08-06 A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance Fardin Afdideh et.al. 2608.06246 null
2026-08-06 DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models ZhiYan Hou et.al. 2608.06243 null
2026-08-06 TS-RAG: Retrieval Augmented Generation for Time Series Forecasting Yixiong Xiao et.al. 2608.06223 null
2026-08-06 What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Ro Encarnación et.al. 2608.06202 null
2026-08-06 Using LLMs to Detect Growth in Computational Thinking in Introductory Physics Sean Savage et.al. 2608.06200 null
2026-08-06 EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Zishan Xu et.al. 2608.06197 null
2026-08-06 Routing LLM Inference to the Cleanest Grid in Real Time Aleks Bernhard et.al. 2608.06188 null
2026-08-06 SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models Hoda Fakharzadehjahromy et.al. 2608.06179 null
2026-08-05 Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning Boxiu Li et.al. 2608.05144 null
2026-08-05 OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling Indraneil Paul et.al. 2608.05141 null
2026-08-05 Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains Ayoub Kirouane et.al. 2608.05138 null
2026-08-05 SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding Yue Zhang et.al. 2608.05137 null
2026-08-05 OPD-V: Visual On-Policy Self-Distillation with Modality Balance Aniri et.al. 2608.05131 null
2026-08-05 Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models Yuezhang Peng et.al. 2608.05126 null
2026-08-05 Chained Recursive Language Models for Multi-Iteration Reasoning Purbesh Mitra et.al. 2608.05124 null
2026-08-05 DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery Roberto Aliaga Medina et.al. 2608.05120 null
2026-08-05 BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning Sajib Hossain et.al. 2608.05104 null
2026-08-05 Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications? Réemi Andrieu et.al. 2608.05097 null
2026-08-05 MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning Tongle Wu et.al. 2608.05088 null
2026-08-05 Item Response Theory for AI Safety Joshua Fonseca Rivera et.al. 2608.05086 null
2026-08-05 Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Zheyuan Zhang et.al. 2608.05080 null
2026-08-05 Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models Jianru Shen et.al. 2608.05064 null
2026-08-05 Hardware Design and Security in the Era of Chiplets and LLMs Johann Knechtel et.al. 2608.05063 null
2026-08-05 The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations Sandra C. Sandoval et.al. 2608.05050 null
2026-08-05 OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing Chenxuan Miao et.al. 2608.05049 null
2026-08-05 Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning Yuxuan Huang et.al. 2608.05045 null
2026-08-05 BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation Peiyan Li et.al. 2608.05042 null
2026-08-05 Private Direct Preference Optimization for LLM Alignment Yangfan Jiang et.al. 2608.05040 null
2026-08-04 ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Yang Yang et.al. 2608.04010 null
2026-08-04 SocietyBench: Forecasting Counterfactual Social-World Evolution Zhenran Wang et.al. 2608.04009 null
2026-08-04 WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament Zhenran Wang et.al. 2608.04008 null
2026-08-04 Semantic Bundling: Interactive Node and Edge Bundling to Simplify Knowledge Graphs using Large Language Models Adam Coscia et.al. 2608.04002 null
2026-08-04 Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Mohsen Hariri et.al. 2608.04001 null
2026-08-04 Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation Junhao Chen et.al. 2608.03999 null
2026-08-04 Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss? Hailong Jiang et.al. 2608.03983 null
2026-08-04 ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning Jinhe Bi et.al. 2608.03972 null
2026-08-04 Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations Zizhao Hu et.al. 2608.03970 null
2026-08-04 HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification Salah Eddine Bekhouche et.al. 2608.03966 null
2026-08-04 Separating quantum circuits from classical LLMs Srinivasan Arunachalam et.al. 2608.03962 null
2026-08-04 Interpretable Adaptive Sampling for LLM Test-Time Scaling Mobina Kashaniyan et.al. 2608.03961 null
2026-08-04 TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring Dongjie Yang et.al. 2608.03952 null
2026-08-04 Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility Jo-Ku Cheng et.al. 2608.03930 null
2026-08-04 The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections Marco Giunti et.al. 2608.03921 null
2026-08-04 Equivariant Music Transformer Zixun Guo et.al. 2608.03920 null
2026-08-04 When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding Ke Li et.al. 2608.03918 null
2026-08-04 UniEvo-RS: Omni-Prompt Unified Remote Sensing Segmentation with Representative Exemplar-Driven Prototype Evolution Kunquan Zhang et.al. 2608.03911 null
2026-08-04 ATLAS: Learning to Recommend Across Unseen Domains Pervez Shaik et.al. 2608.03899 null
2026-08-04 Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition Michal Mráz et.al. 2608.03892 null
2026-08-04 MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models Tong Ling et.al. 2608.03769 null
2026-08-04 TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding Qingxi Du et.al. 2608.03763 null
2026-08-04 Delay Attacks on the German Smart Metering Infrastructure: A Security Analysis of CLS Channel Timing Constraints Fabio Stoll et.al. 2608.03751 null
2026-08-04 Risky Business: Measuring The Faithfulness-Safety Tension Dominik Meier et.al. 2608.03745 null
2026-08-04 Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems Sebastián Andrés Cajas Ordóñez et.al. 2608.03744 null
2026-08-04 Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement Chunyang Jiang et.al. 2608.03733 null
2026-08-04 GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models Yujia Hu et.al. 2608.03729 null
2026-08-04 Detecting Hallucinations and Recovering Verified Answers in Arabic Islamic Question Answering Khaled Ziani et.al. 2608.03720 null
2026-08-04 Attention is Case-Sensitive Maximilian Dillitzer et.al. 2608.03711 null
2026-08-04 Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation Khai-Nguyen Nguyen et.al. 2608.03691 null
2026-08-04 LiveEvalBench: Toward Open-World Evaluation for Web Generation Yiyao Wang et.al. 2608.03689 null
2026-08-04 TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training Lingyun Zhang et.al. 2608.03676 null
2026-08-04 CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning Jian Zhang et.al. 2608.03673 null
2026-08-04 Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training Yibei Liu et.al. 2608.03660 null
2026-08-04 How Closely Do LLM Reviews Align with Human Peer Review? Abraham Camelo-Guerrero et.al. 2608.03659 null
2026-08-04 AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery Zhijing Hu et.al. 2608.03653 null
2026-08-04 Group Perspective Matters: Regulating Debate Relationships Can Mitigate Blind Conformity in Multi-Agent Debate Hao Wu et.al. 2608.03648 null
2026-08-04 MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble Haoze Lv et.al. 2608.03636 null
2026-08-04 When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation Yinuo Jiang et.al. 2608.03632 null
2026-08-04 Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection Razieh Chalehchaleh et.al. 2608.03627 null
2026-08-03 AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Jiajun Liang et.al. 2608.02602 null
2026-08-03 Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework Junjie Yin et.al. 2608.02599 null
2026-08-03 onepot-Bench 0: towards lab-aware in silico chemistry benchmarks Brandon Wang et.al. 2608.02595 null
2026-08-03 GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning Zhaoxin Yu et.al. 2608.02585 null
2026-08-03 ACEM: A Cost Estimation Model for Agentic Software Engineering Mohammad El-Ramly et.al. 2608.02582 null
2026-08-03 Pairwise-Independent Dithering for Single-Stage Hadamard Quantization Honghao Lin et.al. 2608.02564 null
2026-08-03 Structured Memory for Edge Language Models: Persistent Context and Corpus Retrieval via O(1) SSM State Injection Anusha Madan Gopal et.al. 2608.02560 null
2026-08-03 MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs Saman Sarker Joy et.al. 2608.02520 null
2026-08-03 Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions Nicole Mitchell et.al. 2608.02491 null
2026-08-03 CTRAG: An In-Context Retrieval-based Framework for Automated Compliance Checking using LLMs Muhammad Roman et.al. 2608.02472 null
2026-08-03 Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment Vishwajeet Shivaji Hogale et.al. 2608.02470 null
2026-08-03 MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving Ambarish Govindarajulu Kaliamurthi et.al. 2608.02449 null
2026-08-03 Advancing Relevance Measurement with Vision-Language Models for Web-Scale Search Han Wang et.al. 2608.02446 null
2026-08-03 Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks Xuan Ren et.al. 2608.02442 null
2026-08-03 Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning Yiran Gao et.al. 2608.02422 null
2026-08-03 WIP: Chat-Debugging: Large Language Model as a Hardware Debugging Assistant Andrew Ash et.al. 2608.02420 null
2026-08-03 Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes Nan Chen et.al. 2608.02415 null
2026-08-03 Why Large Language Models Fail at Tabular Prediction Marta Garnelo et.al. 2608.02412 null
2026-08-03 MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models Victor Ojewale et.al. 2608.02409 null
2026-08-03 Antares: Foundation Models for Agentic Vulnerability Localization Supriti Vijay et.al. 2608.02407 null
2026-08-02 ReACT-CLIP: Response-Aware Test-Time Defense for Vision–Language Models Hashmat Shadab Malik et.al. 2608.01067 null
2026-08-02 One Query, Many Scales: Sparse Mixture-of-Experts for Efficient Hierarchical Cross-View Geo-Localization Ruijie Fan et.al. 2608.01060 null
2026-08-02 Control Under Compression: Reliability Frontiers for Tool-Using Agents Yinghan Hou et.al. 2608.01056 null
2026-08-02 Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception Xinheng Han et.al. 2608.01055 null
2026-08-02 DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text Muhammad Yousaf Rehman et.al. 2608.01046 null
2026-08-02 Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks Haoyu Zhang et.al. 2608.01043 null
2026-08-02 Opt.Gear Technical Report Juneyoung Park et.al. 2608.01034 null
2026-08-02 CallScreenBench: Benchmarking On-Device Models as Phone Secretaries Simiao Ren et.al. 2608.01033 null
2026-08-02 Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking Timothee Mickus et.al. 2608.01021 null
2026-08-02 Inverting the Hidden: Unveiling Multimodal Privacy Leakage in Collaborative LVLM Inference Shuaifan Jin et.al. 2608.01020 null
2026-08-02 Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy Kaike Ping et.al. 2608.01017 null
2026-08-02 Cloud-ScPO: Hidden-State Geometry for Semi-Supervised Preference Optimization in LLM Reasoning Yuzhou Liu et.al. 2608.01014 null
2026-08-02 MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models Ofir Ben Shoham et.al. 2608.01012 null
2026-08-02 Toward Fine-Grained Forgetting:Attribute Unlearning for Multimodal Large Language Models Junkai Lin et.al. 2608.01008 null
2026-08-02 Hierarchical Solomonoff Induction: An Unbounded Machine Learning Model Nathan Young et.al. 2608.01005 null
2026-08-02 Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms Tezan Sahu et.al. 2608.01004 null
2026-08-02 Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Wenhui Chen et.al. 2608.01000 null
2026-08-02 SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling Shrenil Shaun Sharma et.al. 2608.00991 null
2026-08-02 Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel Yohei Nakajima et.al. 2608.00979 null
2026-08-02 Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models Myeongkyun Kang et.al. 2608.00976 null
2026-07-31 CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding Wenxin Tang et.al. 2607.29637 null
2026-07-31 When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Luca Viano et.al. 2607.29617 null
2026-07-31 CWEEP: A Lexical Static Analysis Framework for CWE Early Prevention Bryan Kwan et.al. 2607.29604 null
2026-07-31 FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models Jeffrey M. Girard et.al. 2607.29602 null
2026-07-31 The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs Jiajia Tang et.al. 2607.29601 null
2026-07-31 Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks Rupak Sarkar et.al. 2607.29585 null
2026-07-31 DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat Ismayil Ismayilov et.al. 2607.29577 null
2026-07-31 SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Pol G. Recasens et.al. 2607.29575 null
2026-07-31 MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models Boxiao Wang et.al. 2607.29561 null
2026-07-31 AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction Rui Zou et.al. 2607.29549 null
2026-07-31 MoRoute: Dynamic Routing for In-Context Multimodal Video Generation Chong Gao et.al. 2607.29545 null
2026-07-31 ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation Gaetano Perrone et.al. 2607.29539 null
2026-07-31 Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction? Xining Xun et.al. 2607.29484 null
2026-07-31 Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification Stanislaw Janik et.al. 2607.29463 null
2026-07-31 MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification Sebastian Doerrich et.al. 2607.29462 null
2026-07-31 QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models Xiang Chen et.al. 2607.29445 null
2026-07-31 Know It, Act on It: Investigating Memory Utilization in LLM Personalization Zhaoxin Feng et.al. 2607.29433 null
2026-07-31 ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models Penglin Zhu et.al. 2607.29431 null
2026-07-31 Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs Mingyu Wang et.al. 2607.29412 null
2026-07-31 Bridging the Question-Answer Gap in Retrieval-Augmented Generation: Hypothetical Prompt Embeddings Domen Vake et.al. 2607.29402 null
2026-07-30 ReToken: One Token to Improve Vision-Language Models for Visual Retrieval Yao Xiao et.al. 2607.28627 null
2026-07-30 AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Xiangning Lin et.al. 2607.28617 null
2026-07-30 Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Chongjian Ge et.al. 2607.28611 null
2026-07-30 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Qiushi Sun et.al. 2607.28609 null
2026-07-30 Inducing language models to assert their own consciousness restores human beliefs and values Junsol Kim et.al. 2607.28607 null
2026-07-30 Beacon: Knowing When and How to Perform Agentic Visual Reasoning Qixun Wang et.al. 2607.28595 null
2026-07-30 $β$ -OPSD: Deriving with Policy Optimization, Training with Self-Distillation Jiawei Xu et.al. 2607.28582 null
2026-07-30 Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Iliya Mirzaei et.al. 2607.28576 null
2026-07-30 Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating Oscar Heath et.al. 2607.28550 null
2026-07-30 ORCA-bench: How Ready Are Language Model Agents for Oncall? Albert Gong et.al. 2607.28545 null
2026-07-30 ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs Ruman Wang et.al. 2607.28538 null
2026-07-30 MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition Alex Andonian et.al. 2607.28532 null
2026-07-30 CoGate: Confidence-Gated Co-Decoding for Secure Code Generation Minghao Hu et.al. 2607.28529 null
2026-07-30 AI systems and the reproduction of (standard) language ideologies in World Englishes Kingsley Ugwuanyi et.al. 2607.28528 null
2026-07-30 MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems Mao-xun Huang et.al. 2607.28527 null
2026-07-30 InfoOps Bench: A live information operations safety benchmark Dorian Quelle et.al. 2607.28503 null
2026-07-30 A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks Ngoc Thai Le et.al. 2607.28481 null
2026-07-30 Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning Zheng Wu et.al. 2607.28478 null
2026-07-30 A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports Jia Yu et.al. 2607.28466 null
2026-07-30 Can Vision-Language Models Reason about AI Edits in Images? Darsha Udayanga et.al. 2607.28464 null
2026-07-29 TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Hengyi Xie et.al. 2607.27205 null
2026-07-29 GraphQAG: A Knowledge-Graph-Guided Visual Analytics Framework for Question-Answer Pairs Generation Yize Li et.al. 2607.27182 null
2026-07-29 HumanCLAW: Can Vision-Language Models Act Through a Body? Siyao Li et.al. 2607.27180 null
2026-07-29 Improving Item Discoverability in e-Commerce Search via Related Intent Generation Ji Xin et.al. 2607.27172 null
2026-07-29 OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding Jingbo Zhou et.al. 2607.27155 null
2026-07-29 MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis Yihao Chen et.al. 2607.27146 null
2026-07-29 Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications Piyush Jain et.al. 2607.27145 null
2026-07-29 Linguistic Monoculture in LLM-Assisted Language Use Suhas Thejaswi et.al. 2607.27134 null
2026-07-29 AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching Yiping Song et.al. 2607.27130 null
2026-07-29 Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs Itbaan Safwan et.al. 2607.27122 null
2026-07-29 Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection Hao Tan et.al. 2607.27113 null
2026-07-29 MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning Weijie Wu et.al. 2607.27109 null
2026-07-29 Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment Yuxuan Cai et.al. 2607.27100 null
2026-07-29 Sky sphere representation in language models Aleksandr Berdnikov et.al. 2607.27092 null
2026-07-29 InferScale: GPU-Native KV Injection for Personalized LLM Serving Peter Li et.al. 2607.27090 null
2026-07-29 SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context Zihan Deng et.al. 2607.27084 null
2026-07-29 On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment Yongjian Guo et.al. 2607.27081 null
2026-07-29 Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data Lingyang Zeng et.al. 2607.27056 null
2026-07-29 CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation Fengming Yu et.al. 2607.27054 null
2026-07-29 Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making Jiayuan Di et.al. 2607.27022 null
2026-07-29 Qwen-Audio-3.0-Gen-Preview Technical Report Junyu Dai et.al. 2607.27011 null
2026-07-29 IMFuse: Instance-Aware Multi-Layer Fusion for LLM-Enhanced Sequential Recommendation Yuheng Zheng et.al. 2607.27002 null
2026-07-29 AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents Ruoyu Wang et.al. 2607.26998 null
2026-07-29 How Developers Experience Debugging Unfamiliar Codebases with Code Tours Generated and Evaluated by Local LLMs Balfroid Martin et.al. 2607.26987 null
2026-07-29 Using large language models to probe the limits of atom-centered structural descriptors Michelangelo Domina et.al. 2607.26984 null
2026-07-29 OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment Seonglae Cho et.al. 2607.26981 null
2026-07-29 Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning? Arnav Hiray et.al. 2607.26952 null
2026-07-29 Progressive Multimodal Alignment for Continual Instruction Tuning Duzhen Zhang et.al. 2607.26947 null
2026-07-29 Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States Weiyi Kong et.al. 2607.26929 null
2026-07-29 Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models Ashish Prajapati et.al. 2607.26922 null
2026-07-28 Spend Experts Where You Are Unsure: Confidence-Adaptive Routing for Mixture-of-Experts LoRA Tom Saliencro et.al. 2607.26052 null
2026-07-28 VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening Syed Mhamudul Hasan et.al. 2607.26042 null
2026-07-28 LLM4OSC: Profile-Bound Natural Language Control with Deterministic Validation for Open Sound Control Yuan-Yi Fan et.al. 2607.26024 null
2026-07-28 CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer Ankang Yang et.al. 2607.26023 null
2026-07-28 Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do Zandi Eberstadt et.al. 2607.26015 null
2026-07-28 RepoReasoner: Evaluating Repository-Level Code Reasoning Ability of Long-Context Language Models Yanlin Wang et.al. 2607.25996 null
2026-07-28 Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches? Farooq Shaikh et.al. 2607.25995 null
2026-07-28 Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Fengxiang Wang et.al. 2607.25993 null
2026-07-28 Untangling Co-Drift: Proactive Multi-Intent Failure Prediction and Root-Cause Disambiguation for Self-Driving Networks Md. Kamrul Hossain et.al. 2607.25989 null
2026-07-28 \textsc{IH-Benchmark}: A Conflict-Centered Benchmark for Instruction-Hierarchy Robustness in LLM Applications Conor McCauley et.al. 2607.25987 null
2026-07-28 Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts? Mike Thelwall et.al. 2607.25965 null
2026-07-28 Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition Podakanti Satyajith Chary et.al. 2607.25961 null
2026-07-28 Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation Jintao Xu et.al. 2607.25956 null
2026-07-28 MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Mingqiao Ye et.al. 2607.25948 null
2026-07-28 A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series Frank Nie et.al. 2607.25947 null
2026-07-28 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Rui Yang et.al. 2607.25933 null
2026-07-28 Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA Carlos Celemin et.al. 2607.25921 null
2026-07-28 Penelope: Localized Latent Recurrence for Efficient Structured Reasoning Yutong Chen et.al. 2607.25915 null
2026-07-28 AnnoBench: A Benchmark for Visualization Annotation Generation Md Rahat-uz-Zaman et.al. 2607.25911 null
2026-07-28 Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models Deepanshu Mody et.al. 2607.25907 null
2026-07-27 ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding Hangjie Yuan et.al. 2607.24743 null
2026-07-27 KANEx: Translating Kolmogorov-Arnold Networks’ Interpretability to Medical Explainability Krithi Shailya et.al. 2607.24730 null
2026-07-27 DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data Zhen Huang et.al. 2607.24717 null
2026-07-27 Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures Fabian Kreppel et.al. 2607.24714 null
2026-07-27 ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams Ali Ansari et.al. 2607.24707 null
2026-07-27 Beyond Scale and Generation: Understanding Language Model-based Entity Matching Zeyu Zhang et.al. 2607.24688 null
2026-07-27 Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating Maruthi Vemula et.al. 2607.24667 null
2026-07-27 MMOE: Modernizing Diffusion Transformers with Efficient Expert Design Yanhao Jia et.al. 2607.24665 null
2026-07-27 Kimi K3: Open Frontier Intelligence Kimi Team et.al. 2607.24653 null
2026-07-27 Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels Zhuchenyang Liu et.al. 2607.24651 null
2026-07-27 Reason-Mediated Behavioral Models for Auditing LLM Social Simulators Atharva Pandey et.al. 2607.24649 null
2026-07-27 LaRec: Unleashing LLM-based Latent Reasoning for Generative Recommendation Yu Xia et.al. 2607.24617 null
2026-07-27 Test-Time Adaptation via Dual Distillation for Videos Under Severe Distribution Shifts André Sacilotti et.al. 2607.24611 null
2026-07-27 CAP-DO: Learned Contextual Action Proposals for Certified Double-Oracle Solving Across Related Zero-Sum Games Mu Wang et.al. 2607.24610 null
2026-07-27 Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review Zhenhan Gao et.al. 2607.24601 null
2026-07-27 SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents Hang Ni et.al. 2607.24588 null
2026-07-27 D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models Bianca Raimondi et.al. 2607.24586 null
2026-07-27 From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference Darina Gold et.al. 2607.24585 null
2026-07-27 CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding Jinlong Yang et.al. 2607.24582 null
2026-07-27 LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports Jonas Schröder et.al. 2607.24573 null
2026-07-26 Zing: Social Mind for LLMs Zing Team et.al. 2607.23740 null
2026-07-26 Separating Clicks from Baits: Using Large Language Models to Detect Misleading YouTube Thumbnails Wajiha Naveed et.al. 2607.23739 null
2026-07-26 E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios Weihuang Zheng et.al. 2607.23722 null
2026-07-26 Formally Verified Synthesizable Floating-Point Data Types in ARCH HDL Shuqing Zhao et.al. 2607.23715 null
2026-07-26 The Intruder Threshold: A Spectral Law for LoRA Fine-Tuning Peng Xie et.al. 2607.23711 null
2026-07-26 The Illusion of Secure LLM Code: Closing the Security Gap via Iterative Reprompting Ishpuneet Singh et.al. 2607.23710 null
2026-07-26 LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratories Haobo Wang et.al. 2607.23704 null
2026-07-26 Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization Haizhou Ge et.al. 2607.23702 null
2026-07-26 Offline-Online Curriculum RL for Multimodal Reasoning Wendi Deng et.al. 2607.23700 null
2026-07-26 Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems Mingzhou Fan et.al. 2607.23678 null
2026-07-26 Multi-level Code Optimization via Mixture of Prompts Yun Peng et.al. 2607.23665 null
2026-07-26 EmoTrace: An Emotion Trajectory-Centered Framework for Psychological Support Dialogue Generation Kaitong Weng et.al. 2607.23648 null
2026-07-26 CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation Gengyu Zhan et.al. 2607.23647 null
2026-07-26 PathSelect: Sequential Token Selection for Whole Slide Pathology Jingzhi Chen et.al. 2607.23631 null
2026-07-26 DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory Xingyang Yu et.al. 2607.23614 null
2026-07-26 MS-GPT: Rethinking MS/MS De Novo Structure Elucidation as Spectrum-Induced Posterior Querying of a Molecule-Language Model Xin Zhao et.al. 2607.23607 null
2026-07-26 Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning Wenxuan Zhang et.al. 2607.23605 null
2026-07-26 ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion Guo Yurong et.al. 2607.23600 null
2026-07-26 HiTMS: A High-Throughput Multi-Stream Linguistic Steganography Framework Ruiyi Yan et.al. 2607.23597 null
2026-07-26 Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration Fatema Tuj Johora Faria et.al. 2607.23538 null
2026-07-24 Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science Davide Scarso et.al. 2607.22513 null
2026-07-24 CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference Jiyuan Tan et.al. 2607.22511 null
2026-07-24 MineValiCoder: Reliable Code Generation with Test Case Quality Mining and Bipartite Graph-Based Mutual Validation Zhen Zhao et.al. 2607.22471 null
2026-07-24 TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI Ritik Raj et.al. 2607.22465 null
2026-07-24 Vibe Coding: An Experiment with Test-Driven Development Moritz Mock et.al. 2607.22406 null
2026-07-24 A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation Fin Gentzen et.al. 2607.22400 null
2026-07-24 SceneActBench: Can Agents Act on the 3D Scenes They See? Yifei Zhao et.al. 2607.22393 null
2026-07-24 HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding Chao Fang et.al. 2607.22389 null
2026-07-24 Agentic Root Cause Analysis through Evidence-Grounded Reasoning Amaury Wei et.al. 2607.22385 null
2026-07-24 A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books Varun Ghat Ravikumar et.al. 2607.22376 null
2026-07-24 IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation Varun Gumma et.al. 2607.22375 null
2026-07-24 Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education Stephan Vonschallen et.al. 2607.22345 null
2026-07-24 Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization Hao Wang et.al. 2607.22334 null
2026-07-24 Towards Trustworthy and Cost-Efficient Data Integration: From Naïve RAG to Agentic RAG Chuangtao Ma et.al. 2607.22319 null
2026-07-24 Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs Zhangzhi Xiong et.al. 2607.22314 null
2026-07-24 Biomedical Machine Translation for Low-Resource Arabic-Script Languages via Cross-Lingual Transfer and LoRA Adapter Merging Abdullah Alabdullah et.al. 2607.22300 null
2026-07-24 RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding Jianqin Liu et.al. 2607.22293 null
2026-07-24 IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning Wei Zhang et.al. 2607.22251 null
2026-07-24 Offline Vision-Language Navigation with Geometric Goal Localization for Outdoor Environments Ali Salmasi et.al. 2607.22226 null
2026-07-24 Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity Pengzhao Lyu et.al. 2607.22218 null
2026-07-23 3D-Aware VLMs with Implicit and Explicit Geometries Wenhao Li et.al. 2607.21595 null
2026-07-23 Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning Rogerio Guimaraes et.al. 2607.21591 null
2026-07-23 Surprisal Theory is Tautological (without Rational Grounding) Ryan Cotterell et.al. 2607.21574 null
2026-07-23 MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education Qian Wu et.al. 2607.21570 null
2026-07-23 Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning Baihui Wang et.al. 2607.21558 null
2026-07-23 MIRROR: Learning from the Other View for Multi-Modal Reasoning Wen Ye et.al. 2607.21552 null
2026-07-23 X $^3$ -OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Dongjie Fu et.al. 2607.21550 null
2026-07-23 From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs Kaiwen Zhang et.al. 2607.21530 null
2026-07-23 Diffusion Language Model for Recommendation Chengyi Liu et.al. 2607.21519 null
2026-07-23 Improved lower bounds for the Shannon capacity of odd cycles Nathaniel Itty et.al. 2607.21517 null
2026-07-23 Transparent by Design, Usable in Practice? A Formative Usability Study of a Conversational Product Advisor Kevin Schott et.al. 2607.21513 null
2026-07-23 Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it Federico Boggia et.al. 2607.21498 null
2026-07-23 Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models Yingchao Huang et.al. 2607.21496 null
2026-07-23 What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations Piotr Wilam et.al. 2607.21491 null
2026-07-23 Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks Mack Nixon et.al. 2607.21482 null
2026-07-23 Thinkink: 2D Spatial Ink-native Interaction with LLMs Mohammad Hasan Payandeh et.al. 2607.21468 null
2026-07-23 AREX: Towards a Recursively Self-Improving Agent for Deep Research Shuqi Lu et.al. 2607.21461 null
2026-07-23 Detecting LLM-Generated Tokens in Human–LLM Coauthored Text Yangjun Lu et.al. 2607.21458 null
2026-07-23 Test-Time Scaling via Error Localization Rajiv Shailesh Chitale et.al. 2607.21453 null
2026-07-23 When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs Anna Mosolova et.al. 2607.21445 null
2026-07-21 Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning Lizhe Fang et.al. 2607.19345 null
2026-07-21 Agents in the Wild: Where Research Meets Deployment Grace Hui Yang et.al. 2607.19336 null
2026-07-21 ISO: An RLVR-Native Optimization Stack Hanqing Zhu et.al. 2607.19331 null
2026-07-21 Selective State-Space Adaptation and Retrieval for Language Model Reasoning Atahan Dokme et.al. 2607.19326 null
2026-07-21 ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D Lena Libon et.al. 2607.19321 null
2026-07-21 Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field Dylan M. Diaz et.al. 2607.19316 null
2026-07-21 Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Priyank Agrawal et.al. 2607.19313 null
2026-07-21 EmbeddedKittens: An Evaluation of Code Embeddings for Scratch Benedikt Fein et.al. 2607.19291 null
2026-07-21 No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation Feinan Cheng et.al. 2607.19288 null
2026-07-21 PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image Dankai Liao et.al. 2607.19261 null
2026-07-21 Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks Guy Stephane Waffo Dzuyo et.al. 2607.19259 null
2026-07-21 Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models Netanel Eliav et.al. 2607.19257 null
2026-07-21 Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks Rawaa Alatrash et.al. 2607.19253 null
2026-07-21 Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs Alexander Manev et.al. 2607.19243 null
2026-07-21 MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings Ziyi Wang et.al. 2607.19235 null
2026-07-21 The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation Michael Jungo et.al. 2607.19226 null
2026-07-21 AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters Yu-Yang Qian et.al. 2607.19223 null
2026-07-21 Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards Xuefeng Jin et.al. 2607.19219 null
2026-07-21 HACO: Hedged Agent Computing for Reliable LLM Systems Enhan Li et.al. 2607.19215 null
2026-07-21 Do LLMs Ask the Right Questions? Evaluating GPT-Generated Surveys as Instruments for Measuring Social Attitudes Tina Behzad et.al. 2607.19211 null
2026-07-20 The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric Sheng-Yu Wang et.al. 2607.18237 null
2026-07-20 Patch Policy: Efficient Embodied Control via Dense Visual Representations Gaoyue Zhou et.al. 2607.18236 null
2026-07-20 It’s Not What You Say, It’s How You Say It: Evaluating LLM Responses to Expressions of Belief Kevin Du et.al. 2607.18232 null
2026-07-20 Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs Yi Tang et.al. 2607.18230 null
2026-07-20 PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning Hang Zhang et.al. 2607.18199 null
2026-07-20 OR Else: A Differentiable Trust Region for Policy Optimization Chinmay Rane et.al. 2607.18163 null
2026-07-20 Testing Retrieval-Augmented Generation Systems with Chunk Coverage Jinhan Kim et.al. 2607.18155 null
2026-07-20 LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications Daniela Rojas et.al. 2607.18147 null
2026-07-20 Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints Thomas MacDougall et.al. 2607.18144 null
2026-07-20 Manifold-Constrained Hyper-Connections for Parameter-Efficient Finetuning Valentijn Oldenburg et.al. 2607.18130 null
2026-07-20 SGA: Plug&Play Geometric Verification for Educational Video Synthesis Lopez Jhon et.al. 2607.18116 null
2026-07-20 Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering Sheldon Yu et.al. 2607.18100 null
2026-07-20 VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval Yu-Chien Tang et.al. 2607.18098 null
2026-07-20 Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs Koyar Afrasyab et.al. 2607.18086 null
2026-07-20 WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting Zhaokai Wang et.al. 2607.18084 null
2026-07-20 SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs Huzaifa Shaaban Kabakibo et.al. 2607.18081 null
2026-07-20 Human Grounded Evaluation of Large Language Models for Optical Network Automation Kiarash Rezaei et.al. 2607.18068 null
2026-07-20 Pancasila-Dilemmas: Evaluating Large Language Models on Indonesian Human Value Dilemmas Grounded in Pancasila Supryadi et.al. 2607.18066 null
2026-07-20 An Early Warning of Emerging Biosecurity Risks in Frontier LLMs Zhida He et.al. 2607.18056 null
2026-07-20 SEE: Structure-aware Exploring \& Exploiting for Long-horizon GUI Agent Trajectory Synthesis Zhuohang Fan et.al. 2607.18046 null
2026-07-17 Knowing the Self, Understanding the World: A Dual-Cognition Benchmark for UAV Spatio-temporal Reasoning with MLLMs Like Liu et.al. 2607.16193 null
2026-07-17 PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization Yuchen Yang et.al. 2607.16184 null
2026-07-17 Vision-Language Assistant for Emotional Reactions to Risky Driving Harine Choi et.al. 2607.16181 null
2026-07-17 Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities Md Erfan et.al. 2607.16175 null
2026-07-17 An Exam for Active Observers Jiarui Zhang et.al. 2607.16165 null
2026-07-17 CADAQUES: A Cost-Aware Dual Architecture for Query-Efficient Autonomous Discovery Jorge Bravo-Abad et.al. 2607.16127 null
2026-07-17 Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content Ingo Ziegler et.al. 2607.16117 null
2026-07-17 Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Sreyan Ghosh et.al. 2607.16107 null
2026-07-17 Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs Maeve Hutchinson et.al. 2607.16105 null
2026-07-17 Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Siyuan Shen et.al. 2607.16100 null
2026-07-17 Understanding Reasoning from Pretraining to Post-Training Jingyan Shen et.al. 2607.16097 null
2026-07-17 How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA Navya Gupta et.al. 2607.16094 null
2026-07-17 What Does It Take to Research with AI? A Rapid Review of Competencies to Train LLM-Literate Researchers Danilo Monteiro Ribeiro et.al. 2607.16083 null
2026-07-17 Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D Haodong Wen et.al. 2607.16072 null
2026-07-17 LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization Mazene Ameur et.al. 2607.16066 null
2026-07-17 Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning Ajay Patel et.al. 2607.16057 null
2026-07-17 Loop the Loopies! Zitian Gao et.al. 2607.16051 null
2026-07-17 Network-Induced Strategic Communication in Opinion Dynamics Hassan Munif et.al. 2607.16036 null
2026-07-17 Revisiting data-driven dynamic security assessment with a tabular foundation model Olayiwola Arowolo et.al. 2607.16031 null
2026-07-17 BayesPO: Bayesian Prompt Optimization via Parallel-Tempered Gradient-Guided Discrete MCMC Junjie Zhou et.al. 2607.16001 null
2026-07-16 Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Patrik Wolf et.al. 2607.15277 null
2026-07-16 Pretraining Data Can Be Poisoned through Computational Propaganda Victoria Graf et.al. 2607.15267 null
2026-07-16 Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents Paul Kassianik et.al. 2607.15263 null
2026-07-16 SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Yuyao Zhang et.al. 2607.15257 null
2026-07-16 HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning Pengcheng Zhou et.al. 2607.15255 null
2026-07-16 Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search Debayan Mukhopadhyay et.al. 2607.15253 null
2026-07-16 ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors Christos Korgialas et.al. 2607.15246 null
2026-07-16 In-Place Tokenizer Expansion for Pre-trained LLMs Jimmy T. H. Smith et.al. 2607.15232 null
2026-07-16 When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space Weimeng Wang et.al. 2607.15218 null
2026-07-16 Symbal: Detecting Systematic Misalignments in Model-Generated Captions Maya Varma et.al. 2607.15216 null
2026-07-16 Expanding the Lexicon of Ge’ez Based African Languages: A Comparative Study of Amharic and Tigrinya Hailay Kidu Teklehaymanot et.al. 2607.15209 null
2026-07-16 Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation Hoang-Loc Cao et.al. 2607.15202 null
2026-07-16 Mask-Aware Policy Gradients for Diffusion Language Models Haran Raajesh et.al. 2607.15200 null
2026-07-16 Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy Patrick Phuoc Do et.al. 2607.15176 null
2026-07-16 Linear representations of grammaticality in neural language models Jane Li et.al. 2607.15175 null
2026-07-16 On-Policy Delta Distillation Byeongho Heo et.al. 2607.15161 null
2026-07-16 Learning in Infinitesimal Non-Compositional Sketches Sridhar Mahadevan et.al. 2607.15107 null
2026-07-16 Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents Dylan Van Mulders et.al. 2607.15095 null
2026-07-16 Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence Haocheng Yang et.al. 2607.15092 null
2026-07-16 DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment Zefeng Wu et.al. 2607.15081 null
2026-07-15 Building Shor’s Algorithm in Lean: An Agentic Formalization of Quantum Attacks on RSA-2048 and P-256 Lei Zhang et.al. 2607.14082 null
2026-07-15 VisualRepair: Dynamic Tool Calling and Region Focusing for Visual Software Issue Repair Jingyu Xiao et.al. 2607.14075 null
2026-07-15 Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models Hefeng Zhou et.al. 2607.14049 null
2026-07-15 LLMs for Qualitative and Mixed-Methods Social Network Analysis Moses Boudourides et.al. 2607.14045 null
2026-07-15 Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation Sanket Badhe et.al. 2607.13987 null
2026-07-15 Screening Is Effective for Visual Recognition Shunya Shimomura et.al. 2607.13983 null
2026-07-15 Music-to-Dance Generation via Atomic Movements Xinhao Cai et.al. 2607.13978 null
2026-07-15 Measuring Sentiment News with Transformer-Based Language Models Maria Saveria Mavillonio et.al. 2607.13968 null
2026-07-15 Pezego-HITL: A policy-grounded large language model architecture for agricultural extension in Ghana Shunbao Li et.al. 2607.13934 null
2026-07-15 SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning Cheng Tang et.al. 2607.13931 null
2026-07-15 NNStar: An end-to-end AI agent for nuclear matter and neutron star physics Yao Ma et.al. 2607.13930 null
2026-07-15 S-squared-VLA: Decoupling Semantic and Spatial Streams in Vision-Language-Action Models for Autonomous Driving Jianguo Yu et.al. 2607.13926 null
2026-07-15 How to Guide LLM Generation: Dual-Surrogate Guided Search for Automated Heuristic Design Yuhan Wang et.al. 2607.13911 null
2026-07-15 High-Order Question Generation in a Multilingual Educational Context Suna-Şeyma Uçar et.al. 2607.13901 null
2026-07-15 AIMO Interpretability Challenge Michal Štefánik et.al. 2607.13899 null
2026-07-15 Fine-Grained Vision-Language Pretraining with Organ-Conditioned Pattern Tokens for CT Understanding Guoliang You et.al. 2607.13892 null
2026-07-15 Experience Memory Graph: One-Shot Error Correction for Agents Wenjun Wang et.al. 2607.13884 null
2026-07-15 Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild Ting Lei et.al. 2607.13881 null
2026-07-15 Towards Enhancing 3D Spatial Reasoning in Medical Multimodal Large Language Models Zhuoyuan Fu et.al. 2607.13860 null
2026-07-15 PROBE: Benchmarking Code Generation in Large Language Models Rodrigo Pato Nogueira et.al. 2607.13820 null
2026-07-14 Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution Junjie Yin et.al. 2607.13034 null
2026-07-14 PalmClaw: A Native On-Device Agent Framework for Mobile Phones Hongru Cai et.al. 2607.13027 null
2026-07-14 Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model Harsha Vardhan Khurdula et.al. 2607.13013 null
2026-07-14 Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs Sen Yang et.al. 2607.12985 null
2026-07-14 FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation Ruoran Xu et.al. 2607.12982 null
2026-07-14 The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context Yanzhe Zhang et.al. 2607.12963 null
2026-07-14 Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition Bruce Coburn et.al. 2607.12911 null
2026-07-14 UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation Yunzhou Li et.al. 2607.12896 null
2026-07-14 Hy-Embodied-VLM-1.0: Efficient Physical-World Agents Ziyi Wang et.al. 2607.12894 null
2026-07-14 Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations Monica Munnangi et.al. 2607.12884 null
2026-07-14 Toward Localizing and Repairing Bias in Transformer Attention Heads Sigma Jahan et.al. 2607.12863 null
2026-07-14 ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation Jhen-Ke Lin et.al. 2607.12857 null
2026-07-14 Knowledgeless Language Models: Suppressing Parametric Recall for Evidence-Grounded Language Modeling Roi Cohen et.al. 2607.12831 null
2026-07-14 Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques Daehoon Gwak et.al. 2607.12829 null
2026-07-14 Breaking Déjà Vu: Independent Auditing of Visual Place Recognition through Vision-Language Reasoning Sania Waheed et.al. 2607.12818 null
2026-07-14 Visual Access Boundaries in Vision-Language Model Reasoning Hiroto Osaka et.al. 2607.12815 null
2026-07-14 The One-Word Census: Answer-Choice Conformity Across 44 Language Models Tapan Parikh et.al. 2607.12796 null
2026-07-14 Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters? Kaiwen Zheng et.al. 2607.12787 null
2026-07-14 CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models Lin Peng et.al. 2607.12786 null
2026-07-14 Learning Mechanistic Reasoning for Chemical Reactions with Large Language Models Xingyu Dang et.al. 2607.12771 null
2026-07-14 EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval Jiashi Lin et.al. 2607.12764 null
2026-07-14 VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression Yupeng Zheng et.al. 2607.12756 null
2026-07-14 Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Hongbo Wang et.al. 2607.12752 null
2026-07-14 Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models Binwen Liu et.al. 2607.12739 null
2026-07-14 LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos Julius Steiglechner et.al. 2607.12733 null
2026-07-14 Bulkhead: Automated Semantic Detection and Remediation of Container Escape Vulnerabilities Qiyuan Fan et.al. 2607.12723 null
2026-07-13 Mixture of Frames Policy: Multi-Frame Action Denoising for Bimanual Mobile Manipulation Dian Wang et.al. 2607.11884 null
2026-07-13 Requential Coding: Pushing the Limits of Model Compression with Self-Generated Training Data Shikai Qiu et.al. 2607.11883 null
2026-07-13 Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks Tiberiu Musat et.al. 2607.11875 null
2026-07-13 A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol Esteban U. Vega Barajas et.al. 2607.11873 null
2026-07-13 Evidence-Backed Video Question Answering Shijie Wang et.al. 2607.11862 null
2026-07-13 Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers? Nishant Aggarwal et.al. 2607.11859 null
2026-07-13 AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Lingkai Kong et.al. 2607.11849 null
2026-07-13 Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding Kerui Chen et.al. 2607.11844 null
2026-07-13 Supporting Reflection in LLM-based Exploratory Search Giulia Di Fede et.al. 2607.11810 null
2026-07-13 Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models Yu-Han Huang et.al. 2607.11801 null
2026-07-13 StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description Seung Hyun Hahm et.al. 2607.11798 null
2026-07-13 How Temperature Shapes Ideological Discourse in Retrieval-Augmented Generation? Elmira Salari et.al. 2607.11783 null
2026-07-13 From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP Michael Rizvi-Martel et.al. 2607.11760 null
2026-07-13 Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement Ziv Ben-Zion et.al. 2607.11749 null
2026-07-13 MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning Ayoung Lee et.al. 2607.11736 null
2026-07-13 STEP: Career-Path Recommendation via Temporal and Educational Trajectory Modeling Iman Johary et.al. 2607.11722 null
2026-07-13 JobHop v2: A Large-Scale Career Trajectory Dataset from Unstructured Resumes Iman Johary et.al. 2607.11715 null
2026-07-13 Production and Perception in LLMs: A Token Probability Approach Anna Marklová et.al. 2607.11703 null
2026-07-13 Qwen-Music Technical Report Jin Xu et.al. 2607.11699 null
2026-07-13 Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction Huan Zhu et.al. 2607.11696 null
2026-07-10 Scalable Visual Pretraining for Language Intelligence Yiming Zhang et.al. 2607.09657 null
2026-07-10 Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models Shravan Murlidaran et.al. 2607.09654 null
2026-07-10 VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents Katherine Swinea et.al. 2607.09653 null
2026-07-10 Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection Cláudio Lúcio do Val Lopes et.al. 2607.09641 null
2026-07-10 LLM for EDA in Front-End Design: Challenges and Opportunities Kangwei Xu et.al. 2607.09616 null
2026-07-10 Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation Kaiji Zhou et.al. 2607.09600 null
2026-07-10 TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models Tianyou Jiang et.al. 2607.09562 null
2026-07-10 The Count Is There, but Misaligned: Understanding and Correcting Counting Failures in VLMs Ahmed Oumar El-Shangiti et.al. 2607.09544 null
2026-07-10 Balancing Usefulness and Naturalness: An LLM-based Curation Pipeline for Code Review Comments Oussama Ben Sghaier et.al. 2607.09524 null
2026-07-10 Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference Junfei Zhan et.al. 2607.09520 null
2026-07-10 Failure as a Process: An Anatomy of CLI Coding Agent Trajectories Xiangxin Zhao et.al. 2607.09510 null
2026-07-10 What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility Filippo Ziliotto et.al. 2607.09503 null
2026-07-10 All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models Pan Li et.al. 2607.09502 null
2026-07-10 Multimodal Reward Hacking in Reinforcement Learning Jiayu Yao et.al. 2607.09492 null
2026-07-10 Neural Collapse Is Forbidden: Information Floors in Language Models Bruno Abrahao et.al. 2607.09487 null
2026-07-10 ProofCouncil: An LLM Agent for Solving Open Mathematical Problems Johannes Schmitt et.al. 2607.09474 null
2026-07-10 Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning Charles Edward Gagnon et.al. 2607.09452 null
2026-07-10 Robustifying Vision-Language Models via Test-Time Prompt Adaptation Xingyu Zhu et.al. 2607.09450 null
2026-07-10 Parameter-Efficient Vision-Language Adaptation with Continuous Metadata Conditioning for Animal Re-Identification Anil Osman Tur et.al. 2607.09443 null
2026-07-10 Test-Time Scaling for Small VLMs on Multilingual Visual MCQ Spiros Baxevanakis et.al. 2607.09438 null
2026-07-09 UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks Zhekai Chen et.al. 2607.08768 null
2026-07-09 OpenCoF: Learning to Reason Through Video Generation Xinyan Chen et.al. 2607.08763 null
2026-07-09 AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding Siddharth Damodharan et.al. 2607.08745 null
2026-07-09 Workflow as Knowledge: Semantic Persistence for LLM-Mediated Workflows Emanuele Quinto et.al. 2607.08740 null
2026-07-09 The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs Baha Rababah et.al. 2607.08734 null
2026-07-09 Validity of LLMs as data annotators: AMALIA on authority Manuel Pita et.al. 2607.08731 null
2026-07-09 Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference Chuning Zhu et.al. 2607.08724 null
2026-07-09 Multimodal Digital Biomarker for Asthma: Complementary Roles of Vocal, Clinical and Demographic Factors Vladimir Despotovic et.al. 2607.08714 null
2026-07-09 How YouTube Frames ChatGPT Use in Education: An Epistemic Network Analysis with Supporting Multimodal Metadata Shayla Sharmin et.al. 2607.08698 null
2026-07-09 A Practical Investigation of Training-free Relaxed Speculative Decoding Guoxuan Xia et.al. 2607.08690 null
2026-07-09 Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models Teng-Ruei Chen et.al. 2607.08665 null
2026-07-09 WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search Xiaoshuai Song et.al. 2607.08662 null
2026-07-09 Secure Decentralized Federated Learning via Gossip and Virtual Voting Amirhossein Taherpour et.al. 2607.08651 null
2026-07-09 UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing Xinlong Zhao et.al. 2607.08646 null
2026-07-09 BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model Compression Yuantian Shao et.al. 2607.08643 null
2026-07-09 The complexities of patient-centred conversational artificial intelligence João Matos et.al. 2607.08625 null
2026-07-09 When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities Weiduo Liao et.al. 2607.08605 null
2026-07-09 Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance Peng Cui et.al. 2607.08602 null
2026-07-09 It Takes a MAESTRO To Prune Bad Experts Palaash Goel et.al. 2607.08601 null
2026-07-09 Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Yiyang Fang et.al. 2607.08572 null
2026-07-08 Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Chen Tang et.al. 2607.07708 null
2026-07-08 Co-LMLM: Continuous-Query Limited Memory Language Models Yair Feldman et.al. 2607.07707 null
2026-07-08 From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization Ying Chang et.al. 2607.07702 null
2026-07-08 Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass Victor Giannakouris et.al. 2607.07696 null
2026-07-08 How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization Xinyi Wu et.al. 2607.07678 null
2026-07-08 Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Shuailei Ma et.al. 2607.07675 null
2026-07-08 MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models Hyunjae Kim et.al. 2607.07673 null
2026-07-08 Does Bielik Know What It Doesn’t Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale Grzegorz Brzezinka et.al. 2607.07670 null
2026-07-08 DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation Jordan Painter et.al. 2607.07669 null
2026-07-08 A hierarchical memory architecture overcomes context limits in long-horizon multi-agent computational modeling Shivendra G. Tewari et.al. 2607.07666 null
2026-07-08 From Custom-Fit to Portable: Bridging the Gap Between Synthesized and Engineered GPU Query Execution Ivan Donchev Kabadzhov et.al. 2607.07632 null
2026-07-08 Future Confidence Distillation in Large Language Models Sahil Kale et.al. 2607.07626 null
2026-07-08 Rethinking Code Performance Benchmarks for LLMs Nhat Minh Le et.al. 2607.07619 null
2026-07-08 User identity conditions moral wrongness ratings in non-reasoning large language models Willem Fourie et.al. 2607.07605 null
2026-07-08 Human and LLM Collaboration for Accelerated Materials Synthesis and Discovery Gregory Bassen et.al. 2607.07604 null
2026-07-08 RubriQ: Rubric-Guided Group Relative Policy Optimization for Constraint-Aware Quantum Circuit Synthesis Ziqing Guo et.al. 2607.07554 null
2026-07-08 Think Big, Search Small: Where Capacity Matters in Hierarchical Search Agents? Qinnan Cai et.al. 2607.07548 null
2026-07-08 A Unified Detection Framework for AI-Related Content and Artifacts Xifeng Zhang et.al. 2607.07527 null
2026-07-08 Creativity from Friction: Human-AI Interaction for Exploratory Structural Design Ricardo Maia Avelino et.al. 2607.07521 null
2026-07-08 Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Zhenyu Hou et.al. 2607.07508 null
2026-07-07 MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation Jiaju Han et.al. 2607.06552 null
2026-07-07 Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Zhenyu Liu et.al. 2607.06540 null
2026-07-07 UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation Zhuofan Zhang et.al. 2607.06537 null
2026-07-07 CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models He Liang et.al. 2607.06534 null
2026-07-07 RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation Sambaran Bandyopadhyay et.al. 2607.06527 null
2026-07-07 DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression Anna Cordoba et.al. 2607.06523 null
2026-07-07 Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment Han-Jun Ko et.al. 2607.06522 null
2026-07-07 Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade Kai Ruan et.al. 2607.06503 null
2026-07-07 AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models Cong Su et.al. 2607.06485 null
2026-07-07 Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities So Hasegawa et.al. 2607.06482 null
2026-07-07 A VLM-Enhanced Framework for Comprehensive Traffic Sign Condition Assessment Integrating Daytime Visual Performance and Nighttime Retroreflectivity Evaluation Linlin Zhang et.al. 2607.06478 null
2026-07-07 WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS Sihang Nie et.al. 2607.06461 null
2026-07-07 From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b Taeyun Roh et.al. 2607.06452 null
2026-07-07 Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Yoav Baron et.al. 2607.06445 null
2026-07-07 HoloCount: A Holistic Visual Counting Benchmark for MLLMs Jinhong Deng et.al. 2607.06420 null
2026-07-07 An Experimental Design Approach to Evaluating Agentic AI’s Autonomous Model Discovery Hao He et.al. 2607.06413 null
2026-07-07 What Images Cannot Say: Language-Guided Olfactory Representation Learning Eleftherios Tsonis et.al. 2607.06402 null
2026-07-07 VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery Jiazi Wang et.al. 2607.06374 null
2026-07-07 TMF-RSE: Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty for Lung Severity Scoring Fadi Abdeladhim Zidi et.al. 2607.06356 null
2026-07-07 Harnessing Code Agents for Automatic Software Verification Shuangxiang Kan et.al. 2607.06341 null
2026-07-06 Weak-to-Strong Generalization via Direct On-Policy Distillation Shiyuan Feng et.al. 2607.05394 null
2026-07-06 SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models Thomas Thebaud et.al. 2607.05365 null
2026-07-06 Faithfulness to Refusal: A Causal Audit of Neuron Selectors Ananth Eswar et.al. 2607.05355 null
2026-07-06 Selective Disclosure Watermarking for Large Language Models Xuyang Chen et.al. 2607.05353 null
2026-07-06 Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis Xianhao Chen et.al. 2607.05348 null
2026-07-06 Data-driven atomistic modelling of hybrid halide perovskite passivation Laura-Bianca Paşca et.al. 2607.05321 null
2026-07-06 How Much is Left? LLMs Linearly Encode Their Remaining Output Length Mohamed Amine Merzouk et.al. 2607.05316 null
2026-07-06 Evaluating and Understanding Model Editing for Medical Vision Language Models Guli Zhu et.al. 2607.05310 null
2026-07-06 ChatImage: Navigating Long-Form LLM Answers through Interactive Images Wencan Jiang et.al. 2607.05290 null
2026-07-06 Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models Jaeyoung Kim et.al. 2607.05268 null
2026-07-06 SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments Suryanarayana Reddy Yarrabothula et.al. 2607.05264 null
2026-07-06 Repurposing CLIP to Localize at Pixel Level Jiaxiang Fang et.al. 2607.05253 null
2026-07-06 MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models Zhi Song et.al. 2607.05238 null
2026-07-06 A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication Avina Nakarmi et.al. 2607.05222 null
2026-07-06 Curated retrieval versus open web search in public AI information services: a coverage-trust trade-off Hafsteinn Einarsson et.al. 2607.05217 null
2026-07-06 Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models Raj Jaiswal et.al. 2607.05199 null
2026-07-06 Latent Programming Horizons in Coding Agents André Silva et.al. 2607.05188 null
2026-07-06 Rethinking On-Policy Self-Distillation for Thinking Models Simran Kaur et.al. 2607.05184 null
2026-07-06 VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving Tianjia Yang et.al. 2607.05180 null
2026-07-06 AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments Zhiheng Xi et.al. 2607.05174 null
2026-07-06 PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection Md. Shakhoyat Rahman Shujon et.al. 2607.04690 null
2026-07-06 URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment Bogdan Zagribelnyy et.al. 2607.04688 null
2026-07-06 ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents Harsh Soni et.al. 2607.04686 null
2026-07-06 Does It Fail to See or Fail to Know? Attributing Errors in Vision-Language Models Khang Nhat Hoang Vo et.al. 2607.04683 null
2026-07-06 Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning Matthew Foutter et.al. 2607.04681 null
2026-07-06 DiCE-CIR: Direct Composition Learning for Efficient Zero-Shot Composed Image Retrieval Gwang-Ho Na et.al. 2607.04665 null
2026-07-06 Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations Samira Hajizadeh et.al. 2607.04645 null
2026-07-06 Wrong Before Right: Late Rescue and Interface Failure in Aligned Language Models Jiaqi Deng et.al. 2607.04640 null
2026-07-06 PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving Pin Tang et.al. 2607.04637 null
2026-07-06 Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning Jiaxing Qi et.al. 2607.04623 null
2026-07-06 CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning Ganesh Pavan Kartikeya Bharadwaj Kolluri et.al. 2607.04619 null
2026-07-06 StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation Veeramanohar Avudaiappan et.al. 2607.04612 null
2026-07-06 RoboVista: Evaluating Vision Language Models for Diverse Robot Applications Shuangyu Xie et.al. 2607.04610 null
2026-07-06 Displacement Preserving Relational Distillation for Robust Medical Segmentation Zhicheng Ding et.al. 2607.04599 null
2026-07-06 TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models Riccardo Renzulli et.al. 2607.04593 null
2026-07-06 Attention Limited Reward Learning Wenqian Xing et.al. 2607.04590 null
2026-07-06 Finetuning Lightweight LLMs for Control Flow Graph Generation Hanyu Zhang et.al. 2607.04582 null
2026-07-06 LLM-Driven CI-CD Workflow Intelligence for Cyber Systems Engineering Bonan Shen et.al. 2607.04579 null
2026-07-06 Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing Bonan Shen et.al. 2607.04572 null
2026-07-06 LLMs for Agentic Home Energy Management Sokipriala Jonah et.al. 2607.04569 null
2026-07-02 Program-as-Weights: A Programming Paradigm for Fuzzy Functions Wentao Zhang et.al. 2607.02512 null
2026-07-02 ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning Yanjun Zhao et.al. 2607.02509 null
2026-07-02 DemoPSD: Disagreement-Modulated Policy Self-Distillation Yunhe Li et.al. 2607.02502 null
2026-07-02 Seek to Segment: Active Perception for Panoramic Referring Segmentation Song Tang et.al. 2607.02497 null
2026-07-02 Towards Robustness against Typographic Attack with Training-free Concept Localization Bohan Liu et.al. 2607.02494 null
2026-07-02 Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning Liyan Tang et.al. 2607.02490 null
2026-07-02 EAGLE-360: Embodied Active Global-to-Local Exploration in 360 $^\circ$ Jingtao Xu et.al. 2607.02479 null
2026-07-02 Will Scaling Improve Social Simulation with LLMs? Caleb Ziems et.al. 2607.02464 null
2026-07-02 Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation Zhuowei Chen et.al. 2607.02460 null
2026-07-02 Language Models as Measurement Apparatus for Culture Kent K. Chang et.al. 2607.02459 null
2026-07-02 When Do LLM Personas Support Visualization Design? A Cross-Model Study of Color Assignment and Chart Choice Shahreen Salim et.al. 2607.02455 null
2026-07-02 AgentsCAD: Automated Design for Manufacturing of FDM Parts via Multi-Agent LLM Reasoning and Geometric Feature Recognition Emmanuel George et.al. 2607.02448 null
2026-07-02 Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach Manuel Alonso-Carracedo et.al. 2607.02432 null
2026-07-02 The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing David Jurgens et.al. 2607.02416 null
2026-07-02 Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments Xianhui Meng et.al. 2607.02407 null
2026-07-02 Show Me Examples: Inferring Visual Concepts from Image Sets Nick Stracke et.al. 2607.02402 null
2026-07-02 Fast Multi-dimensional Refusal Subspaces via RFM-AGOP Thomas Winninger et.al. 2607.02396 null
2026-07-02 WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs Mauricio Fadel Argerich et.al. 2607.02391 null
2026-07-02 DecompRL: Solving Harder Problems by Learning Modular Code Generation Juliette Decugis et.al. 2607.02390 null
2026-07-02 Bringing Agentic Search to Earth Observation Data Discovery Minghan Yu et.al. 2607.02387 null
2026-07-01 Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training Zijian Zhang et.al. 2607.01232 null
2026-07-01 The State-Prediction Separation Hypothesis Giovanni Monea et.al. 2607.01218 null
2026-07-01 Touching and Feeling the Data: A Reusable Software Pipeline for Tactile Statistical Graphs in Accessible Education Lawrence Obiuwevwi et.al. 2607.01214 null
2026-07-01 Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation Shayan Talaei et.al. 2607.01208 null
2026-07-01 Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Hongxing Li et.al. 2607.01191 null
2026-07-01 QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling Michael Y. Li et.al. 2607.01179 null
2026-07-01 Diffusion-GR2: Diffusion Generative Reasoning Re-ranker Zhuoxuan Zhang et.al. 2607.01170 null
2026-07-01 Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity Brett Reynolds et.al. 2607.01153 null
2026-07-01 Emergence of Preferential Attachment and Glass-Ceiling Effects in Autonomous Networks of LLMs Yiming Zhang et.al. 2607.01148 null
2026-07-01 Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains Changguo Jia et.al. 2607.01136 null
2026-07-01 Autonomous Scientific Discovery via Iterative Meta-Reflection Bingchen Zhao et.al. 2607.01131 null
2026-07-01 $\text{Log}_\text{b}$ Quant: Quantizing Language Models in Logarithmic Space Jeremias Bohn et.al. 2607.01127 null
2026-07-01 ZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank Subspaces Xun Dong et.al. 2607.01125 null
2026-07-01 MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Jiale Li et.al. 2607.01117 null
2026-07-01 Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach Md Abu Hanif Shaikh et.al. 2607.01115 null
2026-07-01 CausalMix: Data Mixture as Causal Inference for Language Model Training Zinan Tang et.al. 2607.01104 null
2026-07-01 Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking William Philipp et.al. 2607.01103 null
2026-07-01 LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models Arpita Nema et.al. 2607.01086 null
2026-07-01 Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use Song-Lin Lv et.al. 2607.01084 null
2026-07-01 Where Am I? Semantic Map Grounding via Vision-Language Models for Multi-Modal Localization Suraj Borate et.al. 2607.01079 null
2026-06-30 FaceMoE: Mixture of Experts for Low-Resolution Face Recognition Kartik Narayan et.al. 2606.32040 null
2026-06-30 Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision Zifan Carl Guo et.al. 2606.32038 null
2026-06-30 When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors Yuqing Yang et.al. 2606.32029 null
2026-06-30 SemRF: A Semantic Reference Frame for Residual-Stream Dynamics in Language Models Jian Gu et.al. 2606.32022 null
2026-06-30 CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation Sanghyuk Chun et.al. 2606.32012 null
2026-06-30 Surrogate Fidelity: When Can Open LLMs Explain Closed Ones? Philippe Chlenski et.al. 2606.32008 null
2026-06-30 PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines Sameer Malik et.al. 2606.32004 null
2026-06-30 Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA Ekaterina Alimaskina et.al. 2606.32002 null
2026-06-30 Amplifying Membership Signal Through Chained Regeneration Wojciech Łapacz et.al. 2606.31991 null
2026-06-30 CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts Lianyu Hu et.al. 2606.31986 null
2026-06-30 GR2 Technical Report Yufei Li et.al. 2606.31984 null
2026-06-30 ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs Yuhao Wang et.al. 2606.31982 null
2026-06-30 TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models Shiyi Chen et.al. 2606.31976 null
2026-06-30 MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments Qingyun Liu et.al. 2606.31966 null
2026-06-30 InstanceControl: Controllable Complex Image Generation without Instance Labeling Xiaoyu Liu et.al. 2606.31924 null
2026-06-30 Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action Ben Slater et.al. 2606.31916 null
2026-06-30 CoDex: Learning Compositional Dexterous Functional Manipulation without Demonstrations Bowen Jiang et.al. 2606.31909 null
2026-06-30 Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference Zhaoyang Luo et.al. 2606.31903 null
2026-06-30 Harnessing Textual Refusal Directions for Multimodal Safety Moreno D’Incà et.al. 2606.31876 null
2026-06-30 Explicit Fuzzy Logic in the Feed-Forward Layer: Self-Forgetting Quantifiers Discover Legible Grammatical-Licensing Detectors Mark Oskin et.al. 2606.31845 null
2026-06-29 LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training Shun Lei et.al. 2606.30642 null
2026-06-29 GROW $^2$ : Grounding Which and Where for Robot Tool Use Yuhong Deng et.al. 2606.30632 null
2026-06-29 DOPD: Dual On-policy Distillation Xinlei Yu et.al. 2606.30626 null
2026-06-29 Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent Lei Bai et.al. 2606.30616 null
2026-06-29 PyMETA: A Benchmark Dataset for Hierarchical Student Code Error Classification with Python-Interpreter-Based Labels Chuyue Li et.al. 2606.30610 null
2026-06-29 C $^{2}$ R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders Haoran Jin et.al. 2606.30609 null
2026-06-29 Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection Asif Shahriar et.al. 2606.30587 null
2026-06-29 AI Premium Nicola Borri et.al. 2606.30583 null
2026-06-29 Uncertainty-Aware Generation and Decision-Making Under Ambiguity Nico Daheim et.al. 2606.30578 null
2026-06-29 A Multi-task Mixture of Experts Framework for Malware Classification, Packing Detection, and Family Attribution Jithin S. et.al. 2606.30572 null
2026-06-29 Attractor States Emerge in Multi-Turn LLM Conversations Ting-Wen Ko et.al. 2606.30571 null
2026-06-29 Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task? Shanshan Wang et.al. 2606.30556 null
2026-06-29 Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing Dvir Alsheich et.al. 2606.30555 null
2026-06-29 COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices Yilong Zhao et.al. 2606.30553 null
2026-06-29 Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision Haoyang Li et.al. 2606.30552 null
2026-06-29 Teaching Prompt-Based Programming with LLMs: A 45-Minute Lesson with Guided Practice for End-User Programmers Keith Tran et.al. 2606.30547 null
2026-06-29 Entity Binding Failures in Tool-Augmented Agents Rahul Suresh Babu et.al. 2606.30531 null
2026-06-29 The Illusion of Agentic Complexity in README.md Generation: Evaluating Single-Agent vs. Multi-Agent RAG Systems Abu Saleh et.al. 2606.30524 null
2026-06-29 Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts Bo Wang et.al. 2606.30518 null
2026-06-29 On the Faithfulness of Post-Hoc Concept Bottleneck Models Laines Schmalwasser et.al. 2606.30498 null
2026-06-29 MCP Server Architecture Patterns for LLM-Integrated Applications Carson Rodrigues et.al. 2606.30317 null
2026-06-29 DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information Roland Roller et.al. 2606.30312 null
2026-06-29 ManimAgent: Self-Evolving Multimodal Agents for Visual Education Wenjia Jiang et.al. 2606.30296 null
2026-06-29 VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context Xiaoqian Shen et.al. 2606.30288 null
2026-06-29 Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation Bertram Taetz et.al. 2606.30266 null
2026-06-29 When Is a Draft Accepted? A Theory of Acceptance in Speculative Decoding Aaryam Sharma et.al. 2606.30265 null
2026-06-29 Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation Threats Sebastian Kula et.al. 2606.30259 null
2026-06-29 Grounding LLM Reasoning under Incomplete Graph Evidence Jiaqi Li et.al. 2606.30247 null
2026-06-29 CaresAI at CT-DEB26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Models Leon Hamnett et.al. 2606.30236 null
2026-06-29 From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA Sena Korkut et.al. 2606.30220 null
2026-06-29 Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoning Yinan Zhou et.al. 2606.30217 null
2026-06-29 SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation Filippo Ruffini et.al. 2606.30201 null
2026-06-29 DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning Xinxin Chen et.al. 2606.30189 null
2026-06-29 Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents Yutao Sun et.al. 2606.30185 null
2026-06-29 CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph Chengtao Gan et.al. 2606.30175 null
2026-06-29 Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models Kai Jiang et.al. 2606.30168 null
2026-06-29 End-to-End Abstraction-Based Control with LLM-Enhanced NL-to-LTL Translation Amir Bayat et.al. 2606.30163 null
2026-06-29 Estimating Grammatical Gender Directions in Contextual Embeddings under Controlled and Natural Contexts Huanping Xiao et.al. 2606.30152 null
2026-06-29 AERIS: Aerial-Edge Role-Driven Intelligence at Runtime via Orchestrated Language-Model Swarm Jiabin Lou et.al. 2606.30151 null
2026-06-29 DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks Romain Karpinsky et.al. 2606.30140 null
2026-06-26 Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models Niclas Lietzow et.al. 2606.28273 null
2026-06-26 Agent-Native Immune System: Architecture, Taxonomy, and Engineering Bo Shen et.al. 2606.28270 null
2026-06-26 RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning Yelin Wang et.al. 2606.28266 null
2026-06-26 Push Puppet Networks: Structured Bayesian Pruning Algorithm for Language Model Compression Robert Kubinec et.al. 2606.28251 null
2026-06-26 HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech Sihang Nie et.al. 2606.28249 null
2026-06-26 GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems Xiaocheng Yang et.al. 2606.28187 null
2026-06-26 Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction Chenguang Wang et.al. 2606.28186 null
2026-06-26 LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior Qinhong Zhou et.al. 2606.28182 null
2026-06-26 Tandem Reinforcement Learning with Verifiable Rewards Difan Jiao et.al. 2606.28166 null
2026-06-26 EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography Darya Taratynova et.al. 2606.28164 null
2026-06-26 Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Yanchen Yin et.al. 2606.28153 null
2026-06-26 From Tokens to States: LLMs as a Special Case of World Models and the Continuous Path Beyond Paul Dubois et.al. 2606.28127 null
2026-06-26 Mechanism-Driven Monitors for Preemptive Detection of LLM Training Instability Ruixuan Huang et.al. 2606.28116 null
2026-06-26 Scaling limit of the Random Language Model Eric De Giuli et.al. 2606.28105 null
2026-06-26 Phase structure of the Random Language Model Alessio Giorlandino et.al. 2606.28103 null
2026-06-26 Typing Behavior in Human-LLM Interaction: Keystroke Dynamics Reveal Cognitive Effort During Prompting Laura Schütz et.al. 2606.28090 null
2026-06-26 Single and Multi Truth Data Fusion using Large Language Models Hira Beril Kucuk et.al. 2606.28062 null
2026-06-26 ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents Shijing Hu et.al. 2606.28061 null
2026-06-26 ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures Haoran Xu et.al. 2606.28060 null
2026-06-26 MultiHashFormer: Hash-based Generative Language Models Huiyin Xue et.al. 2606.28057 null
2026-06-25 Autoregressive Boltzmann Generators Danyal Rehman et.al. 2606.27361 null
2026-06-25 When are likely answers right? On Sequence Probability and Correctness in LLMs Johannes Zenn et.al. 2606.27359 null
2026-06-25 Mapping Political-Elite Networks in Europe with a Multilingual Joint Entity-Relation Extraction Pipeline Kirill Solovev et.al. 2606.27347 null
2026-06-25 Language-Based Digital Twins for Elderly Cognitive Assistance Mohammad Mehdi Hosseini et.al. 2606.27334 null
2026-06-25 LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank Serhii Hamotskyi et.al. 2606.27316 null
2026-06-25 ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Xumin Yu et.al. 2606.27313 null
2026-06-25 When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models Josef Chen et.al. 2606.27288 null
2026-06-25 Prompt Injection in Automated Résumé Screening with Large Language Models: Single and Multi-Injection Settings Preet Baxi et.al. 2606.27287 null
2026-06-25 Resource-Aware Neuro-Symbolic Reasoning for Local Small Language Models Carlos Ramírez Ovalle et.al. 2606.27281 null
2026-06-25 How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation Maria Levchenko et.al. 2606.27275 null
2026-06-25 CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs Hashmat Shadab Malik et.al. 2606.27264 null
2026-06-25 RSPC: A Benchmark for Modeling Stress and Psychiatric Conditions in Digitally Mediated Relationships using Psychiatrist Annotations Parmitha Vangapandu et.al. 2606.27247 null
2026-06-25 LMs as Task-Specific Knowledge Bases: An Interpretability Analysis Amit Elhelo et.al. 2606.27237 null
2026-06-25 A hardware-safety-gated system for LLM-written native ARTIQ control code on a trapped-ion platform Duanyang Wang et.al. 2606.27231 null
2026-06-25 Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text Manjinder Singh et.al. 2606.27215 null
2026-06-25 Smaller Models, Unexpected Costs: Trade-offs in LLM Quantization for Automated Program Repair Fernando Vallecillos-Ruiz et.al. 2606.27205 null
2026-06-25 Graph Neural Networks Applications Across Domains: All Insights You Need Abderaouf Bahi et.al. 2606.27202 null
2026-06-25 HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models Jiajun Wu et.al. 2606.27187 null
2026-06-25 Automating Potential-based Reward Shaping with Vision Language Model Guidance Henrik Müller et.al. 2606.27180 null
2026-06-25 TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference Tinghao Wang et.al. 2606.27161 null
2026-06-24 Learning Action Priors for Cross-embodiment Robot Manipulation Dong Jing et.al. 2606.26095 null
2026-06-24 Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models Akshay Paruchuri et.al. 2606.26079 null
2026-06-24 Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment Aditya Singh et.al. 2606.26071 null
2026-06-24 Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining Juliana Li et.al. 2606.26050 null
2026-06-24 How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations Yuxing Cheng et.al. 2606.26041 null
2026-06-24 AI translation of literary texts is “fine”, but readers still prefer human translations Yves Ferstler et.al. 2606.26040 null
2026-06-24 Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning Poojitha Thota et.al. 2606.26036 null
2026-06-24 TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs Yu-Yang Chen et.al. 2606.26029 null
2026-06-24 Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It Yupu Hao et.al. 2606.26027 null
2026-06-24 In-Context World Modeling for Robotic Control Siyin Wang et.al. 2606.26025 null
2026-06-24 Privacy Vulnerabilities of Attention Layers in Tabular Foundation Models and Protection of High-Risk Queries Tânia Carvalho et.al. 2606.26021 null
2026-06-24 SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models Liang-Yuan Wu et.al. 2606.25990 null
2026-06-24 Weave of Formal Thought Alexandre Bouayad et.al. 2606.25987 null
2026-06-24 InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy Mingguang Chen et.al. 2606.25984 null
2026-06-24 Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study Giulian Biolo et.al. 2606.25973 null
2026-06-24 Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Alexander Hägele et.al. 2606.25971 null
2026-06-24 Mixture-of-Experts RL for Fault-Tolerant Legged Locomotion Giulio Turrisi et.al. 2606.25965 null
2026-06-24 Agentic System as Compressor: Quantifying System Intelligence in Bits Zihan Qin et.al. 2606.25960 null
2026-06-24 Explainable Control Framework (XCF) based on Fuzzy Model-Agnostic Explanation and LLM Agent-Supported Interface Faliang Yin et.al. 2606.25941 null
2026-06-24 Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts Juri Opitz et.al. 2606.25935 null
2026-06-23 BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases Qi Chen et.al. 2606.24883 null
2026-06-23 “Zooming In” on Agentic Web Browsers as Assistive Technologies: A Case Study with a Low-Vision Technology Expert Laura Colazzo et.al. 2606.24870 null
2026-06-23 OpenThoughts-Agent: Data Recipes for Agentic Models Negin Raoof et.al. 2606.24855 null
2026-06-23 IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Zixuan Li et.al. 2606.24849 null
2026-06-23 Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models Ahmad Pouramini et.al. 2606.24841 null
2026-06-23 Vision-Language Model Reasoning for Contextual Semantic Mapping in Intralogistics Marvin Rüdt et.al. 2606.24814 null
2026-06-23 Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution Zidu Liu et.al. 2606.24808 null
2026-06-23 EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence Linpeng Huang et.al. 2606.24797 null
2026-06-23 Grad Detect: Gradient-Based Hallucination Detection in LLMs Anand Kamat et.al. 2606.24790 null
2026-06-23 Are We Ready For An Agent-Native Memory System? Wei Zhou et.al. 2606.24775 null
2026-06-23 Posterior Refinement: Fast Language Generation via Any-Order Flow Maps Manan Agarwal et.al. 2606.24773 null
2026-06-23 UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving Xiaowei Gao et.al. 2606.24759 null
2026-06-23 Can Scale Save Us From Plasticity Loss in Large Language Models? J. Fernando Hernandez-Garcia et.al. 2606.24752 null
2026-06-23 Scaling Laws for Task-Specific LLM Distillation Lavinia Ghita et.al. 2606.24747 null
2026-06-23 Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement Wangyi Pu et.al. 2606.24745 null
2026-06-23 World Value Models for Robotic Manipulation Zhihao Wang et.al. 2606.24742 null
2026-06-23 BioMedVR: Confusion-Aware Mixture-of-Prompt Experts for Biomedical Visual Reprogramming Jiaxiang Liu et.al. 2606.24740 null
2026-06-23 A Grounded Evidence-Retrieval Benchmark and Hybrid RAG Framework for Silicon Pixel Detector R&D Tianqi Gao et.al. 2606.24725 null
2026-06-23 Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations Jonas Klotz et.al. 2606.24716 null
2026-06-23 Automated Summarization of Software Documents: An LLM-based Multi-Agent Approach Duc S. H. Nguyen et.al. 2606.24689 null
2026-06-22 Randomized YaRN Improves Length Generalization for Long-Context Reasoning Manas Mehta et.al. 2606.23687 null
2026-06-22 Semantic Browsing: Controllable Diversity for Image Generation Sara Dorfman et.al. 2606.23679 null
2026-06-22 AIR: Adaptive Interleaved Reasoning with Code in MLLMs Cong Han et.al. 2606.23678 null
2026-06-22 Open Problem: Is AdamW Effective Under Heavy-Tailed Noise? Dingzhi Yu et.al. 2606.23676 null
2026-06-22 Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles Prateek Agnihotri et.al. 2606.23672 null
2026-06-22 Can LLMs Reliably Self-Report Adversarial Prefills, and How? Quang Minh Nguyen et.al. 2606.23671 null
2026-06-22 Tapered Language Models Reza Bayat et.al. 2606.23670 null
2026-06-22 On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners David Mguni et.al. 2606.23668 null
2026-06-22 The Table Says Otherwise: Testing LLMs with Counterfactual Relational Data Xinzhi Wang et.al. 2606.23667 null
2026-06-22 Statistical Proof as a Window into Human-AI Collaboration: Practical Insights and a Community Agenda Xiaojing Sun et.al. 2606.23666 null
2026-06-22 Muown Implicitly Performs Angular Step-size Decay Florian Hübler et.al. 2606.23637 null
2026-06-22 AI Exposure Scores: what they measure, what they miss, and what comes next Campbell Lund et.al. 2606.23633 null
2026-06-22 Data Selection Through Iterative Self-Filtering for Vision-Language Settings Andrei Liviu Nicolicioiu et.al. 2606.23611 null
2026-06-22 Causal Discovery in the Era of Agents Yujia Zheng et.al. 2606.23608 null
2026-06-22 Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers Tianyi Li et.al. 2606.23607 null
2026-06-22 SPIRAL: Learning to Search and Aggregate Jubayer Ibn Hamid et.al. 2606.23595 null
2026-06-22 The Topology of Ill-Posed Questions: Persistent Homology for Detection and Steering in LLMs Guangyu Jiang et.al. 2606.23590 null
2026-06-22 Evaluation Awareness Is Not One Capability: Evidence from Open Language Models Nilesh Nayan et.al. 2606.23583 null
2026-06-22 INCARBench: A Benchmark for Scientific Configuration in VASP INCAR by Large Language Models Bin Shao et.al. 2606.23571 null
2026-06-22 SVD-Surgeon: Optimal Singular-Value Surgery for Large Language Model Compression Mahmoud Safari et.al. 2606.23568 null
2026-06-21 PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement Weiwei Ye et.al. 2606.22610 null
2026-06-21 Sub-Billion, Super-Frontier: Small Language Models Rival Zero-Shot Frontier LLMs on General and Literary Relation Extraction Despina Christou et.al. 2606.22606 null
2026-06-21 MapReason-OSM: Can Vision-Language Models Make Graph-Verifiable Mobility Decisions from Street Maps ? Srinivas Venkatanarayanan et.al. 2606.22597 null
2026-06-21 Context-Aware Distillation and Ablation for Text2DSL Alexander V. Kozachok et.al. 2606.22578 null
2026-06-21 What are Key Factors for Updates in RL for LLM Reasoning? Peidong Wang et.al. 2606.22570 null
2026-06-21 Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do Zhuoran Jin et.al. 2606.22565 null
2026-06-21 Training-Free Semantic Correction for Autoregressive Visual Models Junhao Chen et.al. 2606.22550 null
2026-06-21 ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Weiwei Chen et.al. 2606.22541 null
2026-06-21 NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models Yingjie Zhang et.al. 2606.22537 null
2026-06-21 Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding Yuanhao Ding et.al. 2606.22511 null
2026-06-21 Benchmarking Vision-Language Models for Microscopic Plant Image Understanding Tianqi Wei et.al. 2606.22497 null
2026-06-21 Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing Aygün Varol et.al. 2606.22496 null
2026-06-21 An LLM-Orchestrated Agent for Directional-Coupler Design with Self-Consistent Eigenmode and FDTD Validation Saumya Biswas et.al. 2606.22493 null
2026-06-21 SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments Yundaichuan Zhan et.al. 2606.22488 null
2026-06-21 VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflows Teodoro Baldazzi et.al. 2606.22485 null
2026-06-21 ROMEVA: Geometry-Preserving Vocabulary Expansion for Roman Urdu Language Models Mahnoor Khan et.al. 2606.22478 null
2026-06-21 CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming Ruixun Liu et.al. 2606.22476 null
2026-06-21 All Green, Still Broken: Real-Flow Verification Lessons from an LLM-Integrated, Multi-Market Web Application Muhammad Bilal et.al. 2606.22475 null
2026-06-21 Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form Generation Areeba Hassan et.al. 2606.22474 null
2026-06-21 Interleaved Speech Language Models Latently Work In Text Talia Sternberg et.al. 2606.22473 null
2026-06-18 TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living Arkaprava Sinha et.al. 2606.20561 null
2026-06-18 From Efficiency to Leakage – Privacy Backdoor in Federated Language Model Fine-Tuning Shanghao Shi et.al. 2606.20553 null
2026-06-18 Toward Calibrated Mixture-of-Experts Under Distribution Shift Gina Wong et.al. 2606.20544 null
2026-06-18 SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation Shilong Xiang et.al. 2606.20543 null
2026-06-18 Multi-Task Bayesian In-Context Learning Qingyang Zhu et.al. 2606.20538 null
2026-06-18 StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs Shaghayegh Kolli et.al. 2606.20527 null
2026-06-18 HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining Juncheng Ma et.al. 2606.20521 null
2026-06-18 Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages Maria Ivanova et.al. 2606.20517 null
2026-06-18 What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations? Sihui Dai et.al. 2606.20508 null
2026-06-18 Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems Zewen Liu et.al. 2606.20493 null
2026-06-18 Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users Haw-Shiuan Chang et.al. 2606.20482 null
2026-06-18 Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology Yusuf Salcan et.al. 2606.20477 null
2026-06-18 Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems Reza Soosahabi et.al. 2606.20470 null
2026-06-18 Multi-View Decompilation for LLM-Based Malware Classification Bercan Turkmen et.al. 2606.20436 null
2026-06-18 Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation Karn Tiwari et.al. 2606.20419 null
2026-06-18 LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems Hanwool Lee et.al. 2606.20408 null
2026-06-18 PowerAgentBench-Dyn: A Benchmark for Agentic AI in Power System Dynamic Studies Qian Zhang et.al. 2606.20401 null
2026-06-18 Agentic AutoResearch forSpace Autonomy: An Auditable, LLM-Driven Research Agent for Aerospace Control Problems Amit Jain et.al. 2606.20394 null
2026-06-18 AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning Zepeng Li et.al. 2606.20373 null
2026-06-18 SoftSkill: Behavioral Compression for Contextual Adaptation Xijia Tao et.al. 2606.20333 null
2026-06-17 Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning Jisoo Kim et.al. 2606.19340 null
2026-06-17 Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Shengyuan Ding et.al. 2606.19338 null
2026-06-17 Learning User Simulators with Turing Rewards Yingshan Susan Wang et.al. 2606.19336 null
2026-06-17 Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation Siyi Gu et.al. 2606.19327 null
2026-06-17 Explaining Attention with Program Synthesis Amiri Hayes et.al. 2606.19317 null
2026-06-17 Diffusion-Proof: Recipe for Formal Theorem Proving Beyond Auto-Regressive Generation Ruida Wang et.al. 2606.19315 null
2026-06-17 Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Play Leyang Shen et.al. 2606.19308 null
2026-06-17 A Unified Framework for Efficient Remote Sensing Visual Question Answering: Adapting Dual, Hybrid, and Encoder-Decoder Architectures Timothy Agboada et.al. 2606.19277 null
2026-06-17 Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA Ikram Belmadani et.al. 2606.19266 null
2026-06-17 Structured Inference with Large Language Gibbs Sanghyeok Choi et.al. 2606.19264 null
2026-06-17 A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2 Yijin Wang et.al. 2606.19259 null
2026-06-17 DreamReasoner-8B: Block-Size Curriculum Learning for Diffusion Reasoning Models Zirui Wu et.al. 2606.19257 null
2026-06-17 X+Slides: Benchmarking Audience-Conditioned Slide Generation Haodong Chen et.al. 2606.19256 null
2026-06-17 OneCanvas: 3D Scene Understanding via Panoramic Reprojection Bartłomiej Baranowski et.al. 2606.19253 null
2026-06-17 CodeSentinel: A Three-Layer Defense Against Indirect Prompt Injection in Code Contexts Po-Han Cheng et.al. 2606.19235 null
2026-06-17 Language Models as Interfaces, Not Oracles: A Hybrid LLM-ML System for Pediatric Appendicitis Soheyl Bateni et.al. 2606.19183 null
2026-06-17 User as Engram: Internalizing Per-User Memory as Local Parametric Edits Bojie Li et.al. 2606.19172 null
2026-06-17 Dango: A Strictly L1-Only Large Language Model for Studying Second Language Acquisition Shiho Matta et.al. 2606.19170 null
2026-06-17 Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection Jinhan Li et.al. 2606.19168 null
2026-06-17 Teaching Software Engineering with LLM and MCP Integration: From Classroom to Industry Practice Kehui Chen et.al. 2606.19167 null
2026-06-16 Variable-Width Transformers Zhaofeng Wu et.al. 2606.18246 null
2026-06-16 EventDrive: Event Cameras for Vision-Language Driving Intelligence Dongyue Lu et.al. 2606.18242 null
2026-06-16 Darshana Graph: A Parallel Commentary Corpus for Comparative Indian Philosophy, with Stylometric and Exploratory Graph Analyses Joy Bose et.al. 2606.18222 null
2026-06-16 Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Byung-Kwan Lee et.al. 2606.18216 null
2026-06-16 Learning from the Self-future: On-policy Self-distillation for dLLMs Yifu Luo et.al. 2606.18195 null
2026-06-16 A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models Nicola Franco et.al. 2606.18193 null
2026-06-16 The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data Nick Bettencourt et.al. 2606.18192 null
2026-06-16 Multi-Source Cybersecurity Logs: An ATT&CK-Labeled Dataset and SLM Evaluation Abir Ashab Niloy et.al. 2606.18190 null
2026-06-16 IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction Henry Bodwell et.al. 2606.18181 null
2026-06-16 Evaluating Open-Source LLMs for Multi-Label ATT&CK Technique Classification on CTI Reports Ahmed Ryan et.al. 2606.18166 null
2026-06-16 Learning Dynamics of Chain-of-Thought State Tracking in a Solvable Transformer Model Niklas Forner et.al. 2606.18164 null
2026-06-16 The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act Michèle Finck et.al. 2606.18158 null
2026-06-16 Learning Cardiac Electrophysiology Digital Twins Through Agentic Discovery of Hybrid Structure Ziqi Zhou et.al. 2606.18154 null
2026-06-16 WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning Yuwei Zhang et.al. 2606.18147 null
2026-06-16 Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning Alexander Polok et.al. 2606.18134 null
2026-06-16 Unintended Effects of Geographic Conditioning in Large Language Models Naz Col et.al. 2606.18124 null
2026-06-16 Predicting Immune Biomarkers with MultiModal Mixture-of-Expert Pathology Foundation Models Empowers Precision Oncology Tianyu Liu et.al. 2606.18123 null
2026-06-16 Structural Role Injection in Handlebars-Templated LLM Prompts: Triple-Brace Interpolation, Delimiter Family, and the Limits of HTML Auto-Escaping Mohammadreza Rashidi et.al. 2606.18120 null
2026-06-16 Querying an astronomical database using large language models: the ALeRCE text-to-SQL system P. A. Estevez et.al. 2606.18108 null
2026-06-16 OmniPlan: An Adaptive Framework for Timely and Near-Optimal Network Planning Optimization Longlong Zhu et.al. 2606.18105 null
2026-06-15 The Value Axis: Language Models Encode Whether They’re on the Right Track Nick Jiang et.al. 2606.17056 null
2026-06-15 Context-Aware RL for Agentic and Multimodal LLMs Peiyang Xu et.al. 2606.17053 null
2026-06-15 FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models Jiaju Han et.al. 2606.17020 null
2026-06-15 When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning Nathan Gavenski et.al. 2606.16995 null
2026-06-15 DreamX-World 1.0: A General-Purpose Interactive World Model DreamX Team et.al. 2606.16993 null
2026-06-15 Consensus-based Agentic Large Language Model Framework for Harmonized Tariff Schedule Code Classification Truong Thanh Hung Nguyen et.al. 2606.16987 null
2026-06-15 Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data Kareem Amin et.al. 2606.16952 null
2026-06-15 Scalable Circuit Learning for Interpreting Large Language Models Naiyu Yin et.al. 2606.16939 null
2026-06-15 Exploring Extrinsic and Intrinsic Properties for Effective Reasoning with Code Interpreter Patomporn Payoungkhamdee et.al. 2606.16934 null
2026-06-15 IMPACTeen: Intentions, Manipulation, Persuasion, Annotations, and Consequences in Teen Communication Dataset Aleksander Szczęsny et.al. 2606.16910 null
2026-06-15 LESS Is More: Mutual-Stability Sampling for Diffusion Language Models Amr Mohamed et.al. 2606.16908 null
2026-06-15 Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences Mingyang Li et.al. 2606.16905 null
2026-06-15 Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models Dongbin Na et.al. 2606.16902 null
2026-06-15 Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization Kaiyue Wen et.al. 2606.16899 null
2026-06-15 Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization Dongbin Na et.al. 2606.16898 null
2026-06-15 Contrastive-Difference CKA Reveals Concept-Specific Structural Alignment Across Language Model Architectures Xueping Gao et.al. 2606.16897 null
2026-06-15 Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering Sanjay Basu et.al. 2606.16890 null
2026-06-15 Neuro-Symbolic Software Verification: Hyper-charging Local Language Models with Symbolic Reasoning at Scale Muhammad A. A. Pirzada et.al. 2606.16886 null
2026-06-15 Revisiting the Systematicity in Negation in the Era of In-Context Learning Hitomi Yanaka et.al. 2606.16867 null
2026-06-15 Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens Yizhen Yao et.al. 2606.16847 null
2026-06-12 Gaze Heads: How VLMs Look at What They Describe Rohit Gandikota et.al. 2606.14703 null
2026-06-12 OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains Xinyue Cai et.al. 2606.14702 null
2026-06-12 RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space Xichen Pan et.al. 2606.14700 null
2026-06-12 Instruct-Particulate: Scaling Feed-Forward 3D Object Articulation with Kinematic Control Ruining Li et.al. 2606.14699 null
2026-06-12 ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning Sicheng Yang et.al. 2606.14697 null
2026-06-12 Persona-Pruner: Sculpting Lightweight Models for Role-Playing Jinsu Kim et.al. 2606.14695 null
2026-06-12 CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment Jiayue Cao et.al. 2606.14691 null
2026-06-12 Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows Shikun Liu et.al. 2606.14672 null
2026-06-12 The Self-Aware Body: A User-Centered Framework for Designing Therapeutic Sonic Interactions Prithvi Ravi Kantan et.al. 2606.14664 null
2026-06-12 Abstracting Cross-Domain Action Sequences into Interpretable Workflows Gaurav Verma et.al. 2606.14654 null
2026-06-12 From Self-Supervised Speech Models to Mixture-of-Experts for Robust Anti-Spoofing Hugo Daumain et.al. 2606.14639 null
2026-06-12 When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks Jianzhe Lin et.al. 2606.14629 null
2026-06-12 Neither Parallel Nor Sequential: How DiffusionGemma Actually Commits Tokens Ali Asaria et.al. 2606.14620 null
2026-06-12 Safe Reinforcement Learning of Autonomous Highway Driving: A Unified Framework for Safety and Efficiency Chufei Yan et.al. 2606.14609 null
2026-06-12 Expert-Driven Survival Machines: Improving Stratification and Interpretability in Multiple Clinical Cohorts Farica Zhuang et.al. 2606.14608 null
2026-06-12 Empowering Student Debugging in Parallel Programming with Execution Traces and Large Language Models God’salvation F. Oguibe et.al. 2606.14607 null
2026-06-12 What Robots Do Matters More Than What They Look Like: Task Context Shapes Trust in Educational HRI Anna-Maria Velentza et.al. 2606.14602 null
2026-06-12 AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models Hui Geng et.al. 2606.14591 null
2026-06-12 S $^2$ COPE: Self-Supervised Concept Discovery via Preference Learning Shilong Xiang et.al. 2606.14586 null
2026-06-12 SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model Xiaoxin Lu et.al. 2606.14574 null
2026-06-11 EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments Jundong Xu et.al. 2606.13681 null
2026-06-11 Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning Zilin Xiao et.al. 2606.13680 null
2026-06-11 Improving Robotic Generalist Policies via Flow Reversal Steering Andy Tang et.al. 2606.13675 null
2026-06-11 SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Seokju Cho et.al. 2606.13673 null
2026-06-11 Automated reproducibility assessments in the social and behavioral sciences using large language models Tobias Holtdirk et.al. 2606.13670 null
2026-06-11 Influcoder: Distilling Decoders’ Gradient Influence Rankings into an Encoder for Data Attribution Dimitri Kachler et.al. 2606.13668 null
2026-06-11 Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation Guo Yu et.al. 2606.13657 null
2026-06-11 Operadic consistency: a label-free signal for compositional reasoning failures in LLMs Nathaniel Bottman et.al. 2606.13649 null
2026-06-11 SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation Marek Šuppa et.al. 2606.13647 null
2026-06-11 Recursive Agent Harnesses Elias Lumer et.al. 2606.13643 null
2026-06-11 Beyond Uniform Tokens: Adaptive Compression for Time Series Language Models Jialin Gan et.al. 2606.13624 null
2026-06-11 Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning Zach Studdiford et.al. 2606.13607 null
2026-06-11 Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Daniel Scalena et.al. 2606.13603 null
2026-06-11 Reward Modeling for Multi-Agent Orchestration King Yeung Tsang et.al. 2606.13598 null
2026-06-11 ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages Tanmoy Kanti Halder et.al. 2606.13572 null
2026-06-11 A Three-Layer Framework for AI in Scientific Discovery Guojun Liao et.al. 2606.13566 null
2026-06-11 Adaptive Turn-Taking for Real-time Multi-Party Voice Agents Soumyajit Mitra et.al. 2606.13544 null
2026-06-11 AgentRivet: an automated system for producing Rivet routines from journal publications Antonio J. Costa et.al. 2606.13535 null
2026-06-11 Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data Qixu Chen et.al. 2606.13507 null
2026-06-11 From Traditional Automation to Embodied Wireless Intelligence: Vision-Language-Action Empowered Physics-Aware Communication Networks Genze Jiang et.al. 2606.13458 null
2026-06-10 Reroute, Don’t Remove: Recoverable Visual Token Routing for Vision-Language Models Cheng-Yu Yang et.al. 2606.12412 null
2026-06-10 How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology Kian R. Weihrauch et.al. 2606.12407 null
2026-06-10 DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners? Jadelynn Dao et.al. 2606.12402 null
2026-06-10 Doc-to-Atom: Learning to Compile and Compose Memory Atoms Xingjian Diao et.al. 2606.12400 null
2026-06-10 Redesign Mixture-of-Experts Routers with Manifold Power Iteration Songhao Wu et.al. 2606.12397 null
2026-06-10 System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5 Haotao Xie et.al. 2606.12392 null
2026-06-10 TAHOE: Text-to-SQL with Automated Hint Optimization from Experience Zhiyi Chen et.al. 2606.12387 null
2026-06-10 APPO: Agentic Procedural Policy Optimization Xucong Wang et.al. 2606.12384 null
2026-06-10 Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization Hao Xiang et.al. 2606.12373 null
2026-06-10 Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Yucheng Li et.al. 2606.12370 null
2026-06-10 Should LLM Agents Decide in Social Simulations? Comparing Finite-State and LLM-Based Decision Policies Alejandro Buitrago López et.al. 2606.12369 null
2026-06-10 APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies Kechun Xu et.al. 2606.12366 null
2026-06-10 On Subquadratic Architectures: From Applications to Principles Anamaria-Roberta Hartl et.al. 2606.12364 null
2026-06-10 Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Leon Bergen et.al. 2606.12360 null
2026-06-10 Nonslop: A Gamified Experiment in Human-AI Collaborative Writing Maria Edwards et.al. 2606.12350 null
2026-06-10 ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing Chirag Chawla et.al. 2606.12342 null
2026-06-10 OCELOT: Inference-Leakage Budgets for Privacy-Preserving LLM Agents Jin Xie et.al. 2606.12341 null
2026-06-10 Harness In-Context Operator Learning with Chain of Operators Minghui Yang et.al. 2606.12318 null
2026-06-10 Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering Hyun Joe Jeong et.al. 2606.12299 null
2026-06-10 Bridging the Modality Gap in Forensic Image Retrieval Ricardo González-Gazapo et.al. 2606.12294 null
2026-06-09 Next Forcing: Causal World Modeling with Multi-Chunk Prediction Gangwei Xu et.al. 2606.11187 null
2026-06-09 The Role of Feedback Alignment in Self-Distillation Semih Kara et.al. 2606.11173 null
2026-06-09 Flaws in the LLM Automation Narrative George Perrett et.al. 2606.11166 null
2026-06-09 ReasonAlloc: Hierarchical Decoding-Time KV Cache Budget Allocation for Reasoning Models Wenhao Liu et.al. 2606.11164 null
2026-06-09 DarkAgents Michele Lucente et.al. 2606.11157 null
2026-06-09 P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning Yikang Yang et.al. 2606.11152 null
2026-06-09 JOIN: Anchor-Grasp-Conditioned Joining via Opposition, Inference, and Navigation for Bimanual Assistive Manipulation Drake Moore et.al. 2606.11151 null
2026-06-09 ABC-Bench: An Agentic Bio-Capabilities Benchmark for Biosecurity Andrew Bo Liu et.al. 2606.11150 null
2026-06-09 OpenPCC: Open and Confidential LLM Serving on Commodity TEEs Haoling Zhou et.al. 2606.11145 null
2026-06-09 TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Heming Zou et.al. 2606.11119 null
2026-06-09 Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA Vinamra Sharma et.al. 2606.11117 null
2026-06-09 A Neurosymbolic Prolog Skill for LLM-Driven Service Placement Jacopo Massa et.al. 2606.11113 null
2026-06-09 FADA: Accessible fetal ultrasound interpretation and annotation with a selectively distilled unified vision-language model Mahmood Alzubaidi et.al. 2606.11106 null
2026-06-09 PhantomBench: Benchmarking the Non-existential Threat of Language Models Haeji Jung et.al. 2606.11105 null
2026-06-09 The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models Hakan Mehmetcik et.al. 2606.11082 null
2026-06-09 Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models Peiqi Jia et.al. 2606.11074 null
2026-06-09 T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains Genta Indra Winata et.al. 2606.11070 null
2026-06-09 The social consequences of AI delegation Henrique Ferraz de Arruda et.al. 2606.11058 null
2026-06-09 Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Xinyu Zhou et.al. 2606.11052 null
2026-06-09 LLM-Mediated Demand Response Coordination in Smart Microgrids J. de Curtò et.al. 2606.11050 null
2026-06-08 OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics Mingxian Lin et.al. 2606.09826 null
2026-06-08 Causally Evaluating the Learnability of Formal Language Tasks Vésteinn Snæbjarnarson et.al. 2606.09822 null
2026-06-08 Rethinking the Divergence Regularization in LLM RL Jiarui Yao et.al. 2606.09821 null
2026-06-08 Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models Seongbin Park et.al. 2606.09749 null
2026-06-08 HDSL: A Hierarchical Domain-Specific Language for Structured 3D Indoor Scene Generation and Localized Editing with LLM Agents Letian Li et.al. 2606.09738 null
2026-06-08 The Neutral Mask: How RLHF Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model Wendy K. Tam et.al. 2606.09735 null
2026-06-08 Tight Sample Complexity of Transformers Chenxiao Yang et.al. 2606.09731 null
2026-06-08 SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research Pu Ning et.al. 2606.09730 null
2026-06-08 Beyond Probabilistic Similarity: Structural, Temporal, and Causal Limitations of Retrieval-Augmented Generation in the Legal Domain Hudson de Martim et.al. 2606.09724 null
2026-06-08 Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Mohammad Beigi et.al. 2606.09711 null
2026-06-08 IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking Zechen Sun et.al. 2606.09709 null
2026-06-08 Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO Blake Bullwinkel et.al. 2606.09701 null
2026-06-08 What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks Qin Yang et.al. 2606.09700 null
2026-06-08 PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models Gianluca Barmina et.al. 2606.09697 null
2026-06-08 Correlation Is Not Enough: Embedding Human Metadata for Individual Causal Discovery Suraj Biswas et.al. 2606.09672 null
2026-06-08 SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Hongcheng Gao et.al. 2606.09669 null
2026-06-08 In-Context Learning for Latent Space Bayesian Optimization Tuan A. Vu et.al. 2606.09664 null
2026-06-08 End-to-End Context Compression at Scale Ang Li et.al. 2606.09659 null
2026-06-08 Muon Learns More Robust and Transferable Features than Adam Tianyu Ruan et.al. 2606.09658 null
2026-06-08 Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving Yimu Wang et.al. 2606.09644 null
2026-06-08 Gradient-Guided Reward Optimization for Inference-time Alignment Hankun Lin et.al. 2606.09635 null
2026-06-08 Civil Court Simulation with Large Language Models Yifan Chen et.al. 2606.09632 null
2026-06-08 ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies Haodi Hu et.al. 2606.09630 null
2026-06-08 Closure-Validated Circuit Discovery in Attention Heads: Co-activation Proposes, Ablation Disposes Yongzhong Xu et.al. 2606.09607 null
2026-06-08 Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation Ali Tourani et.al. 2606.09595 null
2026-06-08 Clinically Grounded Privacy Evaluation of Medical LMs Sasha Ronaghi et.al. 2606.09590 null
2026-06-08 Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text Yutong Bian et.al. 2606.09585 null
2026-06-08 TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs Momina Ahsan et.al. 2606.09578 null
2026-06-08 Code Is More Than Text: Uncertainty Estimation for Code Generation Yuling Shi et.al. 2606.09577 null
2026-06-08 UXBench: Benchmarking User Experience in AI Assistants Mengze Hong et.al. 2606.09570 null
2026-06-08 PRISM: Recovering Instruction Sets from Language Model Activations Gilad Gressel et.al. 2606.09563 null
2026-06-08 FuseFSS: Efficient Secure LLM Inference with Function Secret Sharing Yuhan Ma et.al. 2606.09551 null
2026-06-08 SecureClaw: Clawing Back Control of LLM Agents Yuhan Ma et.al. 2606.09549 null
2026-06-08 Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur? Apratim Bhattacharyya et.al. 2606.09547 null
2026-06-05 Counterintuitive problems in discrete probability Luca Avena et.al. 2606.07516 null
2026-06-05 How reliable are LLMs when it comes to playing dice? Luca Avena et.al. 2606.07515 null
2026-06-05 MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism Cong Chen et.al. 2606.07512 null
2026-06-05 Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings Songhao Wu et.al. 2606.07502 null
2026-06-05 Sparse Subspace-to-Expert Sharing for Task-Agnostic Continual Learning Fatema Siddika et.al. 2606.07500 null
2026-06-05 Supervision versus Demonstration-Based In-Context Learning for Multiword Expression Classification Sercan Karakaş et.al. 2606.07479 null
2026-06-05 TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment Sweta Mahajan et.al. 2606.07451 null
2026-06-05 Sycophantic Praise: Evaluating Excessive Praise in Language Models Daniel Vennemeyer et.al. 2606.07441 null
2026-06-05 Watch, Remember, Reason: Human-View Video Understanding with MLLMs Jiahao Meng et.al. 2606.07433 null
2026-06-05 Rapid co-design of Buoyancy-assisted robots for Challenging Locomotion using Gaussian Evolutionary Specialists Ankit Sinha et.al. 2606.07424 null
2026-06-05 The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs Yang Zhang et.al. 2606.07422 null
2026-06-05 Lost in Migration: Exposing Android Framework Vulnerabilities in Parallel Java-Kotlin Implementations Rui Li et.al. 2606.07420 null
2026-06-05 Sparsely gated tiny linear experts Simon Schug et.al. 2606.07414 null
2026-06-05 Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills Chuan Xiao et.al. 2606.07412 null
2026-06-05 A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning Yuxiang Chen et.al. 2606.07410 null
2026-06-05 Reversible Foundations: Training a 120B Sparse MoE through State-Preserving Scaling Rohan Shravan et.al. 2606.07404 null
2026-06-05 Online Pandora’s Box for Contextual LLM Cascading Alexandre Belloni et.al. 2606.07392 null
2026-06-05 Self-evolving LLM agents with in-distribution Optimization Yudi Zhang et.al. 2606.07367 null
2026-06-05 On the Shoulders of Giants: Empowering Automated Smart Contract Auditing via the GiAnt Corpus Xiaoting Zhang et.al. 2606.07363 null
2026-06-05 TabSwift: An Efficient Tabular Foundation Model with Row-Wise Attention Si-Yang Liu et.al. 2606.07345 null
2026-06-04 HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers Lizhi Yang et.al. 2606.06493 null
2026-06-04 Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution Liliana Hotsko et.al. 2606.06492 null
2026-06-04 TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies Dong Jing et.al. 2606.06491 null
2026-06-04 PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding Shaohui Dai et.al. 2606.06485 null
2026-06-04 Pretraining Recurrent Networks without Recurrence Akarsh Kumar et.al. 2606.06479 null
2026-06-04 Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators Chenming Zhu et.al. 2606.06476 null
2026-06-04 RREDCoT: Segment-Level Reward Redistribution for Reasoning Models Mykyta Ielanskyi et.al. 2606.06475 null
2026-06-04 Self-Augmenting Retrieval for Diffusion Language Models Paul Jünger et.al. 2606.06474 null
2026-06-04 MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery Shangheng Du et.al. 2606.06473 null
2026-06-04 You Only Index Once: Cross-Layer Sparse Attention with Shared Routing Yutao Sun et.al. 2606.06467 null
2026-06-04 Human Adults and LLMs as Scientists: Who Benefits from Active Exploration? Mandana Samiei et.al. 2606.06464 null
2026-06-04 Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill Mehmet Iscan et.al. 2606.06454 null
2026-06-04 Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Zhuoming Chen et.al. 2606.06453 null
2026-06-04 Latent Reasoning with Normalizing Flows Guancheng Tu et.al. 2606.06447 null
2026-06-04 USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding Heng-Jui Chang et.al. 2606.06444 null
2026-06-04 Revising Context, Shifting Simulated Stance: Auditing LLM-Based Stance Simulation in Online Discussions Xinnong Zhang et.al. 2606.06443 null
2026-06-04 Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation Hanxu Hu et.al. 2606.06428 null
2026-06-04 Annotation of Positive vs Negative User Interactions for Social Sign Prediction Biancamaria Bombino et.al. 2606.06425 null
2026-06-04 A Komi-Yazva–Russian Parallel Corpus and Evaluation Protocol for Zero- and Few-Shot LLM Translation Petr Parshakov et.al. 2606.06420 null
2026-06-04 Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Thomas T. Zhang et.al. 2606.06418 null
2026-05-29 Linear Scaling Video VLMs for Long Video Understanding Cristobal Eyzaguirre et.al. 2605.31598 null
2026-05-29 SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models Olaf Dünkel et.al. 2605.31597 null
2026-05-29 Stateful Online Monitoring Catches Distributed Agent Attacks Davis Brown et.al. 2605.31593 null
2026-05-29 Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions Wesley Scivetti et.al. 2605.31586 null
2026-05-29 LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards Nianyi Lin et.al. 2605.31584 null
2026-05-29 Can Generative AI help people navigate Radical Moral Disagreements? The CONSIDER prototype William Hohnen-Ford et.al. 2605.31574 null
2026-05-29 What Gets Unmasked First? Trajectory Analysis of Diffusion Models for Graph-to-Text Generation Qing Wang et.al. 2605.31564 null
2026-05-29 What Am I Missing? Question-Answering as Hidden State Probing Chu Fei Luo et.al. 2605.31561 null
2026-05-29 Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization Felipe Urrutia et.al. 2605.31558 null
2026-05-29 Vision-Language Models Suppress Female Representations Under Ambiguous Input Arnau Marin-Llobet et.al. 2605.31556 null
2026-05-29 Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models Yibin Zhao et.al. 2605.31550 null
2026-05-29 Preference-Aware Rubric Learning for Personalized Evaluation Yilun Qiu et.al. 2605.31545 null
2026-05-29 If LLMs Have Human-Like Attributes, Then So Does Age of Empires II Adrian de Wynter et.al. 2605.31514 null
2026-05-29 Personalize Your Large Vision-language Models With In-context Prompt Tuning Yanshu Li et.al. 2605.31513 null
2026-05-29 Reliable Multilingual Orthopedic Decision Support from Clinical Narratives: Language-Aware Adaptation and Verification-Guided Deferral Danish Ali et.al. 2605.31512 null
2026-05-29 Skill Reuse as Compression in Agentic RL Zhikun Xu et.al. 2605.31509 null
2026-05-29 Assign and Add: A Mechanistic Study of Compositional Arithmetic Brady Exoo et.al. 2605.31497 null
2026-05-29 General-purpose LLMs as Constrained Crystal Composition Generators Hedda Oschinski et.al. 2605.31495 null
2026-05-29 Consolidating Rewarded Perturbations for LLM Post-Training Zheyu Zhang et.al. 2605.31494 null
2026-05-29 LinTree: Improving LLM Reasoning with Explicitly Structured Search Histories Liwei Kang et.al. 2605.31492 null
2026-05-28 VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion Hidir Yesiltepe et.al. 2605.30351 null
2026-05-28 LLMSurgeon: Diagnosing Data Mixture of Large Language Models Yaxin Luo et.al. 2605.30348 null
2026-05-28 SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations Qinpei Luo et.al. 2605.30345 null
2026-05-28 Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection Xiaona Zhou et.al. 2605.30344 null
2026-05-28 Unlocking the Working Memory of Large Language Models for Latent Reasoning Lukas Aichberger et.al. 2605.30343 null
2026-05-28 GPIC: A Giant Permissive Image Corpus for Visual Generation Keshigeyan Chandrasegaran et.al. 2605.30341 null
2026-05-28 Efficient Test-Time Finetuning of LLMs via Convex Reconstruction and Gradient Caching Alaa Khamis et.al. 2605.30337 null
2026-05-28 Demystifying Data Organization for Enhanced LLM Training Yalun Dai et.al. 2605.30334 null
2026-05-28 COMPOSE: Composing Future Theorems from Citations and Formal Structure David Busbib et.al. 2605.30333 null
2026-05-28 SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? Sy-Tuyen Ho et.al. 2605.30329 null
2026-05-28 Reasoning with Sampling: Cutting at Decision Points Felix Zhou et.al. 2605.30327 null
2026-05-28 In-Context Reward Adaptation for Robust Preference Modeling Zhenyu Sun et.al. 2605.30323 null
2026-05-28 Grounded 3D-Aware Spatial Vision-Language Modeling An-Chieh Cheng et.al. 2605.30307 null
2026-05-28 MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings Valentina Bui Muti et.al. 2605.30295 null
2026-05-28 Statistical Embeddings for Similarity, Retrieval, and Interpretable Alignment of Numeric Tabular Datasets M. Ross Kunz et.al. 2605.30289 null
2026-05-28 ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure A. J. Lew et.al. 2605.30284 null
2026-05-28 Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Qiuyue Wang et.al. 2605.30280 null
2026-05-28 Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection Yutong Wang et.al. 2605.30274 null
2026-05-28 LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback Jiwon Kim et.al. 2605.30273 null
2026-05-28 LoMo: Local Modality Substitution for Deeper Vision-Language Fusion Feng Han et.al. 2605.30265 null
2026-05-25 From Model Scaling to System Scaling: Scaling the Harness in Agentic AI Shangding Gu et.al. 2605.26112 null
2026-05-25 Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation Shuhong Zheng et.al. 2605.26111 null
2026-05-25 Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning Jun-Tao Tang et.al. 2605.26110 null
2026-05-25 Looped Diffusion Language Models Sanghyun Lee et.al. 2605.26106 null
2026-05-25 InstructSAM: Segment Any Instance with Any Instructions Yuqian Yuan et.al. 2605.26102 null
2026-05-25 Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models Bar Weiss et.al. 2605.26100 null
2026-05-25 Language Models Need Sleep Sangyun Lee et.al. 2605.26099 null
2026-05-25 Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay Martin Marek et.al. 2605.26097 null
2026-05-25 OrpQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization Maoyang Xiang et.al. 2605.26092 null
2026-05-25 Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User’s Digital World Yusong Lin et.al. 2605.26086 null
2026-05-25 Automated Benchmark Auditing for AI Agents and Large Language Models Junlin Wang et.al. 2605.26079 null
2026-05-25 WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification Lingyu Gao et.al. 2605.26070 null
2026-05-25 Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals Federico Torrielli et.al. 2605.26045 null
2026-05-25 L2IR: Revealing Latent Intent in Graph Fraud Detection Jinsheng Guo et.al. 2605.26040 null
2026-05-25 DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models Xinrui Shi et.al. 2605.26038 null
2026-05-25 Toward General Quantum Control with Physics-Informed Large Language Models Yusheng Zhao et.al. 2605.26021 null
2026-05-25 Retrieval-Augmented Detection of Potentially Abusive Clauses in Chilean Terms of Service Christoffer Loeffler et.al. 2605.26019 null
2026-05-25 STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models Yiming Liang et.al. 2605.26014 null
2026-05-25 MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models Shristi Das Biswas et.al. 2605.26004 null
2026-05-25 Causal methods for LLM development and evaluation Dennis Frauen et.al. 2605.25998 null
2026-05-22 LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws Xu Ouyang et.al. 2605.23901 null
2026-05-22 SPACENUM: Revisiting Spatial Numerical Understanding in VLMs Jianshu Zhang et.al. 2605.23898 null
2026-05-22 ETCHR: Editing To Clarify and Harness Reasoning Beichen Zhang et.al. 2605.23897 null
2026-05-22 Complete-muE: Optimal Hyperparameter Transfer and Scaling for MoE Models Hongwu Peng et.al. 2605.23893 null
2026-05-22 Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework Xiao Cao et.al. 2605.23891 null
2026-05-22 Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions Anastasiia Sedova et.al. 2605.23885 null
2026-05-22 PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs Rim Assouel et.al. 2605.23883 null
2026-05-22 Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer Aratrika Mustafi et.al. 2605.23871 null
2026-05-22 Human Decision-Making with Persuasive and Narrative LLM Explanations Laura R. Marusich et.al. 2605.23867 null
2026-05-22 Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection – Submission for WildSpoof 2026 TTS Track Renhe Sun et.al. 2605.23859 null
2026-05-22 Strong Teacher Not Needed? On Distillation in LLM Pretraining Taiming Lu et.al. 2605.23857 null
2026-05-22 Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion Remko Proesmans et.al. 2605.23847 null
2026-05-22 Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval Michal Shlapentokh-Rothman et.al. 2605.23826 null
2026-05-22 It’s the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt Stuart Bladon et.al. 2605.23825 null
2026-05-22 Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence Andres Nava et.al. 2605.23821 null
2026-05-22 Inferential Privacy Leakage in Anonymized Conversational AI Logs S M Mehedi Zaman et.al. 2605.23820 null
2026-05-22 Advanced AI Service Provisioning in O-RAN through LLM Engine Integration Seyed Bagher Hashemi Natanzi et.al. 2605.23809 null
2026-05-22 Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language Models Bo Peng et.al. 2605.23797 null
2026-05-22 Benchmarking LLMs for Community Governance Simulation with Life-history Narratives Xu Chen et.al. 2605.23783 null
2026-05-22 Beyond Binary Edits Robust Multimodal Knowledge Editing with Adversarial Subspace Alignment Haoyuan Wang et.al. 2605.23780 null
2026-05-21 Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs Jongseo Lee et.al. 2605.22823 null
2026-05-21 Tokenisation via Convex Relaxations Jan Tempus et.al. 2605.22821 null
2026-05-21 Vector Policy Optimization: Training for Diversity Improves Test-Time Search Ryan Bahlous-Boldi et.al. 2605.22817 null
2026-05-21 AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation Wenxuan Guo et.al. 2605.22816 null
2026-05-21 GS-QA: A Benchmark for Geospatial Question Answering Majid Saeedan et.al. 2605.22811 null
2026-05-21 Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention Ali Hatamizadeh et.al. 2605.22791 null
2026-05-21 LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems Sadia Asif et.al. 2605.22786 null
2026-05-21 FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection Huanchi Wang et.al. 2605.22779 null
2026-05-21 Reducing Political Manipulation with Consistency Training Long Phan et.al. 2605.22771 null
2026-05-21 Understanding Data Temporality Impact on Large Language Models Pre-training Pilchen Hippolyte et.al. 2605.22769 null
2026-05-21 Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation Samson Gourevitch et.al. 2605.22765 null
2026-05-21 Advancing Mathematics Research with AI-Driven Formal Proof Search George Tsoukalas et.al. 2605.22763 null
2026-05-21 Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models Juergen Dietrich et.al. 2605.22732 null
2026-05-21 Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation Dong Nie et.al. 2605.22731 null
2026-05-21 AMEL: Accumulated Message Effects on LLM Judgments Sid-ali Temkit et.al. 2605.22714 null
2026-05-21 Tokenization with Split Trees Craig W. Schmidt et.al. 2605.22705 null
2026-05-21 Machine Learning Interatomic Potentials: Advancing Open-Source Software for Efficient and Scalable Molecular Simulation Christoph Brunken et.al. 2605.22698 null
2026-05-21 Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models Piotr Kubaty et.al. 2605.22679 null
2026-05-21 Self-Policy Distillation via Capability-Selective Subspace Projection Guangya Hao et.al. 2605.22675 null
2026-05-21 Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Nick Merrill et.al. 2605.22672 null
2026-05-20 Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Dayal Singh Kalra et.al. 2605.21486 null
2026-05-20 EvoStruct: Bridging Evolutionary and Structural Priors for Antibody CDR Design via Protein Language Model Adaptation Mansoor Ahmed et.al. 2605.21485 null
2026-05-20 DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation Sixiong Xie et.al. 2605.21482 null
2026-05-20 WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata Basel Shbita et.al. 2605.21479 null
2026-05-20 You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories Zhepei Wei et.al. 2605.21468 null
2026-05-20 DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Kaiyi Zhang et.al. 2605.21467 null
2026-05-20 Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution Weixing Zhang et.al. 2605.21465 null
2026-05-20 Mem- $π$ : Adaptive Memory through Learning When and What to Generate Xiaoqiang Wang et.al. 2605.21463 null
2026-05-20 TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos Yakun Yu et.al. 2605.21443 null
2026-05-20 PALS: Power-Aware LLM Serving for Mixture-of-Experts Models Can Hankendi et.al. 2605.21427 null
2026-05-20 AIGaitor: Privacy-preserving and cloud-free motion analysis for everyone, using edge computing Lauhitya Reddy et.al. 2605.21421 null
2026-05-20 RoadTones: Tone Controllable Text Generation from Road Event Videos Chirag Parikh et.al. 2605.21411 null
2026-05-20 Quantifying the cross-linguistic effects of syncretism on agreement attraction Utku Turk et.al. 2605.21403 null
2026-05-20 Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment Roland Pihlakas et.al. 2605.21401 null
2026-05-20 Post-Hoc Understanding of Metaphor Processing in Decoder-Only Language Models via Conditional Scale Entropy Lawhori Chakrabarti et.al. 2605.21391 null
2026-05-20 Combating Harms of Generative AI in CS1 with Code Review Interviews and a Flipped Classroom Peter Fowles et.al. 2605.21374 null
2026-05-20 “I didn’t Make the Micro Decisions”: Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration Eunsu Kim et.al. 2605.21363 null
2026-05-20 LASH: Adaptive Semantic Hybridization for Black-Box Jailbreaking of Large Language Models Abdullah Al Nomaan Nafi et.al. 2605.21362 null
2026-05-20 SymbolicLight V1: Spike-Gated Dual-Path Language Modeling with High Activation Sparsity and Sub-Billion-Scale Pre-Training Evidence Ting Liu et.al. 2605.21333 null
2026-05-20 TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization Lucheng Fu et.al. 2605.21318 null
2026-05-19 TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload Zhiben Chen et.al. 2605.20179 null
2026-05-19 From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models Juncheng Wu et.al. 2605.20177 null
2026-05-19 ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning Juncheng Wu et.al. 2605.20176 null
2026-05-19 KoRe: Compact Knowledge Representations for Large Language Models Davide Cavicchini et.al. 2605.20170 null
2026-05-19 One in Eight OpenAlex Abstracts Has Integrity Issues Seorin Kim et.al. 2605.20168 null
2026-05-19 CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Hsiang-Wei Huang et.al. 2605.20165 null
2026-05-19 Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models Guangzhi Xiong et.al. 2605.20158 null
2026-05-19 Less Back-and-Forth: A Comparative Study of Structured Prompting Saurav Ghosh et.al. 2605.20149 null
2026-05-19 PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset Haojun Chen et.al. 2605.20147 null
2026-05-19 MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models Yuanqing Cai et.al. 2605.20128 null
2026-05-19 SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction Zhixiong Zhang et.al. 2605.20110 null
2026-05-19 Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding Yuhao Shen et.al. 2605.20104 null
2026-05-19 ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions Chuanyang Jin et.al. 2605.20087 null
2026-05-19 BalanceRAG: Joint Risk Calibration for Cascaded Retrieval-Augmented Generation Zijun Jia et.al. 2605.20084 null
2026-05-19 VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving Zhefan Xu et.al. 2605.20082 null
2026-05-19 CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning Dachuan Shi et.al. 2605.20075 null
2026-05-19 Probing Embodied LLMs: When Higher Observation Fidelity Hurts Problem Solving Oussama Zenkri et.al. 2605.20072 null
2026-05-19 Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP Jann Pfeifer et.al. 2605.20066 null
2026-05-19 Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Wenjie Tang et.al. 2605.20061 null
2026-05-19 PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling Ying-Jia Lin et.al. 2605.20052 null
2026-05-18 DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention Yuxiang Huang et.al. 2605.18753 null
2026-05-18 Traditional statistical representations outperform generative AI in identifying expert peer reviewers Vicente Amado Olivo et.al. 2605.18752 null
2026-05-18 Aurora: Unified Video Editing with a Tool-Using Agent Yongsheng Yu et.al. 2605.18748 null
2026-05-18 Code as Agent Harness Xuying Ning et.al. 2605.18747 null
2026-05-18 SURGE: Approximation-free Training Free Particle Filter for Diffusion Surrogate Lifu Wei et.al. 2605.18745 null
2026-05-18 Actionable World Representation Kunqi Xu et.al. 2605.18743 null
2026-05-18 Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation Qianhao Yuan et.al. 2605.18740 null
2026-05-18 What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models Payal Chandak et.al. 2605.18738 null
2026-05-18 Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory Jinzhuo Liu et.al. 2605.18733 null
2026-05-18 Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency Matthew L. Smith et.al. 2605.18732 null
2026-05-18 General Preference Reinforcement Learning Muhammad Umer et.al. 2605.18721 null
2026-05-18 Democratizing Large-Scale Re-Optimization with LLM-Guided Model Patches Tinghan Ye et.al. 2605.18692 null
2026-05-18 Reversa: A Reverse Documentation Engineering Framework for Converting Legacy Software into Operational Specifications for AI Agents Sanderson Oliveira de Macedo et.al. 2605.18684 null
2026-05-18 CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation Rajeev Goel et.al. 2605.18680 null
2026-05-18 Lance: Unified Multimodal Modeling by Multi-Task Synergy Fengyi Fu et.al. 2605.18678 null
2026-05-18 Generative AI Advertising as a Problem of Trustworthy Commercial Intervention Jingyi Qiu et.al. 2605.18673 null
2026-05-18 Evaluating Multi-turn Human-AI Interaction Shi Ding et.al. 2605.18660 null
2026-05-18 Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees Aditya Tanna et.al. 2605.18654 null
2026-05-18 Language-Switching Triggers Take a Latent Detour Through Language Models Francis Kulumba et.al. 2605.18646 null
2026-05-18 Post-Trained MoE Can Skip Half Experts via Self-Distillation Xingtai Lv et.al. 2605.18643 null
2026-05-18 Are Sparse Autoencoder Benchmarks Reliable? David Chanin et.al. 2605.18229 null
2026-05-18 Context Memorization for Efficient Long Context Generation Yasuyuki Okoshi et.al. 2605.18226 null
2026-05-18 SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning Pawat Chunhachatrachai et.al. 2605.18209 null
2026-05-18 PIPER: Content-Based Table Search via profiling and LLM-Generated Pseudoqueries Riccardo Terrenzi et.al. 2605.18199 null
2026-05-18 Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks Yajing Zhou et.al. 2605.18194 null
2026-05-18 Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method Abdurakhmon Sadiev et.al. 2605.18174 null
2026-05-18 Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models Yanyun Wang et.al. 2605.18168 null
2026-05-18 Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency Junming Liu et.al. 2605.18162 null
2026-05-18 Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models Xinpeng Dong et.al. 2605.18160 null
2026-05-18 FOL2NS: Generating Natural Sentences from First-Order Logic Mei Jia et.al. 2605.18155 null
2026-05-18 Three Heads Are Better Than One: A Multi-perspective Reasoning Framework for Enhanced Vulnerability Detection Xin Peng et.al. 2605.18153 null
2026-05-18 Foundation Models for Credit Risk Prediction: A Game Changer? Bart Baesens et.al. 2605.18147 null
2026-05-18 Evidence-Grounded Frontier Mapping and Agentic Hypothesis Generation in Nanomedicine Christiaan G. A. Viviers et.al. 2605.18144 null
2026-05-18 Generative AI and the Productivity Divide: Human-AI Complementarities in Education Lihi Idan et.al. 2605.18143 null
2026-05-18 A Brief Overview: On-Policy Self-Distillation In Large Language Models Fangming Cui et.al. 2605.18141 null
2026-05-18 Faculty Orientations Shape Adoption of AI in Research and Teaching Timothy J. Atherton et.al. 2605.18140 null
2026-05-18 How Good LLMs Are at Answering Bangla Medical Visual Questions? Dataset and Benchmarking Rafid Ahmed et.al. 2605.18111 null
2026-05-18 Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers Tim Tsz-Kit Lau et.al. 2605.18106 null
2026-05-18 Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction Jiahe Guo et.al. 2605.18104 null
2026-05-18 A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM $Δ$ Integration into Upcycled MoE Hao Zhou et.al. 2605.18083 null
2026-05-14 Articraft: An Agentic System for Scalable Articulated 3D Asset Generation Matt Zhou et.al. 2605.15187 null
2026-05-14 Is Grep All You Need? How Agent Harnesses Reshape Agentic Search Sahil Sen et.al. 2605.15184 null
2026-05-14 Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing Ellwil Sharma et.al. 2605.15179 null
2026-05-14 MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs Rui Wen et.al. 2605.15172 null
2026-05-14 Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval-Augmented Multimodal Alignment Sayantan Kumar et.al. 2605.15168 null
2026-05-14 Does Synthetic Layered Design Data Benefit Layered Design Decomposition? Kam Man Wu et.al. 2605.15167 null
2026-05-14 MeMo: Memory as a Model Ryan Wei Heng Quek et.al. 2605.15156 null
2026-05-14 Pelican-Unified 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action Yi Zhang et.al. 2605.15153 null
2026-05-14 Forgetting That Sticks: Quantization-Permanent Unlearning via Circuit Attribution Saisab Sadhu et.al. 2605.15138 null
2026-05-14 Deep Mixture of Experts Network for Resource Optimization in Aerial-Terrestrial CF-mMIMO Systems under URLLC Donggen Li et.al. 2605.15135 null
2026-05-14 Training ML Models with Predictable Failures Will Schwarzer et.al. 2605.15134 null
2026-05-14 Causal Foundation Models with Continuous Treatments Christopher Stith et.al. 2605.15133 null
2026-05-14 APWA: A Distributed Architecture for Parallelizable Agentic Workflows Evan Rose et.al. 2605.15132 null
2026-05-14 Improving Multi-turn Dialogue Consistency with Self-Recall Thinking Renning Pang et.al. 2605.15102 null
2026-05-14 Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling Rongman Xu et.al. 2605.15100 null
2026-05-14 On the Cultural Anachronism and Temporal Reasoning in Vision Language Models Mukul Ranjan et.al. 2605.15071 null
2026-05-14 LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection Mitchell Piehl et.al. 2605.15054 null
2026-05-14 TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale Anurup Ganguli et.al. 2605.15053 null
2026-05-14 An Interpretable Latency Model for Speculative Decoding in LLM Serving Linghao Kong et.al. 2605.15051 null
2026-05-14 SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning KiHyun Nam et.al. 2605.15044 null
2026-05-13 WARDEN: Endangered Indigenous Language Transcription and Translation with 6 Hours of Training Data Ziheng Zhang et.al. 2605.13846 null
2026-05-13 Unlocking Patch-Level Features for CLIP-Based Class-Incremental Learning Hao Sun et.al. 2605.13835 null
2026-05-13 Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context Zhaowei Wang et.al. 2605.13831 null
2026-05-13 Neurosymbolic Auditing of Natural-Language Software Requirements Bethel Hall et.al. 2605.13817 null
2026-05-13 Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling Deepak Pandita et.al. 2605.13801 null
2026-05-13 An LLM-Based System for Argument Reconstruction Paulo Pirozelli et.al. 2605.13793 null
2026-05-13 ENSEMBITS: an alphabet of protein conformational ensembles Kaiwen Shi et.al. 2605.13789 null
2026-05-13 LMPath: Language-Mediated Priors and Path Generation for Aerial Exploration Jonathan A. Diller et.al. 2605.13782 null
2026-05-13 RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data Harold Haodong Chen et.al. 2605.13775 null
2026-05-13 (How) Do Large Language Models Understand High-Level Message Sequence Charts? Mohammad Reza Mousavi et.al. 2605.13773 null
2026-05-13 Where Does Reasoning Break? Step-Level Hallucination Detection via Hidden-State Transport Geometry Tyler Alvarez et.al. 2605.13772 null
2026-05-13 Dense vs Sparse Pretraining at Tiny Scale: Active-Parameter vs Total-Parameter Matching Abdalrahman Wael et.al. 2605.13769 null
2026-05-13 EconAI: Dynamic Persona Evolution and Memory-Aware Agents in Evolving Economic Environments Annie Liu et.al. 2605.13762 null
2026-05-13 Learning POMDP World Models from Observations with Language-Model Priors Valentin Six et.al. 2605.13740 null
2026-05-13 Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs Trung Nguyen Quang et.al. 2605.13737 null
2026-05-13 ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles Yitian Yang et.al. 2605.13725 null
2026-05-13 SkillOps: Managing LLM Agent Skill Libraries as Self-Maintaining Software Ecosystems Hongji Pu et.al. 2605.13716 null
2026-05-13 MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling Hsing-Huan Chung et.al. 2605.13711 null
2026-05-13 Children’s English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety Qian Shen et.al. 2605.13709 null
2026-05-13 Identifying AI Web Scrapers Using Canary Tokens Steven Seiden et.al. 2605.13706 null
2026-05-12 SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Haiwen Diao et.al. 2605.12500 null
2026-05-12 Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation Kexuan Shi et.al. 2605.12492 null
2026-05-12 Learning, Fast and Slow: Towards LLMs That Adapt Continually Rishabh Tiwari et.al. 2605.12484 null
2026-05-12 Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training Yuanda Xu et.al. 2605.12483 null
2026-05-12 Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts Sagi Ahrac et.al. 2605.12476 null
2026-05-12 Solve the Loop: Attractor Models for Language and Reasoning Jacob Fein-Ashley et.al. 2605.12466 null
2026-05-12 Search Your Block Floating Point Scales! Tanmaey Gupta et.al. 2605.12464 null
2026-05-12 Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs Guinan Su et.al. 2605.12460 null
2026-05-12 TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection Tom Sander et.al. 2605.12456 null
2026-05-12 The Algorithmic Caricature: Auditing LLM-Generated Political Discourse Across Crisis Events Gunjan et.al. 2605.12452 null
2026-05-12 ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models Chen Li et.al. 2605.12446 null
2026-05-12 A Causal Language Modeling Detour Improves Encoder Continued Pretraining Rian Touchent et.al. 2605.12438 null
2026-05-12 Geometric Factual Recall in Transformers Shauli Ravfogel et.al. 2605.12426 null
2026-05-12 Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals Yo Ehara et.al. 2605.12422 null
2026-05-12 Formalize, Don’t Optimize: The Heuristic Trap in LLM-Generated Combinatorial Solvers Haoyu Wang et.al. 2605.12421 null
2026-05-12 ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging Neha Verma et.al. 2605.12419 null
2026-05-12 Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images Yuangong Chen et.al. 2605.12413 null
2026-05-12 Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Eric Bigelow et.al. 2605.12412 null
2026-05-12 Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems William Parris et.al. 2605.12406 null
2026-05-12 OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning Yuxiao Yang et.al. 2605.12400 null
2026-05-11 ELF: Embedded Language Flows Keya Hu et.al. 2605.10938 null
2026-05-11 Personal Visual Context Learning in Large Multimodal Models Zihui Xue et.al. 2605.10936 null
2026-05-11 DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices Chenyang Song et.al. 2605.10933 null
2026-05-11 Evaluating the False Trust engendered by LLM Explanations Vardhan Palod et.al. 2605.10930 null
2026-05-11 Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning Junhao Shen et.al. 2605.10923 null
2026-05-11 RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark Huashuo Lei et.al. 2605.10921 null
2026-05-11 WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Shuangrui Ding et.al. 2605.10912 null
2026-05-11 Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers Nikita Kezins et.al. 2605.10901 null
2026-05-11 Grounded or Guessing? LVLM Confidence Estimation via Blind-Image Contrastive Ranking Reza Khanmohammadi et.al. 2605.10893 null
2026-05-11 Count Anything at Any Granularity Chang Liu et.al. 2605.10887 null
2026-05-11 LoKA: Low-precision Kernel Applications for Recommendation Models At Scale Liang Luo et.al. 2605.10886 null
2026-05-11 Compute Where it Counts: Self Optimizing Language Models Yash Akhauri et.al. 2605.10875 null
2026-05-11 BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD Haozhe Zhang et.al. 2605.10865 null
2026-05-11 DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization Mengyi Deng et.al. 2605.10863 null
2026-05-11 RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems Joel Rorseth et.al. 2605.10862 null
2026-05-11 Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding Jianzhu Bao et.al. 2605.10855 null
2026-05-11 Grounded Satirical Generation with RAG Oona Itkonen et.al. 2605.10853 null
2026-05-11 Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA Ruinan Jin et.al. 2605.10850 null
2026-05-11 Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient? Tz-Huan Hsu et.al. 2605.10848 null
2026-05-11 Training-Free Cultural Alignment of Large Language Models via Persona Disagreement Huynh Trung Kiet et.al. 2605.10843 null
2026-05-11 Extending Confidence-Based Text2Cypher with Grammar and Schema Aware Filtering Makbule Gulcin Ozsoy et.al. 2605.10318 null
2026-05-11 AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks Baraa Al Jorf et.al. 2605.10286 null
2026-05-11 DP-LAC: Lightweight Adaptive Clipping for Differentially Private Federated Fine-tuning of Language Models Haaris Mehmood et.al. 2605.10272 null
2026-05-11 Teaching LLMs to See Graphs: Unifying Text and Structural Reasoning Dario Vajda et.al. 2605.10247 null
2026-05-11 Route Before Retrieve: Activating Latent Routing Abilities of LLMs for RAG vs. Long-Context Selection Yiwen Chen et.al. 2605.10235 null
2026-05-11 Social Policy of Large Language Models: How GPT, Claude, DeepSeek and Grok Allocate Social Budgets in Spain and Germany Claudia Benavides Cantos et.al. 2605.10234 null
2026-05-11 FORGE: Fragment-Oriented Ranking and Generation for Context-Aware Molecular Optimization Qingchuan Zhang et.al. 2605.10230 null
2026-05-11 FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries Qijie You et.al. 2605.10228 null
2026-05-11 Hypothesis-Driven Deep Research with Large Language Models: A Structured Methodology for Automated Knowledge Discovery Michael Chin et.al. 2605.10224 null
2026-05-11 Beyond Autonomy: A Dynamic Tiered AgentRunner Framework for Governable and Resilient Enterprise AI Execution Kai Pan et.al. 2605.10223 null
2026-05-11 Relative Score Policy Optimization for Diffusion Language Models Zichao Yu et.al. 2605.10218 null
2026-05-11 The Impact of Editorial Intervention on Detecting Native Language Traces Ahmet Yavuz Uluslu et.al. 2605.10216 null
2026-05-11 To Redact, or not to Redact? A Local LLM Approach to Deliberative Process Privilege Classification Maik Larooij et.al. 2605.10211 null
2026-05-11 LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation Yiwen Chen et.al. 2605.10207 null
2026-05-11 How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue Hui Lu et.al. 2605.10199 null
2026-05-11 Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Shuzhang Zhong et.al. 2605.10195 null
2026-05-11 ProteinOPD: Towards Effective and Efficient Preference Alignment for Protein Design Yulin Zhang et.al. 2605.10189 null
2026-05-11 SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation Longteng Guo et.al. 2605.10187 null
2026-05-11 LegalCiteBench: Evaluating Citation Reliability in Legal Language Models Sijia Chen et.al. 2605.10186 null
2026-05-11 When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications Farzad Nourmohammadzadeh Motlagh et.al. 2605.10176 null
2026-05-08 LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling Tong Zheng et.al. 2605.08083 null
2026-05-08 Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment Jerry Jiang et.al. 2605.08064 null
2026-05-08 Flow-OPD: On-Policy Distillation for Flow Matching Models Zhen Fang et.al. 2605.08063 null
2026-05-08 The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents Jiayuan Liu et.al. 2605.08060 null
2026-05-08 CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocation James Petullo et.al. 2605.08057 null
2026-05-08 Fast Byte Latent Transformer Julie Kallini et.al. 2605.08044 null
2026-05-08 Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph Ning Liu et.al. 2605.08037 null
2026-05-08 Weak Order on the MacNeille Completion of Bruhat Order Colin Defant et.al. 2605.08033 null
2026-05-08 Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models Kaidi Jia et.al. 2605.08031 null
2026-05-08 STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation Ying Shen et.al. 2605.08029 null
2026-05-08 Abductive Reasoning with Probabilistic Commonsense Joseph Cotnareanu et.al. 2605.08011 null
2026-05-08 SphereVAD: Training-Free Video Anomaly Detection via Geodesic Inference on the Unit Hypersphere Chao Huang et.al. 2605.08003 null
2026-05-08 Semantic Smoothing for Language Models via Distribution Estimation and Embeddings Haricharan Balasundaram et.al. 2605.07994 null
2026-05-08 Tool Calling is Linearly Readable and Steerable in Language Models Zekun Wu et.al. 2605.07990 null
2026-05-08 Where’s the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions Nicole Ma et.al. 2605.07984 null
2026-05-08 GLiGuard: Schema-Conditioned Classification for LLM Safeguard Urchade Zaratiana et.al. 2605.07982 null
2026-05-08 Graph Representation Learning Augmented Model Manipulation on Federated Fine-Tuning of LLMs Hanlin Cai et.al. 2605.07961 null
2026-05-08 Similar Pattern Annotation via Retrieval Knowledge for LLM-Based Test Code Fault Localization Golnaz Gharachorlu et.al. 2605.07957 null
2026-05-08 Prototype Guided Post-pretraining for Single-Cell Representation Learning Sachini Weerasekara et.al. 2605.07938 null
2026-05-08 TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples Shuren Xia et.al. 2605.07935 null
2026-05-07 UniPool: A Globally Shared Expert Pool for Mixture-of-Experts Minbin Huang et.al. 2605.06665 null
2026-05-07 EMO: Pretraining Mixture of Experts for Emergent Modularity Ryan Wang et.al. 2605.06663 null
2026-05-07 Verifier-Backed Hard Problem Generation for Mathematical Reasoning Yuhang Lai et.al. 2605.06660 null
2026-05-07 Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Yuxing Liu et.al. 2605.06654 null
2026-05-07 When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels Sushant Gautam et.al. 2605.06652 null
2026-05-07 Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Mingwei Xu et.al. 2605.06650 null
2026-05-07 Edge-specific signal propagation on mature chromophore-region 3D mechanism graphs for fluorescent protein quantum-yield prediction Yuchen Xiong et.al. 2605.06644 null
2026-05-07 StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction Xiangyuan Xue et.al. 2605.06642 null
2026-05-07 GlazyBench: A Benchmark for Ceramic Glaze Property Prediction and Image Generation Ziyu Zhai et.al. 2605.06641 null
2026-05-07 Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key Tianle Wang et.al. 2605.06638 null
2026-05-07 Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents Hailey Onweller et.al. 2605.06635 null
2026-05-07 Crafting Reversible SFT Behaviors in Large Language Models Yuping Lin et.al. 2605.06632 null
2026-05-07 Task-Aware Answer Preservation under Audio Compression for Large Audio Language Models Amir Ivry et.al. 2605.06631 null
2026-05-07 MASPO: Joint Prompt Optimization for LLM-based Multi-Agent Systems Zhexuan Wang et.al. 2605.06623 null
2026-05-07 Algospeak, Hiding in the Open: The Trade-off Between Legible Meaning and Detection Avoidance Jan Fillies et.al. 2605.06619 null
2026-05-07 The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity Siquan Li et.al. 2605.06611 null
2026-05-07 SoftSAE: Dynamic Top-K Selection for Adaptive Sparse Autoencoders Jakub Stępień et.al. 2605.06610 null
2026-05-07 Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent Chenyang Zhang et.al. 2605.06609 null
2026-05-07 How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Shai Feldman et.al. 2605.06605 null
2026-05-07 Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches Isaac David et.al. 2605.06601 null
2026-05-06 Implicit Representations of Grammaticality in Language Models Yingshan Susan Wang et.al. 2605.05197 null
2026-05-06 Almost-Orthogonality in Lp Spaces: A Case Study with Grok Ziang Chen et.al. 2605.05192 null
2026-05-06 Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer Alexander Hsu et.al. 2605.05176 null
2026-05-06 The First Token Knows: Single-Decode Confidence for Hallucination Detection Mina Gabriel et.al. 2605.05166 null
2026-05-06 Geometry-Aware State Space Model: A New Paradigm for Whole-Slide Image Representation Enhui Chai et.al. 2605.05164 null
2026-05-06 Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging Bernhard Kainz et.al. 2605.05161 null
2026-05-06 PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation Srikar Kashyap Pulipaka et.al. 2605.05159 null
2026-05-06 Superposition Is Not Necessary: A Mechanistic Interpretability Analysis of Transformer Representations for Time Series Forecasting Alper Yıldırım et.al. 2605.05151 null
2026-05-06 Low-Cost Black-Box Detection of LLM Hallucinations via Dynamical System Prediction Dan Wilson et.al. 2605.05134 null
2026-05-06 Joint Treatment Effect Estimation from Incomplete Healthcare Data: Temporal Causal Normalizing Flows with LLM-driven Evolutionary MNAR Imputation Olivia Jullian Parra et.al. 2605.05125 null
2026-05-06 Beyond Semantics: An Evidential Reasoning-Aware Multi-View Learning Framework for Trustworthy Mental Health Prediction Yucheng Ruan et.al. 2605.05121 null
2026-05-06 On the Hardness of Junking LLMs Marco Rando et.al. 2605.05116 null
2026-05-06 Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Daniel Wurgaft et.al. 2605.05115 null
2026-05-06 Think-Aloud Reshapes Automated Cognitive Model Discovery Beyond Behavior Hanbo Xie et.al. 2605.05091 null
2026-05-06 Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models Quintin Pope et.al. 2605.05090 null
2026-05-06 The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences Hubert Plisiecki et.al. 2605.05080 null
2026-05-06 Heterogeneous Judge-Aware Ranking with Sensitivity, Disagreement, and Confidence Shibo Yu et.al. 2605.05073 null
2026-05-06 SoK: Robustness in Large Language Models against Jailbreak Attacks Feiyue Xu et.al. 2605.05058 null
2026-05-06 Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation Hongxu Chen et.al. 2605.05054 null
2026-05-06 Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism Sajal Dash et.al. 2605.05049 null
2026-05-05 Large Language Models are Universal Reasoners for Visual Generation Sucheng Ren et.al. 2605.04040 null
2026-05-05 Safety and accuracy follow different scaling laws in clinical large language models Sebastian Wind et.al. 2605.04039 null
2026-05-05 OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories Yuwen Du et.al. 2605.04036 null
2026-05-05 Stayin’ Aligned Over Time: Towards Longitudinal Human-LLM Alignment via Contextual Reflection and Privacy-Preserving Behavioral Data Simret Araya Gebreegziabher et.al. 2605.04029 null
2026-05-05 SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment Joseph Breda et.al. 2605.04012 null
2026-05-05 Physics-Grounded Multi-Agent Architecture for Traceable, Risk-Aware Human-AI Decision Support in Manufacturing Danny Hoang et.al. 2605.04003 null
2026-05-05 RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction Renjie He et.al. 2605.03999 null
2026-05-05 EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage Richard J. Young et.al. 2605.03998 null
2026-05-05 Logical Consistency as a Bridge: Improving LLM Hallucination Detection via Label Constraint Modeling between Responses and Self-Judgments Hao Mi et.al. 2605.03971 null
2026-05-05 Generating Proof-of-Vulnerability Tests to Help Enhance the Security of Complex Software Shravya Kanchi et.al. 2605.03956 null
2026-05-05 Beyond Rules: LLM-Powered Linting for Quantum Programs Pietro Cassieri et.al. 2605.03943 null
2026-05-05 MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model Jingyao Gong et.al. 2605.03937 null
2026-05-05 The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Models Daniel Drucker et.al. 2605.03936 null
2026-05-05 StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning Xiaowen Sun et.al. 2605.03927 null
2026-05-05 Atomic Fact-Checking Increases Clinician Trust in Large Language Model Recommendations for Oncology Decision Support: A Randomized Controlled Trial Lisa C. Adams et.al. 2605.03916 null
2026-05-05 Task-Aware Scanning Parameter Configuration for Robotic Inspection Using Vision Language Embeddings and Hyperdimensional Computing Zhiling Chen et.al. 2605.03909 null
2026-05-05 Steer Like the LLM: Activation Steering that Mimics Prompting Geert Heyman et.al. 2605.03907 null
2026-05-05 Deco: Extending Personal Physical Objects into Pervasive AI Companion through a Dual-Embodiment Framework Zhihan Jiang et.al. 2605.03882 null
2026-05-05 EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics Shuyue Stella Li et.al. 2605.03871 null
2026-05-05 On Adaptivity in Zeroth-Order Optimization Hassan Dbouk et.al. 2605.03869 null
2026-05-04 Semantic Risk-Aware Heuristic Planning for Robotic Navigation in Dynamic Environments: An LLM-Inspired Approach Hamza Ahmed Durrani et.al. 2605.02862 null
2026-05-04 Standing on the Shoulders of Giants: Stabilized Knowledge Distillation for Cross–Language Code Clone Detection Mohamad Khajezade et.al. 2605.02860 null
2026-05-04 Trust, but Verify: Peeling Low-Bit Transformer Networks for Training Monitoring Arian Eamaz et.al. 2605.02853 null
2026-05-04 VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition Tanush Yadav et.al. 2605.02834 null
2026-05-04 When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs Haorui Li et.al. 2605.02821 null
2026-05-04 SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering Jiujiu Chen et.al. 2605.02819 null
2026-05-04 Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces Chenchen Zhang et.al. 2605.02801 null
2026-05-04 FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework Mario Rodríguez Béjar et.al. 2605.02789 null
2026-05-04 When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition Pehuén Moure et.al. 2605.02782 null
2026-05-04 Mitigating Misalignment Contagion by Steering with Implicit Traits Maria Chang et.al. 2605.02751 null
2026-05-04 Bolek: A Multimodal Language Model for Molecular Reasoning Frederic Grabowski et.al. 2605.02745 null
2026-05-04 AI-Generated Smells: An Analysis of Code and Architecture in LLM and Agent-Driven Development Yuecai Zhu et.al. 2605.02741 null
2026-05-04 Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs Xin Zhang et.al. 2605.02735 null
2026-05-04 Perceptual Flow Network for Visually Grounded Reasoning Yangfu Li et.al. 2605.02730 null
2026-05-04 PubMed-Ophtha: An open resource for training ophthalmology vision-language models on scientific literature Verena Jasmin Hallitschke et.al. 2605.02720 null
2026-05-04 Hybrid Inspection and Task-Based Access Control in Zero-Trust Agentic AI Majed El Helou et.al. 2605.02682 null
2026-05-04 Fuzzy Fingerprinting Encoder Pre-trained Language Models for Emotion Recognition in Conversations: Human Assessment and Validity Study Patrícia Pereira et.al. 2605.02665 null
2026-05-04 ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming Mario Rodríguez Béjar et.al. 2605.02647 null
2026-05-04 Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE Yangming Shi et.al. 2605.02641 null
2026-05-04 AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding Ruilin Yao et.al. 2605.02630 null
2026-05-01 When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models Sailesh Panda et.al. 2605.00817 null
2026-05-01 Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs Siyuan Huang et.al. 2605.00814 null
2026-05-01 Let ViT Speak: Generative Language-Image Pre-training Yan Fang et.al. 2605.00809 null
2026-05-01 Can Coding Agents Reproduce Findings in Computational Materials Science? Ziyang Huang et.al. 2605.00803 null
2026-05-01 GMGaze: MoE-Based Context-Aware Gaze Estimation with CLIP and Multiscale Transformer Xinyuan Zhao et.al. 2605.00799 null
2026-05-01 RunAgent: Interpreting Natural-Language Plans with Constraint-Guided Execution Arunabh Srivastava et.al. 2605.00798 null
2026-05-01 Make Your LVLM KV Cache More Lightweight Xihao Chen et.al. 2605.00789 null
2026-05-01 Characterizing the Expressivity of Local Attention in Transformers Jiaoda Li et.al. 2605.00768 null
2026-05-01 Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Indraneil Paul et.al. 2605.00754 null
2026-05-01 Self-Adaptive Multi-Agent LLM-Based Security Pattern Selection for IoT Systems Saeid Jamshidi et.al. 2605.00741 null
2026-05-01 FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios Yutao Hou et.al. 2605.00706 null
2026-05-01 Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory Derong Xu et.al. 2605.00702 null
2026-05-01 STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack Xutao Mao et.al. 2605.00699 null
2026-05-01 Adaptive Querying with AI Persona Priors Kaizheng Wang et.al. 2605.00696 null
2026-05-01 Learning to Act and Cooperate for Distributed Black-Box Consensus Optimization Zi-Bo Qin et.al. 2605.00691 null
2026-05-01 ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models Yunhan Zhao et.al. 2605.00689 null
2026-05-01 Eliminating Hidden Serialization in Multi-Node Megakernel Communication Byungsoo Oh et.al. 2605.00686 null
2026-05-01 Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game Lixing Li et.al. 2605.00677 null
2026-05-01 Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs Jasper Dekoninck et.al. 2605.00674 null
2026-05-01 SENECA: Small-Sample Discrete Entropy Estimation via Self-Consistent Missing Mass Lucas H. McCabe et.al. 2605.00668 null
2026-04-30 HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation Xin Zhou et.al. 2604.28196 null
2026-04-30 LaST-R1: Reinforcing Action via Adaptive Physical Latent Reasoning for VLA Models Hao Chen et.al. 2604.28192 null
2026-04-30 Exploration Hacking: Can LLMs Learn to Resist RL Training? Eyon Jang et.al. 2604.28182 null
2026-04-30 LLM as Clinical Graph Structure Refiner: Enhancing Representation Learning in EEG Seizure Diagnosis Lincan Li et.al. 2604.28178 null
2026-04-30 AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images Bo Zhang et.al. 2604.28177 null
2026-04-30 PhyCo: Learning Controllable Physical Priors for Generative Motion Sriram Narayanan et.al. 2604.28169 null
2026-04-30 FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption Yanting Wang et.al. 2604.28157 null
2026-04-30 On the Proper Treatment of Units in Surprisal Theory Samuel Kiegeland et.al. 2604.28147 null
2026-04-30 PRISM: Pre-alignment via Black-box On-policy Distillation for Multimodal Reinforcement Learning Sudong Wang et.al. 2604.28123 null
2026-04-30 FreeOcc: Training-Free Embodied Open-Vocabulary Occupancy Prediction Zeyu Jiang et.al. 2604.28115 null
2026-04-30 What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design Ivan Bercovich et.al. 2604.28093 null
2026-04-30 Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles Zainab Rehan et.al. 2604.28087 null
2026-04-30 Characterizing the Consistency of the Emergent Misalignment Persona Anietta Weckauff et.al. 2604.28082 null
2026-04-30 TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering An-Yang Ji et.al. 2604.28076 null
2026-04-30 Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling Ansar Aynetdinov et.al. 2604.28075 null
2026-04-30 RHyVE: Competence-Aware Verification and Phase-Aware Deployment for LLM-Generated Reward Hypotheses Feiyu Wu et.al. 2604.28056 null
2026-04-30 Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception Neemias B da Silva et.al. 2604.28048 null
2026-04-30 Collaborative Agent Reasoning Engineering (CARE): A Three-Party Design Methodology for Systematically Engineering AI Agents with Subject Matter Experts, Developers, and Helper Agents Rahul Ramachandran et.al. 2604.28043 null
2026-04-30 SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images Jialu Shen et.al. 2604.28039 null
2026-04-30 Models Recall What They Violate: Constraint Adherence in Multi-Turn LLM Ideation Garvin Kruthof et.al. 2604.28031 null
2026-04-29 Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models Gongbo Zhang et.al. 2604.26951 null
2026-04-29 Three-Step Nav: A Hierarchical Global-Local Planner for Zero-Shot Vision-and-Language Navigation Wanrong Zheng et.al. 2604.26946 null
2026-04-29 Select to Think: Unlocking SLM Potential with Local Sufficiency Wenxuan Ye et.al. 2604.26940 null
2026-04-29 World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning Wanyue Zhang et.al. 2604.26934 null
2026-04-29 FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving Minghe Wang et.al. 2604.26881 null
2026-04-29 HealthNLP_Retrievers at ArchEHR-QA 2026: Cascaded LLM Pipeline for Grounded Clinical Question Answering Md Biplob Hosen et.al. 2604.26880 null
2026-04-29 MoRFI: Monotonic Sparse Autoencoder Feature Identification Dimitris Dimakopoulos et.al. 2604.26866 null
2026-04-29 Cognitive Atrophy and Systemic Collapse in AI-Dependent Software Engineering Frank Ginac et.al. 2604.26855 null
2026-04-29 What Kind of Language is Easy to Language-Model Under Curriculum Learning? Nadine El-Naggar et.al. 2604.26844 null
2026-04-29 Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance Lingfeng Zhang et.al. 2604.26839 null
2026-04-29 Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel Yiqi Liu et.al. 2604.26821 null
2026-04-29 Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding Hayate Iso et.al. 2604.26779 null
2026-04-29 Domain-Adapted Small Language Models for Reliable Clinical Triage Manar Aljohani et.al. 2604.26766 null
2026-04-29 Factorized Latent Reasoning for LLM-based Recommendation Tianqi Gao et.al. 2604.26760 null
2026-04-29 GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents GLM-V Team et.al. 2604.26752 null
2026-04-29 FutureWorld: A Live Environment for Training Predictive Agents with Real-World Outcome Rewards Zhixin Han et.al. 2604.26733 null
2026-04-29 COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Akhmed Sakip et.al. 2604.26687 null
2026-04-29 When Model Editing Meets Service Evolution: A Knowledge-Update Perspective for Service Recommendation Guodong Fan et.al. 2604.26686 null
2026-04-29 Differentially-Private Text Rewriting reshapes Linguistic Style Stefan Arnold et.al. 2604.26656 null
2026-04-29 AgentSim: A Platform for Verifiable Agent-Trace Simulation Saber Zerhoudi et.al. 2604.26653 null
2026-04-28 Recursive Multi-Agent Systems Xiyuan Yang et.al. 2604.25917 null
2026-04-28 Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language Models Ajmain Inqiad Alam et.al. 2604.25903 null
2026-04-28 Three Models of RLHF Annotation: Extension, Evidence, and Authority Steve Coyne et.al. 2604.25895 null
2026-04-28 Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers Jan Dubiński et.al. 2604.25891 null
2026-04-28 QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding Shuxiang Cao et.al. 2604.25884 null
2026-04-28 When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Shuning Shang et.al. 2604.25872 null
2026-04-28 From Syntax to Emotion: A Mechanistic Analysis of Emotion Inference in LLMs Bangzhao Shu et.al. 2604.25866 null
2026-04-28 Luminol-AIDetect: Fast Zero-shot Machine-Generated Text Detection based on Perplexity under Text Shuffling Lucio La Cava et.al. 2604.25860 null
2026-04-28 Investigation into In-Context Learning Capabilities of Transformers Rushil Chandrupatla et.al. 2604.25858 null
2026-04-28 SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring Hector G. Rodriguez et.al. 2604.25855 null
2026-04-28 G-Loss: Graph-Guided Fine-Tuning of Language Models Sharma Aditya et.al. 2604.25853 null
2026-04-28 From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling Jianghao Lin et.al. 2604.25847 null
2026-04-28 Towards Agentic Investigation of Security Alerts Even Eilertsen et.al. 2604.25846 null
2026-04-28 Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning Yashwant Pravinrao Bangde et.al. 2604.25809 null
2026-04-28 Barriers to Universal Reasoning With Transformers (And How to Overcome Them) Oliver Kraus et.al. 2604.25800 null
2026-04-28 Subliminal Steering: Stronger Encoding of Hidden Signals George Morgulis et.al. 2604.25783 null
2026-04-28 CGU-ILALab at FoodBench-QA 2026: Comparing Traditional and LLM-based Approaches for Recipe Nutrient Estimation Wei-Chun Chen et.al. 2604.25774 null
2026-04-28 SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing? Noam Tarshish et.al. 2604.25737 null
2026-04-28 Toward Multimodal Conversational AI for Age-Related Macular Degeneration Ran Gu et.al. 2604.25720 null
2026-04-28 Step-Audio-R1.5 Technical Report Yuxin Zhang et.al. 2604.25719 null
2026-04-27 AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents Yixiang Zhang et.al. 2604.24657 null
2026-04-27 DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference Zahra Dehghanighobadi et.al. 2604.24647 null
2026-04-27 K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology Soyeon Kim et.al. 2604.24645 null
2026-04-27 Cortex-Inspired Continual Learning: Unsupervised Instantiation and Recovery of Functional Task Networks Kevin McKee et.al. 2604.24637 null
2026-04-27 Less Is More: Engineering Challenges of On-Device Small Language Model Integration in a Mobile Application William Oliveira et.al. 2604.24636 null
2026-04-27 Meta-CoT: Enhancing Granularity and Generalization in Image Editing Shiyi Zhang et.al. 2604.24625 null
2026-04-27 XGRAG: A Graph-Native Framework for Explaining KG-based Retrieval-Augmented Generation Zhuoling Li et.al. 2604.24623 null
2026-04-27 Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions Utku Boran Torun et.al. 2604.24621 null
2026-04-27 Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models Yuxing Tian et.al. 2604.24608 null
2026-04-27 Majorization-Guided Test-Time Adaptation for Vision-Language Models under Modality-Specific Shift Lixian Chen et.al. 2604.24602 null
2026-04-27 Skill Retrieval Augmentation for Agentic AI Weihang Su et.al. 2604.24594 null
2026-04-27 A systematic evaluation of vision-language models for observational astronomical reasoning tasks Wenke Ren et.al. 2604.24589 null
2026-04-27 Improving Vision-language Models with Perception-centric Process Reward Models Yingqian Min et.al. 2604.24583 null
2026-04-27 Measuring the Unmeasurable: Markov Chain Reliability for LLM Agents Phat T. Tran-Truong et.al. 2604.24579 null
2026-04-27 MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG Xihang Wang et.al. 2604.24564 null
2026-04-27 Towards Lawful Autonomous Driving: Deriving Scenario-Aware Driving Requirements from Traffic Laws and Regulations Bowen Jian et.al. 2604.24562 null
2026-04-27 STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator Alessio Sordo et.al. 2604.24544 null
2026-04-27 Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models Nay Myat Min et.al. 2604.24542 null
2026-04-27 Generating Place-Based Compromises Between Two Points of View Sumanta Bhattacharyya et.al. 2604.24536 null
2026-04-27 Why AI Harms Can’t Be Fixed One Identity at a Time: What 5300 Incident Reports Reveal About Intersectionality Edyta Bogucka et.al. 2604.24519 null
2026-04-24 Representational Harms in LLM-Generated Narratives Against Global Majority Nationalities Ilana Nguyen et.al. 2604.22749 null
2026-04-24 Code for All: Educational Applications of the “Vibe Coding” Hackathon in Programming Education across All Skill Levels Ashley J. Chen et.al. 2604.22747 null
2026-04-24 Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought Keshav Ramji et.al. 2604.22709 null
2026-04-24 BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering Jinghong Chen et.al. 2604.22678 null
2026-04-24 Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines Negar Arabzadeh et.al. 2604.22661 null
2026-04-24 RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices Jia Li et.al. 2604.22659 null
2026-04-24 GazeVLA: Learning Human Intention for Robotic Manipulation Chengyang Li et.al. 2604.22615 null
2026-04-24 Dharma, Data and Deception: An LLM-Powered Rhetorical Analysis of Cow-Urine Health Claims on YouTube Sheza Munir et.al. 2604.22606 null
2026-04-24 Vibe coding for clinicians: democratising bespoke software development for digital health innovation Ariel Yuhan Ong et.al. 2604.22604 null
2026-04-24 From Natural Language to Verified Code: Toward AI Assisted Problem-to-Code Generation with Dafny-Based Formal Verification Md Erfan et.al. 2604.22601 null
2026-04-24 Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity Erez Yosef et.al. 2604.22597 null
2026-04-24 LARA: Validation-Driven Agentic Supercomputer Workflows for Atomistic Modeling William Dawson et.al. 2604.22571 null
2026-04-24 Learning Evidence Highlighting for Frozen LLMs Shaoang Li et.al. 2604.22565 null
2026-04-24 SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning Jichao Wang et.al. 2604.22558 null
2026-04-24 Controllable Spoken Dialogue Generation: An LLM-Driven Grading System for K-12 Non-Native English Learners Haidong Yuan et.al. 2604.22542 null
2026-04-24 Dr.Sai: An agentic AI for real-world physics analysis at BESIII Mingfeng He et.al. 2604.22541 null
2026-04-24 FeatEHR-LLM: Leveraging Large Language Models for Feature Engineering in Electronic Health Records Hojjat Karami et.al. 2604.22534 null
2026-04-24 RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment Yingfeng Luo et.al. 2604.22520 null
2026-04-24 Benchmarking LLM-Driven Network Configuration Repair Ioannis Protogeros et.al. 2604.22513 null
2026-04-24 Objective Shaping with Hard Negatives: Windowed Partial AUC Optimization for RL-based LLM Recommenders Wentao Shi et.al. 2604.22504 null
2026-04-23 Evaluation of Automatic Speech Recognition Using Generative Large Language Models Thibault Bañeras-Roux et.al. 2604.21928 null
2026-04-23 Seeing Without Eyes: 4D Human-Scene Understanding from Wearable IMUs Hao-Yu Hsu et.al. 2604.21926 null
2026-04-23 MathDuels: Evaluating LLMs as Problem Posers and Solvers Zhiqiu Xu et.al. 2604.21916 null
2026-04-23 When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs Pegah Khayatan et.al. 2604.21911 null
2026-04-23 Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models Chee Wei Tan et.al. 2604.21896 null
2026-04-23 EVENT5Ws: A Large Dataset for Open-Domain Event Extraction from Documents Praval Sharma et.al. 2604.21890 null
2026-04-23 TingIS: Real-time Risk Event Discovery from Noisy Customer Incidents at Enterprise Scale Jun Wang et.al. 2604.21889 null
2026-04-23 A Multimodal Text- and Graph-Based Approach for Open-Domain Event Extraction from Documents Praval Sharma et.al. 2604.21885 null
2026-04-23 Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms Yuto Nishida et.al. 2604.21882 null
2026-04-23 Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions Jiseon Kim et.al. 2604.21871 null
2026-04-23 FAccT-Checked: A Narrative Review of Authority Reconfigurations and Retention in AI-Mediated Journalism Stefano Sorrentino et.al. 2604.21864 null
2026-04-23 Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models Naheed Rayhan et.al. 2604.21860 null
2026-04-23 OptiMat Alloys: A FAIR End-to-End Agent with Living Database for Computational Multi-Principal Alloy Exploration Yang Hu et.al. 2604.21850 null
2026-04-23 Modulating Cross-Modal Convergence with Single-Stimulus, Intra-Modal Dispersion Eghbal A. Hosseini et.al. 2604.21836 null
2026-04-23 Tool Attention Is All You Need: Dynamic Tool Gating and Lazy Schema Loading for Eliminating the MCP/Tools Tax in Scalable Agentic Workflows Anuj Sadani et.al. 2604.21816 null
2026-04-23 Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems Ye Yu et.al. 2604.21794 null
2026-04-23 Agentic AI-Enabled Framework for Thermal Comfort and Building Energy Assessment in Tropical Urban Neighborhoods Po-Yen Lai et.al. 2604.21787 null
2026-04-23 From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media Katharina Prasse et.al. 2604.21786 null
2026-04-23 Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models Larissa Höfling et.al. 2604.21780 null
2026-04-23 Misinformation Span Detection in Videos via Audio Transcripts Breno Matos et.al. 2604.21767 null
2026-04-22 SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation Ruohan Liu et.al. 2604.20842 null
2026-04-22 Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL Zhaofeng Wu et.al. 2604.20835 null
2026-04-22 PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance Yupeng Zheng et.al. 2604.20834 null
2026-04-22 AVISE: Framework for Evaluating the Security of AI Systems Mikko Lempinen et.al. 2604.20833 null
2026-04-22 Stream-CQSA: Avoiding Out-of-Memory in Attention Computation via Flexible Workload Scheduling Yiming Bian et.al. 2604.20819 null
2026-04-22 Convergent Evolution: How Different Language Models Learn Similar Number Representations Deqing Fu et.al. 2604.20817 null
2026-04-22 DNA storage approaching the information-theoretic ceiling James L. Banal et.al. 2604.20810 null
2026-04-22 OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Qiguang Chen et.al. 2604.20806 null
2026-04-22 Synthesizing Multi-Agent Harnesses for Vulnerability Discovery Hanzhi Liu et.al. 2604.20801 null
2026-04-22 LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Inclusion AI et.al. 2604.20796 null
2026-04-22 Automatic Ontology Construction Using LLMs as an External Layer of Memory, Verification, and Planning for Hybrid Intelligent Systems Pavel Salovskii et.al. 2604.20795 null
2026-04-22 Can “AI” Be a Doctor? A Study of Empathy, Readability, and Alignment in Clinical LLMs Mariano Barone et.al. 2604.20791 null
2026-04-22 V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Yubo Jiang et.al. 2604.20755 null
2026-04-22 Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation Dongding Lin et.al. 2604.20749 null
2026-04-22 RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering Marisa Hudspeth et.al. 2604.20738 null
2026-04-22 Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback Guotao Liang et.al. 2604.20730 null
2026-04-22 COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling Noah Flynn et.al. 2604.20720 null
2026-04-22 SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Jiahao Xie et.al. 2604.20705 null
2026-04-22 R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs Jiahao Xie et.al. 2604.20696 null
2026-04-22 MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment Andor Vári-Kakas et.al. 2604.20685 null
2026-04-21 PlayCoder: Making LLM-Generated GUI Code Playable Zhiyuan Peng et.al. 2604.19742 null
2026-04-21 Discovering a Shared Logical Subspace: Steering LLM Logical Reasoning via Alignment of Natural-Language and Symbolic Views Feihao Fang et.al. 2604.19716 null
2026-04-21 Epistemic orientation in parliamentary discourse is associated with deliberative democracy Segun Aroyehun et.al. 2604.19699 null
2026-04-21 Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks Jing Jin et.al. 2604.19697 null
2026-04-21 A-MAR: Agent-based Multimodal Art Retrieval for Fine-Grained Artwork Understanding Shuai Wang et.al. 2604.19689 null
2026-04-21 Exploring Language-Agnosticity in Function Vectors: A Case Study in Machine Translation Nurkhan Laiyk et.al. 2604.19678 null
2026-04-21 InHabit: Leveraging Image Foundation Models for Scalable 3D Human Placement Nikita Kister et.al. 2604.19673 null
2026-04-21 Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Yi Zhong et.al. 2604.19667 null
2026-04-21 Pause or Fabricate? Training Language Models for Grounded Reasoning Yiwen Qiu et.al. 2604.19656 null
2026-04-21 FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training Shuyao Qi et.al. 2604.19654 null
2026-04-21 A Gesture-Based Visual Learning Model for Acoustophoretic Interactions using a Swarm of AcoustoBots Alex Lin et.al. 2604.19643 null
2026-04-21 Micro Language Models Enable Instant Responses Wen Cheng et.al. 2604.19642 null
2026-04-21 SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models Josue Torres-Fonseca et.al. 2604.19638 null
2026-04-21 CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Xiangyang Luo et.al. 2604.19636 null
2026-04-21 Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model Shuhai Peng et.al. 2604.19635 null
2026-04-21 Time Series Augmented Generation for Financial Applications Anton Kolonin et.al. 2604.19633 null
2026-04-21 CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers Weidong Chen et.al. 2604.19632 null
2026-04-21 Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models Kihyuk Lee et.al. 2604.19598 null
2026-04-21 Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI Wenqing Wu et.al. 2604.19578 null
2026-04-21 Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic Chuou Xu et.al. 2604.19567 null
2026-04-21 Cross-Stock Predictability via LLM-Augmented Semantic Networks Yikuan Huang et.al. 2604.19476 null
2026-04-21 LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal Issues Fanyu Wang et.al. 2604.19464 null
2026-04-21 Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4 Alex Polyakov et.al. 2604.19461 null
2026-04-21 Unsupervised Confidence Calibration for Reasoning LLMs from a Single Generation Thomas Zollo et.al. 2604.19444 null
2026-04-21 What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search Xinhao Zhang et.al. 2604.19440 null
2026-04-21 Discerning Authorship in Online Health Communities: Experience, Trust, and Transparency Implications for Moderating AI Yefim Shulman et.al. 2604.19429 null
2026-04-21 MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding Zheng Lian et.al. 2604.19417 null
2026-04-21 VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing Yanbin Huang et.al. 2604.19412 null
2026-04-21 HP-Edit: A Human-Preference Post-Training Framework for Image Editing Fan Li et.al. 2604.19406 null
2026-04-21 Lost in Translation: Do LVLM Judges Generalize Across Languages? Md Tahmid Rahman Laskar et.al. 2604.19405 null
2026-04-21 CASCADE: Detecting Inconsistencies between Code and Documentation with Automatic Test Generation Tobias Kiecker et.al. 2604.19400 null
2026-04-21 GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models Ziyang Wang et.al. 2604.19398 null
2026-04-21 Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain? Niclas Doll et.al. 2604.19394 null
2026-04-21 Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval Zhiheng Fu et.al. 2604.19386 null
2026-04-21 Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture The Flag Challenges Ali Al-Kaswan et.al. 2604.19354 null
2026-04-21 DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing Jinyu Guo et.al. 2604.19351 null
2026-04-21 Quadruped Parkour Learning: Sparsely Gated Mixture of Experts with Visual Input Michael Ziegltrum et.al. 2604.19344 null
2026-04-21 Are Large Language Models Economically Viable for Industry Deployment? Abdullah Mohammad et.al. 2604.19342 null
2026-04-21 Evaluation-driven Scaling for Scientific Discovery Haotian Ye et.al. 2604.19341 null
2026-04-21 Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation Eoghan Cunningham et.al. 2604.19331 null
2026-04-20 Sessa: Selective State Space Attention Liubomyr Horbatko et.al. 2604.18580 null
2026-04-20 ReCap: Lightweight Referential Grounding for Coherent Story Visualization Aditya Arora et.al. 2604.18575 null
2026-04-20 When Can LLMs Learn to Reason with Weak Supervision? Salman Rahman et.al. 2604.18574 null
2026-04-20 Back into Plato’s Cave: Examining Cross-modal Representational Convergence at Scale A. Sophia Koepke et.al. 2604.18572 null
2026-04-20 Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering Manan Gupta et.al. 2604.18567 null
2026-04-20 Benchmarking System Dynamics AI Assistants: Cloud Versus Local LLMs on CLD Extraction and Discussion Terry Leitch et.al. 2604.18566 null
2026-04-20 Dual Alignment Between Language Model Layers and Human Sentence Processing Tatsuki Kuribayashi et.al. 2604.18563 null
2026-04-20 GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling Alireza Dadgarnia et.al. 2604.18556 null
2026-04-20 FUSE: Ensembling Verifiers with Zero Labeled Data Joonhyuk Lee et.al. 2604.18547 null
2026-04-20 OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Xinyu Ma et.al. 2604.18530 null
2026-04-20 S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models Nitish Shukla et.al. 2604.18512 null
2026-04-20 Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks Md Rysul Kabir et.al. 2604.18510 null
2026-04-20 MASS-RAG: Multi-Agent Synthesis Retrieval-Augmented Generation Xingchen Xiao et.al. 2604.18509 null
2026-04-20 Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints Hao Meng et.al. 2604.18489 null
2026-04-20 OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Jinghui Lu et.al. 2604.18486 null
2026-04-20 XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments Kangan Qian et.al. 2604.18484 null
2026-04-20 Multi-Scale Reversible Chaos Game Representation: A Unified Framework for Sequence Classification Sarwan Ali et.al. 2604.18477 null
2026-04-20 SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection Hao Vo et.al. 2604.18476 null
2026-04-20 Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts Jacob Morrison et.al. 2604.18473 null
2026-04-20 NI Sampling: Accelerating Discrete Diffusion Sampling by Token Order Optimization Enshu Liu et.al. 2604.18471 null
2026-04-19 VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech Yi-Cheng Lin et.al. 2604.17248 null
2026-04-19 All Public Voices Are Equal, But Are Some More Equal Than Others to LLMs? Sola Kim et.al. 2604.17247 null
2026-04-19 DORA Explorer: Improving the Exploration Ability of LLMs Without Training Priya Gurjar et.al. 2604.17244 null
2026-04-19 RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation Rui Min et.al. 2604.17243 null
2026-04-19 GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning Kun Wang et.al. 2604.17241 null
2026-04-19 From Language to Action: Enhancing LLM Task Efficiency with Task-Aware MCP Server Recommendation Shiyu He et.al. 2604.17234 null
2026-04-19 Revisiting Auxiliary Losses for Conditional Depth Routing: An Empirical Study Qingwei Lin et.al. 2604.17228 null
2026-04-19 Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models – A Research Agenda Minxian Xu et.al. 2604.17227 null
2026-04-19 A Multi-Agent Approach for Claim Verification from Tabular Data Documents Rudra Ranajee Saha et.al. 2604.17225 null
2026-04-19 Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation Jiuyun Jiang et.al. 2604.17220 null
2026-04-19 Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability Lijie Zhou et.al. 2604.17217 null
2026-04-19 Continual Safety Alignment via Gradient-Based Sample Selection Thong Bach et.al. 2604.17215 null
2026-04-19 Beyond the Basics: Leveraging Large Language Model for Fine-Grained Medical Entity Recognition Nwe Ni Win et.al. 2604.17214 null
2026-04-19 Guardrails in Logit Space: Safety Token Regularization for LLM Alignment Thong Bach et.al. 2604.17210 null
2026-04-19 DREAM: Dynamic Retinal Enhancement with Adaptive Multi-modal Fusion for Expert Precision Medical Report Generation Nagur Shareef Shaik et.al. 2604.17209 null
2026-04-19 Calibrating Model-Based Evaluation Metrics for Summarization Hongye Liu et.al. 2604.17200 null
2026-04-19 Do LLM-derived graph priors improve multi-agent coordination? Nikunj Gupta et.al. 2604.17191 null
2026-04-19 React-ing to Grace Hopper 200: Five Open-Weights Coding Models, One React Native App, One GH200, One Weekend Alex Potanin et.al. 2604.17187 null
2026-04-19 SynthFix: Adaptive Neuro-Symbolic Code Vulnerability Repair Yifan Zhang et.al. 2604.17184 null
2026-04-19 Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy Shun-ichiro Hayashi et.al. 2604.17182 null
2026-04-17 Using Large Language Models and Knowledge Graphs to Improve the Interpretability of Machine Learning Models in Manufacturing Thomas Bayer et.al. 2604.16280 null
2026-04-17 Evaluating the Progression of Large Language Model Capabilities for Small-Molecule Drug Design Shriram Chennakesavalu et.al. 2604.16279 null
2026-04-17 Learning to Reason with Insight for Informal Theorem Proving Yunhe Li et.al. 2604.16278 null
2026-04-17 No Universal Courtesy: A Cross-Linguistic, Multi-Model Study of Politeness Effects on LLMs Using the PLUM Corpus Hitesh Mehta et.al. 2604.16275 null
2026-04-17 VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects Xiangbo Gao et.al. 2604.16272 null
2026-04-17 From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text Van-Truong Le et.al. 2604.16270 null
2026-04-17 FL-MHSM: Spatially-adaptive Fusion and Ensemble Learning for Flood-Landslide Multi-Hazard Susceptibility Mapping at Regional Scale Aswathi Mundayatt et.al. 2604.16265 null
2026-04-17 Information Router for Mitigating Modality Dominance in Vision-Language Models Seulgi Kim et.al. 2604.16264 null
2026-04-17 Semantic Area Graph Reasoning for Multi-Robot Language-Guided Search Ruiyang Wang et.al. 2604.16263 null
2026-04-17 SwanNLP at SemEval-2026 Task 5: An LLM-based Framework for Plausibility Scoring in Narrative Word Sense Disambiguation Deshan Sumanathilaka et.al. 2604.16262 null
2026-04-17 Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap Yige Xu et.al. 2604.16256 null
2026-04-17 Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization Siddhant Bharadwaj et.al. 2604.16248 null
2026-04-17 Joint-Centric Dual Contrastive Alignment with Structure-Preserving and Information-Balanced Regularization Habibeh Naderi et.al. 2604.16247 null
2026-04-17 Detecting and Suppressing Reward Hacking with Gradient Fingerprints Songtao Wang et.al. 2604.16242 null
2026-04-17 BAGEL: Benchmarking Animal Knowledge Expertise in Language Models Jiacheng Shen et.al. 2604.16241 null
2026-04-17 Optimizing Korean-Centric LLMs via Token Pruning Hoyeol Kim et.al. 2604.16235 null
2026-04-17 Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations Yanli Wang et.al. 2604.16217 null
2026-04-17 ChemGraph-XANES: An Agentic Framework for XANES Simulation and Analysis Vitor F. Grizzi et.al. 2604.16205 null
2026-04-17 Bridging the Gap between User Intent and LLM: A Requirement Alignment Approach for Code Generation Jia Li et.al. 2604.16198 null
2026-04-17 Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation Yide Ran et.al. 2604.16197 null
2026-04-16 Generalization in LLM Problem Solving: The Case of the Shortest Path Yao Tong et.al. 2604.15306 null
2026-04-16 AnimationBench: Are Video Models Good at Character-Centric Animation? Leyi Wu et.al. 2604.15299 null
2026-04-16 Why Do Vision Language Models Struggle To Recognize Human Emotions? Madhav Agarwal et.al. 2604.15280 null
2026-04-16 Enhancing Large Language Models with Retrieval Augmented Generation for Software Testing and Inspection Automation Zoe Fingleton et.al. 2604.15270 null
2026-04-16 From Tokens to Steps: Verification-Aware Speculative Decoding for Efficient Multi-Step Reasoning Kiran Purohit et.al. 2604.15244 null
2026-04-16 RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography Mélanie Roschewitz et.al. 2604.15231 null
2026-04-16 UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos Joel Perca et.al. 2604.15225 null
2026-04-16 Context Over Content: Exposing Evaluation Faking in Automated Judges Manan Gupta et.al. 2604.15224 null
2026-04-16 MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events Raunak Agarwal et.al. 2604.15203 null
2026-04-16 VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models Huawei Ji et.al. 2604.15188 null
2026-04-16 Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines Marcel Wagenländer et.al. 2604.15186 null
2026-04-16 OmniLight: One Model to Rule All Lighting Conditions Youngjin Oh et.al. 2604.15170 null
2026-04-16 DPC: Training-Free Text-to-SQL Candidate Selection via Dual-Paradigm Consistency Boyan Li et.al. 2604.15163 null
2026-04-16 Compressing Sequences in the Latent Embedding Space: $K$ -Token Merging for Large Language Models Zihao Xu et.al. 2604.15153 null
2026-04-16 QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Alexey Khoroshilov et.al. 2604.15151 null
2026-04-16 IG-Search: Step-Level Information Gain Rewards for Search-Augmented Reasoning Zihan Liang et.al. 2604.15148 null
2026-04-16 Feedback-Driven Execution for LLM-Based Binary Analysis XiangRui Zhang et.al. 2604.15136 null
2026-04-16 Blinded Multi-Rater Comparative Evaluation of a Large Language Model and Clinician-Authored Responses in CGM-Informed Diabetes Counseling Zhijun Guo et.al. 2604.15124 null
2026-04-16 IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation Haozhi Fan et.al. 2604.15109 null
2026-04-16 OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis Kanzhi Cheng et.al. 2604.15093 null
2026-04-16 Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID Jiaxuan Li et.al. 2604.15090 null
2026-04-16 Autonomous Evolution of EDA Tools: Multi-Agent Self-Evolved ABC Cunxi Yu et.al. 2604.15082 null
2026-04-16 Atropos: Improving Cost-Benefit Trade-off of LLM-based Agents under Self-Consistency with Early Termination and Model Hotswap Naryeong Kim et.al. 2604.15075 null
2026-04-16 HintPilot: LLM-based Compiler Hint Synthesis for Code Optimization Hanyun Jiang et.al. 2604.15041 null
2026-04-16 Towards Faster Language Model Inference Using Mixture-of-Experts Flow Matching Aihua Li et.al. 2604.15009 null
2026-04-16 Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving Tingyang Sun et.al. 2604.14993 null
2026-04-16 Dr.~RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement Wenji Fang et.al. 2604.14989 null
2026-04-16 SAGER: Self-Evolving User Policy Skills for Recommendation Agent Zhen Tao et.al. 2604.14972 null
2026-04-16 Explain the Flag: Contextualizing Hate Speech Beyond Censorship Jason Liartis et.al. 2604.14970 null
2026-04-16 Discovering Novel LLM Experts via Task-Capability Coevolution Andrew Dai et.al. 2604.14969 null
2026-04-16 UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards Jun Wang et.al. 2604.14967 null
2026-04-15 One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Zheyu Zhang et.al. 2604.14149 null
2026-04-15 ROSE: Retrieval-Oriented Segmentation Enhancement Song Tang et.al. 2604.14147 null
2026-04-15 LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning Sumeet Ramesh Motwani et.al. 2604.14140 null
2026-04-15 Don’t Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models Ami Baid et.al. 2604.14129 null
2026-04-15 Rhetorical Questions in LLM Representations: A Linear Probing Study Louie Hong Yao et.al. 2604.14128 null
2026-04-15 HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System Tianshuo Yang et.al. 2604.14125 null
2026-04-15 Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis Zipeng Ling et.al. 2604.14121 null
2026-04-15 TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration Zerun Ma et.al. 2604.14116 null
2026-04-15 Interpretable Stylistic Variation in Human and LLM Writing Across Genres, Models, and Decoding Strategies Swati Rallapalli et.al. 2604.14111 null
2026-04-15 From Weights to Activations: Is Steering the Next Frontier of Adaptation? Simon Ostermann et.al. 2604.14090 null
2026-04-15 Training-Free Semantic Multi-Object Tracking with Vision-Language Models Laurence Bonat et.al. 2604.14074 null
2026-04-15 Towards Unconstrained Human-Object Interaction Francesco Tonini et.al. 2604.14069 null
2026-04-15 From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution Pavel Chizhov et.al. 2604.14053 null
2026-04-15 Enhancing Local Life Service Recommendation with Agentic Reasoning in Large Language Model Shiteng Cao et.al. 2604.14051 null
2026-04-15 Demanding peer review is associated with higher impact in published science Huihuang Jiang et.al. 2604.14047 null
2026-04-15 Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models Xiaohe Li et.al. 2604.14044 null
2026-04-15 Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios Xiaomin Li et.al. 2604.14041 null
2026-04-15 Large Language Models to Enhance Business Process Modeling: Past, Present, and Future Trends João Bettencourt et.al. 2604.14034 null
2026-04-15 Dual-Enhancement Product Bundling: Bridging Interactive Graph and Large Language Model Zhe Huang et.al. 2604.14030 null
2026-04-15 MAny: Merge Anything for Multimodal Continual Instruction Tuning Zijian Gao et.al. 2604.14016 null
2026-04-13 Psychological Concept Neurons: Can Neural Control Bias Probing and Shift Generation in LLMs? Yuto Harada et.al. 2604.11802 null
2026-04-13 CLSGen: A Dual-Head Fine-Tuning Framework for Joint Probabilistic Classification and Verbalized Explanation WonJin Yoon et.al. 2604.11801 null
2026-04-13 C-ReD: A Comprehensive Chinese Benchmark for AI-Generated Text Detection Derived from Real-World Prompts Chenxi Qing et.al. 2604.11796 null
2026-04-13 A Mechanistic Analysis of Looped Reasoning Language Models Hugh Blayney et.al. 2604.11791 null
2026-04-13 ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection Wei Zhao et.al. 2604.11790 null
2026-04-13 General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks Junlin Liu et.al. 2604.11778 null
2026-04-13 Towards Automated Pentesting with Large Language Models Ricardo Bessa et.al. 2604.11772 null
2026-04-13 Enhancing Program Repair with Specification Guidance and Intermediate Behavioral Signals Minh Le-Anh et.al. 2604.11770 null
2026-04-13 Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks Yoonsang Lee et.al. 2604.11753 null
2026-04-13 LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling Yuxin Chen et.al. 2604.11748 null
2026-04-13 Discourse Diversity in Multi-Turn Empathic Dialogue Hongli Zhan et.al. 2604.11742 null
2026-04-13 Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games Keyang Zhong et.al. 2604.11741 null
2026-04-13 Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions Manuela González-González et.al. 2604.11730 null
2026-04-13 Predicting User Satisfaction in Online Education Platforms: A Large Language Model Based Multi-Modal Review Mining Framework Arman Bekov et.al. 2604.11723 null
2026-04-13 SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context Shuquan Lian et.al. 2604.11716 null
2026-04-13 Agentic Driving Coach: Robustness and Determinism of Agentic AI-Powered Human-in-the-Loop Cyber-Physical Systems Deeksha Prahlad et.al. 2604.11705 null
2026-04-13 DreamKG: A KG-Augmented Conversational System for People Experiencing Homelessness Javad M Alizadeh et.al. 2604.11703 null
2026-04-13 Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning Jieying Xue et.al. 2604.11699 null
2026-04-13 EA-Agent: A Structured Multi-Step Reasoning Agent for Entity Alignment Yixuan Nan et.al. 2604.11686 null
2026-04-13 VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification Jiangyou Zhu et.al. 2604.11671 null
2026-04-12 LLMs Should Incorporate Explicit Mechanisms for Human Empathy Xiaoxing You et.al. 2604.10557 null
2026-04-12 Lost in Diffusion: Uncovering Hallucination Patterns and Failure Modes in Diffusion Large Language Models Zhengnan Guo et.al. 2604.10556 null
2026-04-12 Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training? Wanyi Chen et.al. 2604.10547 null
2026-04-12 Enhanced Self-Learning with Epistemologically-Informed LLM Dialogue Yi-Fan Cao et.al. 2604.10545 null
2026-04-12 WaveMoE: A Wavelet-Enhanced Mixture-of-Experts Foundation Model for Time Series Forecasting Shunyu Wu et.al. 2604.10544 null
2026-04-12 IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs Yuzhen Mao et.al. 2604.10539 null
2026-04-12 Evaluating Small Open LLMs for Medical Question Answering: A Practical Framework Avi-ad Avraam Buskila et.al. 2604.10535 null
2026-04-12 Machine Learning-Based Detection of MCP Attacks Tobias Mattsson et.al. 2604.10534 null
2026-04-12 Towards an Appropriate Level of Reliance on AI: A Preliminary Reliance-Control Framework for AI in Software Engineering Samuel Ferino et.al. 2604.10530 null
2026-04-12 BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Aaditya Baranwal et.al. 2604.10528 null
2026-04-12 ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization Suyoung Bae et.al. 2604.10520 null
2026-04-12 From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum Learning Xiaoda Yang et.al. 2604.10517 null
2026-04-12 Structure-Grounded Knowledge Retrieval via Code Dependencies for Multi-Step Data Reasoning Xinyi Huang et.al. 2604.10516 null
2026-04-12 Agent Mentor: Framing Agent Knowledge through Semantic Trajectory Analysis Roi Ben-Gigi et.al. 2604.10513 null
2026-04-12 Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation Yanjie He et.al. 2604.10511 null
2026-04-12 How Many Tries Does It Take? Iterative Self-Repair in LLM Code Generation Across Model Scales and Benchmarks Johin Johny Arimbur et.al. 2604.10508 null
2026-04-12 A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning Xiaoda Yang et.al. 2604.10506 null
2026-04-12 CARO: Chain-of-Analogy Reasoning Optimization for Robust Content Moderation Bingzhe Wu et.al. 2604.10504 null
2026-04-12 CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMs Haotian Lu et.al. 2604.10502 null
2026-04-12 MuSimA: A Tool with Multi-modal Input for Generating Bespoke ABAC Datasets Saket Jha et.al. 2604.10501 null
2026-04-10 Tango: Taming Visual Signals for Efficient Video Large Language Models Shukang Yin et.al. 2604.09547 null
2026-04-10 Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism Hadas Orgad et.al. 2604.09544 null
2026-04-10 EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks Lulin Liu et.al. 2604.09535 null
2026-04-10 Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise Zibin Geng et.al. 2604.09532 null
2026-04-10 VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images Guanyu Zhou et.al. 2604.09531 null
2026-04-10 VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning Wenyi Xiao et.al. 2604.09529 null
2026-04-10 When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation Ahmed Nusayer Ashik et.al. 2604.09515 null
2026-04-10 Many Ways to Be Fake: Benchmarking Fake News Detection Under Strategy-Driven AI Generation Xinyu Wang et.al. 2604.09514 null
2026-04-10 Integrated electro-optic attention nonlinearities for transformers Luis Mickeler et.al. 2604.09512 null
2026-04-10 RIRF: Reasoning Image Restoration Framework Wending Yan et.al. 2604.09511 null
2026-04-10 VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning Yucheng Shen et.al. 2604.09508 null
2026-04-10 Strategic Algorithmic Monoculture:Experimental Evidence from Coordination Games Gonzalo Ballestero et.al. 2604.09502 null
2026-04-10 You Can’t Fight in Here! This is BBS! Richard Futrell et.al. 2604.09501 null
2026-04-10 BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation Hippolyte Gisserot-Boukhlef et.al. 2604.09497 null
2026-04-10 RecaLLM: Addressing the Lost-in-Thought Phenomenon with Explicit In-Context Retrieval Kyle Whitecross et.al. 2604.09494 null
2026-04-10 Policy-Aware Edge LLM-RAG Framework for Internet of Battlefield Things Mission Orchestration Om Solanki et.al. 2604.09493 null
2026-04-10 Dynamic Ranked List Truncation for Reranking Pipelines via LLM-generated Reference-Documents Nilanjan Sinhababu et.al. 2604.09492 null
2026-04-10 Across the Levels of Analysis: Explaining Predictive Processing in Humans Requires More Than Machine-Estimated Probabilities Sathvik Nair et.al. 2604.09466 null
2026-04-10 Adaptor: Advancing Assistive Teleoperation with Few-Shot Learning and Cross-Operator Generalization Yu Liu et.al. 2604.09462 null
2026-04-10 From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models Chenchen Zhang et.al. 2604.09459 null
2026-04-09 Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts Haolei Xu et.al. 2604.08541 null
2026-04-09 AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Ziwei Zhou et.al. 2604.08540 null
2026-04-09 OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Wenbo Hu et.al. 2604.08539 null
2026-04-09 ParseBench: A Document Parsing Benchmark for AI Agents Boyang Zhang et.al. 2604.08538 null
2026-04-09 Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding Mu Nan et.al. 2604.08537 null
2026-04-09 Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models Feng Luo et.al. 2604.08527 null
2026-04-09 Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest Addison J. Wu et.al. 2604.08525 null
2026-04-09 What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal Stephen Cheng et.al. 2604.08524 null
2026-04-09 UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding Joungbin An et.al. 2604.08522 null
2026-04-09 Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts Jiayuan Ye et.al. 2604.08519 null
2026-04-09 What do Language Models Learn and When? The Implicit Curriculum Hypothesis Emmy Liu et.al. 2604.08510 null
2026-04-09 What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric Mohamed Amine Kerkouri et.al. 2604.08494 null
2026-04-09 Figures as Interfaces: Toward LLM-Native Artifacts for Scientific Discovery Yifang Wang et.al. 2604.08491 null
2026-04-09 AI generates well-liked but templatic empathic responses Emma Gueorguieva et.al. 2604.08479 null
2026-04-09 SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions Ashima Suvarna et.al. 2604.08477 null
2026-04-09 Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization Sai Srinivas Kancheti et.al. 2604.08476 null
2026-04-09 LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation Jingjing Wang et.al. 2604.08475 null
2026-04-09 From Safety Risk to Design Principle: Peer-Preservation in Multi-Agent LLM Systems and Its Implications for Orchestrated Democratic Discourse Analysis Juergen Dietrich et.al. 2604.08465 null
2026-04-09 CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning Rui Gan et.al. 2604.08457 null
2026-04-09 Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models Marcel Gröpl et.al. 2604.08456 null
2026-04-09 Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification Xun Zhu et.al. 2604.08333 null
2026-04-09 ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment via Explicit Injection He Geng et.al. 2604.08326 null
2026-04-09 Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data Yuchuan Deng et.al. 2604.08322 null
2026-04-09 A Model Context Protocol Server for Quantum Execution in Hybrid Quantum-HPC Environments Masaki Shiraishi et.al. 2604.08318 null
2026-04-09 Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Directions Yuming Xu et.al. 2604.08304 null
2026-04-09 DMax: Aggressive Parallel Decoding for dLLMs Zigeng Chen et.al. 2604.08302 null
2026-04-09 SeLaR: Selective Latent Reasoning in Large Language Models Renyu Fu et.al. 2604.08299 null
2026-04-09 Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models Weiwei Qi et.al. 2604.08297 null
2026-04-09 Can Vision Language Models Judge Action Quality? An Empirical Evaluation Miguel Monte e Freitas et.al. 2604.08294 null
2026-04-09 CIAO - Code In Architecture Out - Automated Software Architecture Documentation with Large Language Models Marco De Luca et.al. 2604.08293 null
2026-04-09 Tokalator: A Context Engineering Toolkit for Artificial Intelligence Coding Assistants Vahid Farajijobehdar et.al. 2604.08290 null
2026-04-09 Distributed Multi-Layer Editing for Rule-Level Knowledge in Large Language Models Yating Wang et.al. 2604.08284 null
2026-04-09 When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning Ruotao Xu et.al. 2604.08281 null
2026-04-09 Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models Jing Gu et.al. 2604.08266 null
2026-04-09 Neural-Symbolic Knowledge Tracing: Injecting Educational Knowledge into Deep Learning for Responsible Learner Modelling Danial Hooshyar et.al. 2604.08263 null
2026-04-09 Behavior-Aware Item Modeling via Dynamic Procedural Solution Representations for Knowledge Tracing Jun Seo et.al. 2604.08260 null
2026-04-09 Self-Debias: Self-correcting for Debiasing Large Language Models Xuan Feng et.al. 2604.08243 null
2026-04-09 Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering Chenyu Zhou et.al. 2604.08224 null
2026-04-09 MemCoT: Test-Time Scaling through Memory-Driven Chain-of-Thought Haodong Lei et.al. 2604.08216 null
2026-04-09 EditCaption: Human-Aligned Instruction Synthesis for Image Editing via Supervised Fine-Tuning and Direct Preference Optimization Xiangyuan Wang et.al. 2604.08213 null
2026-04-08 Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization Qiyao Ma et.al. 2604.07343 null
2026-04-08 Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images Yuechen Jiang et.al. 2604.07338 null
2026-04-08 Syntax Is Easy, Semantics Is Hard: Evaluating LLMs for LTL Translation Priscilla Kyei Danso et.al. 2604.07321 null
2026-04-08 Evaluating In-Context Translation with Synchronous Context-Free Grammar Transduction Jackson Petty et.al. 2604.07320 null
2026-04-08 Chatbot-Based Assessment of Code Understanding in Automated Programming Assessment Systems Eduard Frankford et.al. 2604.07304 null
2026-04-08 Region-Graph Optimal Transport Routing for Mixture-of-Experts Whole-Slide Image Classification Xin Tian et.al. 2604.07298 null
2026-04-08 Why teaching resists automation in an AI-inundated era: Human judgment, non-modular work, and the limits of delegation Songhee Han et.al. 2604.07285 null
2026-04-08 A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering Nusrat Sultana et.al. 2604.07274 null
2026-04-08 On the Price of Privacy for Language Identification and Generation Xiaoyu Li et.al. 2604.07238 null
2026-04-08 How Much LLM Does a Self-Revising Agent Actually Need? Seongwoo Jeong et.al. 2604.07236 null
2026-04-08 TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories Yen-Shan Chen et.al. 2604.07223 null
2026-04-08 VersaVogue: Visual Expert Orchestration and Preference Alignment for Unified Fashion Synthesis Jian Yu et.al. 2604.07210 null
2026-04-08 Probing 3D Chromatin Structure Awareness in Evo2 DNA Language Model UkJin Lee et.al. 2604.07196 null
2026-04-08 LaScA: Language-Conditioned Scalable Modelling of Affective Dynamics Kosmas Pinitas et.al. 2604.07193 null
2026-04-08 The ATOM Report: Measuring the Open Language Model Ecosystem Nathan Lambert et.al. 2604.07190 null
2026-04-08 Agent-Driven Corpus Linguistics: A Framework for Autonomous Linguistic Discovery Jia Yu et.al. 2604.07189 null
2026-04-08 InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models Hongyu Chen et.al. 2604.07173 null
2026-04-08 Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling Tom A. Lamb et.al. 2604.07172 null
2026-04-08 Critical Inker: Scaffolding Critical Thinking in AI-Assisted Writing Through Socratic Questioning Philipp Hugenroth et.al. 2604.07167 null
2026-04-08 Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization Yu Li et.al. 2604.07165 null
2026-04-08 Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing Ning Yang et.al. 2604.07148 null
2026-04-08 Dynamic Context Evolution for Scalable Synthetic Data Generation Ryan Lingo et.al. 2604.07147 null
2026-04-08 Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering Zhuohong Chen et.al. 2604.07146 null
2026-04-08 Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics Youhe Jiang et.al. 2604.07144 null
2026-04-08 A Utility-preserving De-identification Pipeline for Cross-hospital Radiology Data Sharing Chenhao Liu et.al. 2604.07128 null
2026-04-08 Language Bias under Conflicting Information in Multilingual LLMs Robert Östling et.al. 2604.07123 null
2026-04-08 The Impact of Steering Large Language Models with Persona Vectors in Educational Applications Yongchao Wu et.al. 2604.07102 null
2026-04-08 Selective Neuron Amplification for Training-Free Task Enhancement Ryyan Akhtar et.al. 2604.07098 null
2026-04-07 Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework Komal Kumar et.al. 2604.06170 null
2026-04-07 In-Place Test-Time Training Guhao Feng et.al. 2604.06169 null
2026-04-07 HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models Reihaneh Zohrabi et.al. 2604.06165 null
2026-04-07 MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control Yuchi Wang et.al. 2604.06156 null
2026-04-07 Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement Qimin Zhong et.al. 2604.06155 null
2026-04-07 Exclusive Unlearning Mutsumi Sasaki et.al. 2604.06154 null
2026-04-07 Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents Bowen Ye et.al. 2604.06132 null
2026-04-07 Gym-Anything: Turn any Software into an Agent Environment Pranjal Aggarwal et.al. 2604.06126 null
2026-04-07 Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery Hao Chen et.al. 2604.06124 null
2026-04-07 LLM4CodeRE: Generative AI for Code Decompilation Analysis and Reverse Engineering Hamed Jelodar et.al. 2604.06095 null
2026-04-07 Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives Changgeon Ko et.al. 2604.06091 null
2026-04-07 LAG-XAI: A Lie-Inspired Affine Geometric Framework for Interpretable Paraphrasing in Transformer Latent Spaces Olexander Mazurets et.al. 2604.06086 null
2026-04-07 Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning Juekai Lin et.al. 2604.06079 null
2026-04-07 Stories of Your Life as Others: A Round-Trip Evaluation of LLM-Generated Life Stories Conditioned on Rich Psychometric Profiles Ben Wigler et.al. 2604.06071 null
2026-04-07 Short Data, Long Context: Distilling Positional Knowledge in Transformers Patrick Huber et.al. 2604.06070 null
2026-04-07 From Hallucination to Structure Snowballing: The Alignment Tax of Constrained Decoding in LLM Reflection Hongxu Zhou et.al. 2604.06066 null
2026-04-07 Large Language Model Assisted Discovery of Optimal Dopants for Enhanced Thermoelectric Performance in CoSb $_3$ Based Skutterudites Yagnik Bandyopadhyay et.al. 2604.06048 null
2026-04-07 CoStream: Codec-Guided Resource-Efficient System for Video Streaming Analytics Yulin Zou et.al. 2604.06036 null
2026-04-07 A Multi-Stage Validation Framework for Trustworthy Large-scale Clinical Information Extraction using Large Language Models Maria Mahbub et.al. 2604.06028 null
2026-04-07 CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments Gustav Keppler et.al. 2604.06019 null
2026-04-07 FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks Michael Krumdick et.al. 2604.05912 null
2026-04-07 AICA-Bench: Holistically Examining the Capabilities of VLMs in Affective Image Content Analysis Dong She et.al. 2604.05900 null
2026-04-07 FRENCH-YMCA: A FRENCH Corpus meeting the language needs of Youth, froM Children to Adolescents Cherifa Ben Khelil et.al. 2604.05899 null
2026-04-07 HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference Bowen Zeng et.al. 2604.05887 null
2026-04-07 Mechanistic Circuit-Based Knowledge Editing in Large Language Models Tianyi Zhao et.al. 2604.05876 null
2026-04-07 Joint Knowledge Base Completion and Question Answering by Combining Large Language Models and Small Language Models Yinan Liu et.al. 2604.05875 null
2026-04-07 Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts Fatih Uenal et.al. 2604.05872 null
2026-04-07 JTON: A Token-Efficient JSON Superset with Zen Grid Tabular Encoding for Large Language Models Gowthamkumar Nandakishore et.al. 2604.05865 null
2026-04-07 LoRM: Learning the Language of Rotating Machinery for Self-Supervised Condition Monitoring Xiao Qin et.al. 2604.05863 null
2026-04-07 When Do We Need LLMs? A Diagnostic for Language-Driven Bandits Uljad Berdica et.al. 2604.05859 null
2026-04-07 Deep Researcher Agent: An Autonomous Framework for 24/7 Deep Learning Experimentation with Zero-Cost Monitoring Xiangyue Zhang et.al. 2604.05854 null
2026-04-07 Reading Between the Pixels: An Inscriptive Jailbreak Attack on Text-to-Image Models Zonghao Ying et.al. 2604.05853 null
2026-04-07 AgentGL: Towards Agentic Graph Learning with LLMs via Reinforcement Learning Yuanfu Sun et.al. 2604.05846 null
2026-04-07 Vision-Guided Iterative Refinement for Frontend Code Generation Hannah Sansford et.al. 2604.05839 null
2026-04-07 WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering Yingjian Zhu et.al. 2604.05818 null
2026-04-07 Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents Shuai Zhen et.al. 2604.05808 null
2026-04-07 Measuring What Matters!! Assessing Therapeutic Principles in Mental-Health Conversation Abdullah Mazhar et.al. 2604.05795 null
2026-04-07 An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face Yujian Liu et.al. 2604.05782 null
2026-04-07 What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say “I Don’t Know” Joosung Lee et.al. 2604.05779 null
2026-04-07 PhageBench: Can LLMs Understand Raw Bacteriophage Genomes? Yusen Hou et.al. 2604.05775 null
2026-04-05 Schema-Aware Planning and Hybrid Knowledge Toolset for Reliable Knowledge Graph Triple Verification Xinyan Ma et.al. 2604.04190 null
2026-04-05 AURA: Always-On Understanding and Real-Time Assistance via Video Streams Xudong Lu et.al. 2604.04184 null
2026-04-05 Scale-Aware Vision-Language Adaptation for Extreme Far-Distance Video Person Re-identification Ashwat Rajbhandari et.al. 2604.04183 null
2026-04-05 Comparative reversal learning reveals rigid adaptation in LLMs under non-stationary uncertainty Haomiaomiao Wang et.al. 2604.04182 null
2026-04-05 Position: Logical Soundness is not a Reliable Criterion for Neurosymbolic Fact-Checking with LLMs Jason Chan et.al. 2604.04177 null
2026-04-05 CoALFake: Collaborative Active Learning with Human-LLM Co-Annotation for Cross-Domain Fake News Detection Esma Aïmeur et.al. 2604.04174 null
2026-04-05 GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Yaohan Guan et.al. 2604.04172 null
2026-04-05 A Semi-Automated Annotation Workflow for Paediatric Histopathology Reports Using Small Language Models Avish Vijayaraghavan et.al. 2604.04168 null
2026-04-05 Readable Minds: Emergent Theory-of-Mind-Like Behavior in LLM Poker Agents Hsieh-Ting Lin et.al. 2604.04157 null
2026-04-05 Solar-VLM: Multimodal Vision-Language Models for Augmented Solar Power Forecasting Hang Fan et.al. 2604.04145 null
2026-04-05 Many Preferences, Few Policies: Towards Scalable Language Model Personalization Cheol Woo Kum et.al. 2604.04144 null
2026-04-05 Formalized Information Needs Improve Large-Language-Model Relevance Judgments Jüri Keller et.al. 2604.04140 null
2026-04-05 Learning Robust Visual Features in Computed Tomography Enables Efficient Transfer Learning for Clinical Tasks Rubén Moreno-Aguado et.al. 2604.04133 null
2026-04-05 Profile-Then-Reason: Bounded Semantic Complexity for Tool-Augmented Language Agents Paulo Akira F. Enabe et.al. 2604.04131 null
2026-04-05 SARES-DEIM: Sparse Mixture-of-Experts Meets DETR for Robust SAR Ship Detection Fenghao Song et.al. 2604.04127 null
2026-04-05 Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression Lingjie Zeng et.al. 2604.04120 null
2026-04-05 Hypothesis Graph Refinement: Hypothesis-Driven Exploration with Cascade Error Correction for Embodied Navigation Peixin Chen et.al. 2604.04108 null
2026-04-05 InsTraj: Instructing Diffusion Models with Travel Intentions to Generate Real-world Trajectories Yuanshao Zhu et.al. 2604.04106 null
2026-04-05 Compliance-by-Construction Argument Graphs: Using Generative AI to Produce Evidence-Linked Formal Arguments for Certification-Grade Accountability Mahyar T. Moghaddam et.al. 2604.04103 null
2026-04-05 BadgeX: IoT-Enhanced Wearable Analytics Meets LLMs for Collaborative Learning Zaibei Li et.al. 2604.04093 null
2026-04-03 CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning Ankan Deria et.al. 2604.03231 null
2026-04-03 BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence Sean Wu et.al. 2604.03216 null
2026-04-03 Prosocial Persuasion at Scale? Large Language Models Outperform Humans in Donation Appeals Across Levels of Personalization John Caffier et.al. 2604.03202 null
2026-04-03 Learning the Signature of Memorization in Autoregressive Language Models David Ilić et.al. 2604.03199 null
2026-04-03 The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling Takuya Shiba et.al. 2604.03191 null
2026-04-03 Reflective Context Learning: Studying the Optimization Primitives of Context Space Nikita Vassilyev et.al. 2604.03189 null
2026-04-03 Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models Gengwei Zhang et.al. 2604.03179 null
2026-04-03 Beyond the Parameters: A Technical Survey of Contextual Enrichment in Large Language Models: From In-Context Prompting to Causal Retrieval-Augmented Generation Prakhar Bansal et.al. 2604.03174 null
2026-04-03 Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents Delip Rao et.al. 2604.03173 null
2026-04-03 EffiMiniVLM: A Compact Dual-Encoder Regression Framework Yin-Loon Khor et.al. 2604.03172 null
2026-04-03 BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation Delip Rao et.al. 2604.03159 null
2026-04-03 Chart-RL: Policy Optimization Reinforcement Learning for Enhanced Visual Reasoning in Chart Question Answering with Vision Language Models Yunfei Bai et.al. 2604.03157 null
2026-04-03 Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control Lihao Sun et.al. 2604.03147 null
2026-04-03 InCoder-32B-Thinking: Industrial Code World Model for Thinking Jian Yang et.al. 2604.03144 null
2026-04-03 Beyond Precision: Importance-Aware Recall for Factuality Evaluation in Long-Form LLM Generation Nazanin Jafari et.al. 2604.03141 null
2026-04-03 FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation Mingao Tan et.al. 2604.03139 null
2026-04-03 A Systematic Security Evaluation of OpenClaw and Its Variants Yuhang Wang et.al. 2604.03131 null
2026-04-03 Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patches for Infrared Vision-Language Models Chengyin Hu et.al. 2604.03117 null
2026-04-03 Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning Zhangyun Tan et.al. 2604.03114 null
2026-04-03 PAFT: Preservation Aware Fine-Tuning for Minimal-Edit Program Repair Boyang Yang et.al. 2604.03113 null
2026-04-02 Steerable Visual Representations Jona Ruthardt et.al. 2604.02327 null
2026-04-02 Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation Daiwei Chen et.al. 2604.02324 null
2026-04-02 Batched Contextual Reinforcement: A Task-Scaling Law for Efficient Reasoning Bangji Yang et.al. 2604.02322 null
2026-04-02 Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining Junxuan Li et.al. 2604.02320 null
2026-04-02 Beyond the Assistant Turn: User Turn Generation as a Probe of Interaction Awareness in Language Models Sarath Shekkizhar et.al. 2604.02315 null
2026-04-02 go- $m$ HC: Direct Parameterization of Manifold-Constrained Hyper-Connections via Generalized Orthostochastic Matrices Torque Dandachi et.al. 2604.02309 null
2026-04-02 VOID: Video Object and Interaction Deletion Saman Motamed et.al. 2604.02296 null
2026-04-02 Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation Chongjie Ye et.al. 2604.02289 null
2026-04-02 Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing Gengsheng Li et.al. 2604.02288 null
2026-04-02 Retrieval-Augmented Question Answering over Scientific Literature for the Electron-Ion Collider Tina. J. Jat et.al. 2604.02259 null
2026-04-02 SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation Naomi Kombol et.al. 2604.02252 null
2026-04-02 Do Emotions in Prompts Matter? Effects of Emotional Framing on Large Language Models Minda Zhao et.al. 2604.02236 null
2026-04-02 Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs Abinitha Gourabathina et.al. 2604.02230 null
2026-04-02 When to ASK: Uncertainty-Gated Language Assistance for Reinforcement Learning Juarez Monteiro et.al. 2604.02226 null
2026-04-02 Impact of Multimodal and Conversational AI on Learning Outcomes and Experience Karan Taneja et.al. 2604.02221 null
2026-04-02 VISTA: Visualization of Token Attribution via Efficient Analysis Syed Ahmed et.al. 2604.02217 null
2026-04-02 Multi-Agent Video Recommenders: Evolution, Patterns, and Open Challenges Srivaths Ranganathan et.al. 2604.02211 null
2026-04-02 Towards Position-Robust Talent Recommendation via Large Language Models Silin Du et.al. 2604.02200 null
2026-04-02 Neuro-RIT: Neuron-Guided Instruction Tuning for Robust Retrieval-Augmented Language Model Jaemin Kim et.al. 2604.02194 null
2026-04-02 UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving Yongkang Li et.al. 2604.02190 null
2026-04-02 The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level Jeremy Herbst et.al. 2604.02178 null
2026-04-02 Adam’s Law: Textual Frequency Law on Large Language Models Hongyuan Adam Lu et.al. 2604.02176 null
2026-04-02 Quantifying Self-Preservation Bias in Large Language Models Matteo Migliarini et.al. 2604.02174 null
2026-04-02 Brief Is Better: Non-Monotonic Chain-of-Thought Budget Effects in Function-Calling Language Agents Xuan Qi et.al. 2604.02155 null
2026-04-02 TRACE-Bot: Detecting Emerging LLM-Driven Social Bots via Implicit Semantic Representations and AIGC-Enhanced Behavioral Patterns Zhongbo Wang et.al. 2604.02147 null
2026-04-02 MTI: A Behavior-Based Temperament Profiling System for AI Agents Jihoon Jeong et.al. 2604.02145 null
2026-04-02 GaelEval: Benchmarking LLM Performance for Scottish Gaelic Peter Devine et.al. 2604.02135 null
2026-04-02 Semantic Evolution over Populations for LLM-Guided Automated Program Repair Cuong Chi Le et.al. 2604.02134 null
2026-04-02 AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression Atul Kumar Sinha et.al. 2604.02119 null
2026-04-02 LLM-as-a-Judge for Time Series Explanations Preetham Sivalingam et.al. 2604.02118 null
2026-04-02 Reliable Control-Point Selection for Steering Reasoning in Large Language Models Haomin Zhuang et.al. 2604.02113 null
2026-04-02 FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators Chi Zhang et.al. 2604.02110 null
2026-04-02 GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding Rong Fan et.al. 2604.02093 null
2026-04-02 PLUME: Latent Reasoning Based Universal Multimodal Embedding Chenwei He et.al. 2604.02073 null
2026-04-02 Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection Soo Won Seo et.al. 2604.02071 null
2026-04-02 Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models Issa Sugiura et.al. 2604.02048 null
2026-04-01 HippoCamp: Benchmarking Contextual Agents on Personal Computers Zhe Yang et.al. 2604.01221 null
2026-04-01 Universal YOCO for Efficient Depth Scaling Yutao Sun et.al. 2604.01220 null
2026-04-01 LLM REgression with a Latent Iterative State Head Yiheng Su et.al. 2604.01206 null
2026-04-01 Therefore I am. I Think Esakkivel Esakkiraja et.al. 2604.01202 null
2026-04-01 ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget Nandan Thakur et.al. 2604.01195 null
2026-04-01 AgentWatcher: A Rule-based Prompt Injection Monitor Yanting Wang et.al. 2604.01194 null
2026-04-01 Embarrassingly Simple Self-Distillation Improves Code Generation Ruixiang Zhang et.al. 2604.01193 null
2026-04-01 True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies Graziano Blasilli et.al. 2604.01181 null
2026-04-01 A ROS 2 Wrapper for Florence-2: Multi-Mode Local Vision-Language Inference for Robotic Systems J. E. Domínguez-Vidal et.al. 2604.01179 null
2026-04-01 Screening Is Enough Ken M. Nakanishi et.al. 2604.01178 null
2026-04-01 Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning Cai Zhou et.al. 2604.01170 null
2026-04-01 S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models Jack Young et.al. 2604.01168 null
2026-04-01 AdaLoRA-QAT: Adaptive Low-Rank and Quantization-Aware Segmentation Prantik Deb et.al. 2604.01167 null
2026-04-01 Reasoning Shift: How Context Silently Shortens LLM Reasoning Gleb Rodionov et.al. 2604.01161 null
2026-04-01 FineLAP: Taming Heterogeneous Supervision for Fine-grained Language-Audio Pretraining Xiquan Li et.al. 2604.01155 null
2026-04-01 Brainstacks: Cross-Domain Cognitive Capabilities via Frozen MoE-LoRA Stacks for Continual LLM Learning Mohammad R. Abu Ayyash et.al. 2604.01152 null
2026-04-01 SERSEM: Selective Entropy-Weighted Scoring for Membership Inference in Code Language Models Kıvanç Kuzey Dikici et.al. 2604.01147 null
2026-04-01 Multi-Agent LLM Governance for Safe Two-Timescale Reinforcement Learning in SDN-IoT Defense Saeid Jamshidi et.al. 2604.01127 null
2026-04-01 Lightweight Prompt-Guided CLIP Adaptation for Monocular Depth Estimation Reyhaneh Ahani Manghotay et.al. 2604.01118 null
2026-04-01 CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance Haochen Liu et.al. 2604.01113 null
2026-04-01 Uncertainty-Aware Variational Reward Factorization via Probabilistic Preference Bases for LLM Personalization Gyuseok Lee et.al. 2604.00997 null
2026-04-01 ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration Bei Yan et.al. 2604.00983 null
2026-04-01 VisG AV-HuBERT: Viseme-Guided AV-HuBERT Aristeidis Papadopoulos et.al. 2604.00982 null
2026-04-01 Dual Optimal: Make Your LLM Peer-like with Dignity Xiangqi Wang et.al. 2604.00979 null
2026-04-01 FlexAI: A Multi-modal Solution for Delivering Personalized and Adaptive Fitness Interventions Shivangi Agarwal et.al. 2604.00968 null
2026-04-01 Understanding Transformers and Attention Mechanisms: An Introduction for Applied Mathematicians Michel Fabrice Serret et.al. 2604.00965 null
2026-04-01 Phase transition on a context-sensitive random language model with short range interactions Yuma Toji et.al. 2604.00947 null
2026-04-01 Auditing the Reliability of Multimodal Generative Search Erfan Samieyan Sahneh et.al. 2604.00944 null
2026-04-01 Positional Cognitive Specialization: Where Do LLMs Learn To Comprehend and Speak Your Language? Luis Frentzen Salim et.al. 2604.00923 null
2026-04-01 GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training Jesse van Oort et.al. 2604.00920 null
2026-04-01 Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time Razvan Mihai Popescu et.al. 2604.00917 null
2026-04-01 Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment Zhuchenyang Liu et.al. 2604.00913 null
2026-04-01 ProCap: Projection-Aware Captioning for Spatial Augmented Reality Zimo Cao et.al. 2604.00912 null
2026-04-01 JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation Issa Sugiura et.al. 2604.00909 null
2026-04-01 Beyond Symbolic Solving: Multi Chain-of-Thought Voting for Geometric Reasoning in Large Language Models Md. Abu Bakor Siddique et.al. 2604.00890 null
2026-04-01 PixelPrune: Pixel-Level Adaptive Visual Token Reduction via Predictive Coding Nan Wang et.al. 2604.00886 null
2026-04-01 KUET at StanceNakba Shared Task: StanceMoE: Mixture-of-Experts Architecture for Stance Detection Abdullah Al Shafi et.al. 2604.00878 null
2026-04-01 A 4D Representation for Training-Free Agentic Reasoning from Monocular Laparoscopic Video Maximilian Fehrentz et.al. 2604.00867 null
2026-04-01 Policy Improvement Reinforcement Learning Huaiyang Wang et.al. 2604.00860 null
2026-04-01 Reliability of Large Language Models for Design Synthesis: An Empirical Study of Variance, Prompt Sensitivity, and Method Scaffolding Rabia Iftikhar et.al. 2604.00851 null
2026-03-31 Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought? Max Kaufmann et.al. 2603.30036 null
2026-03-31 Reward-Based Online LLM Routing via NeuralUCB Ming-Hua Tsai et.al. 2603.30035 null
2026-03-31 The Triadic Cognitive Architecture: Bounding Autonomous Action via Spatio-Temporal and Epistemic Friction Davide Di Gioia et.al. 2603.30031 null
2026-03-31 Can Commercial LLMs Be Parliamentary Political Companions? Comparing LLM Reasoning Against Romanian Legislative Expuneri de Motive Iulian Lucău et.al. 2603.30028 null
2026-03-31 ContextClaim: A Context-Driven Paradigm for Verifiable Claim Detection Yufeng Li et.al. 2603.30025 null
2026-03-31 Hybrid Framework for Robotic Manipulation: Integrating Reinforcement Learning and Large Language Models Md Saad et.al. 2603.30022 null
2026-03-31 Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks Chong Xiang et.al. 2603.30016 null
2026-03-31 Performative Scenario Optimization Quanyan Zhu et.al. 2603.29982 null
2026-03-31 Scaling Video Pretraining for Surgical Foundation Models Sicheng Lu et.al. 2603.29966 null
2026-03-31 Think Anywhere in Code Generation Xue Jiang et.al. 2603.29957 null
2026-03-31 EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos Fumihiko Tsuchiya et.al. 2603.29943 null
2026-03-31 Bethe Ansatz with a Large Language Model Balázs Pozsgay et.al. 2603.29932 null
2026-03-31 SISA: A Scale-In Systolic Array for GEMM Acceleration Luigi Altamura et.al. 2603.29913 null
2026-03-31 C-TRAIL: A Commonsense World Framework for Trajectory Planning in Autonomous Driving Zhihong Cui et.al. 2603.29908 null
2026-03-31 ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation Yinuo Liu et.al. 2603.29902 null
2026-03-31 ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training Rui Ai et.al. 2603.29871 null
2026-03-31 SNEAK: Evaluating Strategic Communication and Information Leakage in Large Language Models Adar Avsian et.al. 2603.29846 null
2026-03-31 Cold-Starts in Generative Recommendation: A Reproducibility Study Zhen Zhang et.al. 2603.29845 null
2026-03-31 DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA Yi Chen et.al. 2603.29844 null
2026-03-31 Compiling Code LLMs into Lightweight Executables Jieke Shi et.al. 2603.29813 null
2026-03-29 EvA: An Evidence-First Audio Understanding Paradigm for LALMs Xinyuan Xie et.al. 2603.27667 null
2026-03-29 Investigating the Influence of Language on Sycophantic Behavior of Multilingual LLMs Bayan Abdullah Aldahlawi et.al. 2603.27664 null
2026-03-29 A Benchmarking Methodology to Assess Open-Source Video Large Language Models in Automatic Captioning of News Videos David Miranda Paredes et.al. 2603.27662 null
2026-03-29 V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models Xinying Lin et.al. 2603.27650 null
2026-03-29 PRBench: End-to-end Paper Reproduction in Physics Research Shi Qiu et.al. 2603.27646 null
2026-03-29 OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery Qi Guo et.al. 2603.27645 null
2026-03-29 Umwelt Engineering: Designing the Cognitive Worlds of Linguistic Agents Rodney Jehu-Appiah et.al. 2603.27626 null
2026-03-29 Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling Songchen Ma et.al. 2603.27624 null
2026-03-29 STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding Junho Kim et.al. 2603.27593 null
2026-03-29 Sci-Mind: Cognitively-Inspired Adversarial Debate for Autonomous Mathematical Modeling Ruiying Sun et.al. 2603.27584 null
2026-03-29 LLM-Enabled Low-Altitude UAV Natural Language Navigation via Signal Temporal Logic Specification Translation and Repair Yuqi Ping et.al. 2603.27583 null
2026-03-29 Structured Observation Language for Efficient and Generalizable Vision-Language Navigation Daojie Peng et.al. 2603.27577 null
2026-03-29 RAGent: Physics-Aware Agentic Reasoning for Training-Free mmWave Human Activity Recognition Mingda Han et.al. 2603.27571 null
2026-03-29 Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language Models Baoheng Zhang et.al. 2603.27558 null
2026-03-29 Toward Reliable Evaluation of LLM-Based Financial Multi-Agent Systems: Taxonomy, Coordination Primacy, and Cost Awareness Phat Nguyen et.al. 2603.27539 null
2026-03-29 LongCat-Next: Lexicalizing Modalities as Discrete Tokens Meituan LongCat Team et.al. 2603.27538 null
2026-03-29 Visualization of Machine Learning Models through Their Spatial and Temporal Listeners Siyu Wu et.al. 2603.27527 null
2026-03-29 Q-BIOLAT: Binary Latent Protein Fitness Landscapes for QUBO-Based Optimization Truong-Son Hy et.al. 2603.27526 null
2026-03-29 Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models Duanyi Yao et.al. 2603.27522 null
2026-03-29 Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs Utsav Maskey et.al. 2603.27518 null
2026-03-26 SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding Jiwook Han et.al. 2603.25733 null
2026-03-26 Seeing to Ground: Visual Attention for Hallucination-Resilient MDLLMs Vishal Narnaware et.al. 2603.25711 null
2026-03-26 S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation Ligong Han et.al. 2603.25702 null
2026-03-26 Self-Improvement of Large Language Models: A Technical Overview and Future Outlook Haoyan Yang et.al. 2603.25681 null
2026-03-26 Measuring What Matters – or What’s Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors Cole Walsh et.al. 2603.25674 null
2026-03-26 A Mentalistic Interface for Probing Folk-Psychological Attribution to Non-Humanoid Robots Giulio Pisaneschi et.al. 2603.25646 null
2026-03-26 Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos Abdullah Hamdi et.al. 2603.25645 null
2026-03-26 RenoBench: A Citation Parsing Benchmark Parth Sarin et.al. 2603.25640 null
2026-03-26 Beyond Via: Analysis and Estimation of the Impact of Large Language Models in Academic Papers Mingmeng Geng et.al. 2603.25638 null
2026-03-26 Is Mathematical Problem-Solving Expertise in Large Language Models Associated with Assessment Performance? Liang Zhang et.al. 2603.25633 null
2026-03-26 PICon: A Multi-Turn Interrogation Framework for Evaluating Persona Agent Consistency Minseo Kim et.al. 2603.25620 null
2026-03-26 Demographic Fairness in Multimodal LLMs: A Benchmark of Gender and Ethnicity Bias in Face Verification Ünsal Öztürk et.al. 2603.25613 null
2026-03-26 Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes Yuqian Fu et.al. 2603.25562 null
2026-03-26 Humans vs Vision-Language Models: A Unified Measure of Narrative Coherence Nikolai Ilinykh et.al. 2603.25537 null
2026-03-26 Investigating the Fundamental Limit: A Feasibility Study of Hybrid-Neural Archival Marcus Armstrong et.al. 2603.25526 null
2026-03-26 An Experimental Comparison of the Most Popular Approaches to Fake News Detection Pietro Dell’Oglio et.al. 2603.25501 null
2026-03-26 Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO Manipulation Pei Chen et.al. 2603.25500 null
2026-03-26 EcoThink: A Green Adaptive Inference Framework for Sustainable and Accessible Agents Linxiao Li et.al. 2603.25498 null
2026-03-26 GridVAD: Open-Set Video Anomaly Detection via Spatial Reasoning over Stratified Frame Grids Mohamed Eltahir et.al. 2603.25467 null
2026-03-26 CLAR: CIF-Localized Alignment for Retrieval-Augmented Speech LLM-Based Contextual ASR Shangkun Huang et.al. 2603.25460 null
2026-03-25 Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini Ruofei Du et.al. 2603.24591 null
2026-03-25 MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination Zhuo Li et.al. 2603.24579 null
2026-03-25 Vision-Language Models vs Human: Perceptual Image Quality Assessment Imran Mehmood et.al. 2603.24578 null
2026-03-25 VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models Qijia He et.al. 2603.24575 null
2026-03-25 Scaling Recurrence-aware Foundation Models for Clinical Records via Next-Visit Prediction Haresh Rengaraj Rajamohan et.al. 2603.24562 null
2026-03-25 LensWalk: Agentic Video Understanding by Planning How You See in Videos Keliang Li et.al. 2603.24558 null
2026-03-25 Evaluating Chunking Strategies For Retrieval-Augmented Generation in Oil and Gas Enterprise Documents Samuel Taiwo et.al. 2603.24556 null
2026-03-25 Representation Learning to Study Temporal Dynamics in Tutorial Scaffolding Conrad Borchers et.al. 2603.24535 null
2026-03-25 UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience Zichuan Lin et.al. 2603.24533 null
2026-03-25 Cross-Modal Prototype Alignment and Mixing for Training-Free Few-Shot Classification Dipam Goswami et.al. 2603.24528 null
2026-03-25 AVO: Agentic Variation Operators for Autonomous Evolutionary Search Terry Chen et.al. 2603.24517 null
2026-03-25 Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models Siqi Liu et.al. 2603.24484 null
2026-03-25 Mechanic: Sorrifier-Driven Formal Decomposition Workflow for Automated Theorem Proving Ruichen Qiu et.al. 2603.24465 null
2026-03-25 Unleashing Vision-Language Semantics for Deepfake Video Detection Jiawen Zhu et.al. 2603.24454 null
2026-03-25 Enes Causal Discovery Alexis Kafantaris et.al. 2603.24436 null
2026-03-25 OneSearch-V2: The Latent Reasoning Enhanced Self-distillation Generative Search Framework Ben Chen et.al. 2603.24422 null
2026-03-25 PINGALA: Prosody-Aware Decoding for Sanskrit Poetry Generation Manoj Balaji Jagadeeshan et.al. 2603.24413 null
2026-03-25 AI-Supervisor: Autonomous AI Research Supervision via a Persistent Research World Model Yunbo Long et.al. 2603.24402 null
2026-03-25 3D-Mix for VLA: A Plug-and-Play Module for Integrating VGGT-based 3D Information into Vision-Language-Action Models Bin Yu et.al. 2603.24393 null
2026-03-25 When AI Meets Early Childhood Education: Large Language Models as Assessment Teammates in Chinese Preschools Xingming Li et.al. 2603.24389 null
2026-03-25 PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks Cheng Cui et.al. 2603.24373 null
2026-03-25 LATS: Large Language Model Assisted Teacher-Student Framework for Multi-Agent Reinforcement Learning in Traffic Signal Control Yifeng Zhang et.al. 2603.24361 null
2026-03-25 Enhancing Efficiency and Performance in Deepfake Audio Detection through Neuron-level dropin & Neuroplasticity Mechanisms Yupei Li et.al. 2603.24343 null
2026-03-25 Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual Processing Cheng Cui et.al. 2603.24326 null
2026-03-25 Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning Dogan Urgun et.al. 2603.24324 null
2026-03-25 Language-Assisted Image Clustering Guided by Discriminative Relational Signals and Adaptive Semantic Centers Jun Ma et.al. 2603.24275 null
2026-03-25 Semantic Alignment across Ancient Egyptian Language Stages via Normalization-Aware Multitask Learning He Huang et.al. 2603.24258 null
2026-03-25 Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning Tommaso Galliena et.al. 2603.24257 null
2026-03-25 B-MoE: A Body-Part-Aware Mixture-of-Experts “All Parts Matter” Approach to Micro-Action Recognition Nishit Poddar et.al. 2603.24245 null
2026-03-25 Optimizing Multilingual LLMs via Federated Learning: A Study of Client Language Composition Aleix Sant et.al. 2603.24242 null
2026-03-25 UniScale: Synergistic Entire Space Data and Model Scaling for Search Ranking Liren Yu et.al. 2603.24226 null
2026-03-25 RVLM: Recursive Vision-Language Models with Adaptive Depth Nicanor Mayumu et.al. 2603.24224 null
2026-03-25 Environment-Grounded Multi-Agent Workflow for Autonomous Penetration Testing Michael Somma et.al. 2603.24221 null
2026-03-25 Who Benefits from RAG? The Role of Exposure, Utility and Attribution Bias Mahdi Dehghan et.al. 2603.24218 null
2026-03-25 SumRank: Aligning Summarization Models for Long-Document Listwise Reranking Jincheng Feng et.al. 2603.24204 null
2026-03-25 Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search Yulin Shen et.al. 2603.24203 null
2026-03-25 A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula Cansu Sancaktar et.al. 2603.24202 null
2026-03-25 RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution Yushuai Song et.al. 2603.24198 null
2026-03-25 Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection Adhemar de Senneville et.al. 2603.24181 null
2026-03-25 Towards Automated Crowdsourced Testing via Personified-LLM Shengcheng Yu et.al. 2603.24160 null
2026-03-24 MedObvious: Exposing the Medical Moravec’s Paradox in VLMs via Clinical Triage Ufaq Khan et.al. 2603.23501 null
2026-03-24 VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions Adrian Bulat et.al. 2603.23495 null
2026-03-24 Failure of contextual invariance in gender inference with large language models Sagar Kumar et.al. 2603.23485 null
2026-03-24 SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning Haoyu Huang et.al. 2603.23483 null
2026-03-24 ReqFusion: A Multi-Provider Framework for Automated PEGS Analysis Across Software Domains Muhammad Khalid et.al. 2603.23482 null
2026-03-24 UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation Jiaying Lin et.al. 2603.23478 null
2026-03-24 Evidence of political bias in search engines and language models before major elections Íris Damião et.al. 2603.23474 null
2026-03-24 ConceptCoder: Improve Code Reasoning via Concept Learning Md Mahbubur Rahman et.al. 2603.23470 null
2026-03-24 DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection Gautam Rajendrakumar Gare et.al. 2603.23455 null
2026-03-24 3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding Yiping Chen et.al. 2603.23447 null
2026-03-24 Evaluating LLM-Based Test Generation Under Software Evolution Sabaat Haroon et.al. 2603.23443 null
2026-03-24 Similarity-Aware Mixture-of-Experts for Data-Efficient Continual Learning Connor Mclaughlin et.al. 2603.23436 null
2026-03-24 SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling Yiqi Zhang et.al. 2603.23414 null
2026-03-24 Beyond Preset Identities: How Agents Form Stances and Boundaries in Generative Societies Hanzhong Zhang et.al. 2603.23406 null
2026-03-24 Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning Jiacheng Hua et.al. 2603.23404 null
2026-03-24 Off-Policy Value-Based Reinforcement Learning for Large Language Models Peng-Yuan Wang et.al. 2603.23355 null
2026-03-24 Leveraging LLMs and Social Media to Understand User Perception of Smartphone-Based Earthquake Early Warnings Hanjing Wang et.al. 2603.23322 null
2026-03-24 ARGENT: Adaptive Hierarchical Image-Text Representations Chuong Huynh et.al. 2603.23311 null
2026-03-24 Curriculum-Driven 3D CT Report Generation via Language-Free Visual Grafting and Zone-Constrained Compression V. K. Cody Bumgardner et.al. 2603.23308 null
2026-03-24 Designing Agentic AI-Based Screening for Portfolio Investment Mehmet Caner et.al. 2603.23300 null
2026-03-24 From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service Haoyu He et.al. 2603.23172 null
2026-03-24 Robust Safety Monitoring of Language Models via Activation Watermarking Toluwani Aremu et.al. 2603.23171 null
2026-03-24 Conformal Cross-Modal Active Learning Huy Hoang Nguyen et.al. 2603.23159 null
2026-03-24 Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models Massimiliano Pappa et.al. 2603.23149 null
2026-03-24 Why AI-Generated Text Detection Fails: Evidence from Explainable AI Beyond Benchmark Accuracy Shushanta Pudasaini et.al. 2603.23146 null
2026-03-24 Can Language Models Pass Software Testing Certification Exams? a case study Fitash Ul Haq et.al. 2603.23142 null
2026-03-24 HGNet: Scalable Foundation Model for Automated Knowledge Graph Generation from Scientific Literature Devvrat Joshi et.al. 2603.23136 null
2026-03-24 InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual Guidance Dongwei Pan et.al. 2603.23132 null
2026-03-24 Polaris: A Gödel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair Aditya Kakade et.al. 2603.23129 null
2026-03-24 From Questions to Trust Reports: A LLM-IR Framework for the TREC 2025 DRAGUN Track Ignacy Alwasiak et.al. 2603.23125 null
2026-03-24 SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions Jinzhe Tu et.al. 2603.23118 null
2026-03-24 TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches Zhengxian Huang et.al. 2603.23117 null
2026-03-24 AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image Detection Yangxin Yu et.al. 2603.23115 null
2026-03-24 Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment Adrian Sauter et.al. 2603.23114 null
2026-03-24 When Language Models Lose Their Mind: The Consequences of Brain Misalignment Gabriele Merlin et.al. 2603.23091 null
2026-03-24 MedCausalX: Adaptive Causal Reasoning with Self-Reflection for Trustworthy Medical Vision-Language Models Jianxin Lin et.al. 2603.23085 null
2026-03-24 Good for the Planet, Bad for Me? Intended and Unintended Consequences of AI Energy Consumption Disclosure Michael Klesel et.al. 2603.23075 null
2026-03-24 Can an LLM Detect Instances of Microservice Infrastructure Patterns? Carlos Eduardo Duarte et.al. 2603.23073 null
2026-03-24 MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding Basit Alawode et.al. 2603.23067 null
2026-03-24 Prompt Amplification and Zero-Shot Late Fusion in Audio-Language Models for Speech Emotion Recognition Saurabh Kataria et.al. 2603.23057 null
2026-03-23 VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding Ruoliu Yang et.al. 2603.22285 null
2026-03-23 ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model Haichao Zhang et.al. 2603.22281 null
2026-03-23 DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models Zhide Zhong et.al. 2603.22280 null
2026-03-23 3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing Haoyu Zhen et.al. 2603.22279 null
2026-03-23 The Dual Mechanisms of Spatial Reasoning in Vision-Language Models Kelly Cui et.al. 2603.22278 null
2026-03-23 Scaling DoRA: High-Rank Adaptation via Factored Norms and Fused Kernels Alexandra Zelenin et.al. 2603.22276 null
2026-03-23 Greater accessibility can amplify discrimination in generative AI Carolin Holtermann et.al. 2603.22260 null
2026-03-23 Confidence-Based Decoding is Provably Efficient for Diffusion Language Models Changxiao Cai et.al. 2603.22248 null
2026-03-23 RotorMap and Quantum Fingerprints of DNA Sequences via Rotary Position Embeddings Danylo Yakymenko et.al. 2603.22245 null
2026-03-23 MemDLM: Memory-Enhanced DLM Training Zehua Pei et.al. 2603.22241 null
2026-03-23 SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation Sashuai Zhou et.al. 2603.22228 null
2026-03-23 Gumbel Distillation for Parallel Text Generation Chi Zhang et.al. 2603.22216 null
2026-03-23 Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models Tom Biskupski et.al. 2603.22214 null
2026-03-23 SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection Kexian Tang et.al. 2603.22213 null
2026-03-23 Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance Learning Daniel Shao et.al. 2603.22198 null
2026-03-23 CayleyPy-4: AI-Holography. Towards analogs of holographic string dualities for AI tasks A. Chervov et.al. 2603.22195 null
2026-03-23 Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement Junrong Guo et.al. 2603.22187 null
2026-03-23 Enhancing Document-Level Machine Translation via Filtered Synthetic Corpora and Two-Stage LLM Adaptation Ireh Kim et.al. 2603.22186 null
2026-03-23 Revisiting Quantum Code Generation: Where Should Domain Knowledge Live? Oscar Novo et.al. 2603.22184 null
2026-03-23 MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management Jack W O’Sullivan et.al. 2603.22179 null
2026-03-22 CVT-Bench: Counterfactual Viewpoint Transformations Reveal Unstable Spatial Representations in Multimodal LLMs Shanmukha Vellamcheti et.al. 2603.21114 null
2026-03-22 ResPrune: Text-Conditioned Subspace Reconstruction for Visual Token Pruning in Large Vision-Language Models Xu Li et.al. 2603.21105 null
2026-03-22 Mixture of Chapters: Scaling Learnt Memory in Transformers Tasmay Pankaj Tibrewal et.al. 2603.21096 null
2026-03-22 Evaluating Reasoning-Based Scaffolds for Human-AI Co-Annotation: The ReasonAlign Annotation Protocol Smitha Muthya Sudheendra et.al. 2603.21094 null
2026-03-22 CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models Nan Zhou et.al. 2603.21077 null
2026-03-22 NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection Yupeng Zhang et.al. 2603.21069 null
2026-03-22 LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning Jianing Wang et.al. 2603.21065 null
2026-03-22 When Minor Edits Matter: LLM-Driven Prompt Attack for Medical VLM Robustness in Ultrasound Yasamin Medghalchi et.al. 2603.21047 null
2026-03-22 Left Behind: Cross-Lingual Transfer as a Bridge for Low-Resource Languages in Large Language Models Abdul-Salem Beibitkhan et.al. 2603.21036 null
2026-03-22 TabPFN Extensions for Interpretable Geotechnical Modelling Taiga Saito et.al. 2603.21033 null
2026-03-22 KLDrive: Fine-Grained 3D Scene Reasoning for Autonomous Driving based on Knowledge Graph Ye Tian et.al. 2603.21029 null
2026-03-22 SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration Zihan Guo et.al. 2603.21019 null
2026-03-22 Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO Jinquan Zheng et.al. 2603.21016 null
2026-03-22 CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs Florent Draye et.al. 2603.21014 null
2026-03-22 SkinCLIP-VL: Consistency-Aware Vision-Language Learning for Multimodal Skin Cancer Diagnosis Zhixiang Lu et.al. 2603.21010 null
2026-03-22 ECI: Effective Contrastive Information to Evaluate Hard-Negatives Aarush Sinha et.al. 2603.20990 null
2026-03-22 Can we automatize scientific discovery in the cognitive sciences? Akshay K. Jagadish et.al. 2603.20988 null
2026-03-22 Consistent but Dangerous: Per-Sample Safety Classification Reveals False Reliability in Medical Vision-Language Models Binesh Sadanandan et.al. 2603.20985 null
2026-03-21 Detection of adversarial intent in Human-AI teams using LLMs Abed K. Musaffar et.al. 2603.20976 null
2026-03-21 DiscoUQ: Structured Disagreement Analysis for Uncertainty Quantification in LLM Agent Ensembles Bo Jiang et.al. 2603.20975 null
2026-03-20 LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation Jiazheng Xing et.al. 2603.20192 null
2026-03-20 IndoorR2X: Indoor Robot-to-Everything Coordination with LLM-Driven Planning Fan Yang et.al. 2603.20182 null
2026-03-20 Adaptive Greedy Frame Selection for Long Video Understanding Yuning Huang et.al. 2603.20180 null
2026-03-20 AI Agents Can Already Autonomously Perform Experimental High Energy Physics Eric A. Moreno et.al. 2603.20179 null
2026-03-20 Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation Richard J. Young et.al. 2603.20172 null
2026-03-20 Learning Dynamic Belief Graphs for Theory-of-mind Reasoning Ruxiao Chen et.al. 2603.20170 null
2026-03-20 The Robot’s Inner Critic: Self-Refinement of Social Behaviors through VLM-based Replanning Jiyu Lim et.al. 2603.20164 null
2026-03-20 Evaluating Evidence Grounding Under User Pressure in Instruction-Tuned Language Models Sai Koneru et.al. 2603.20162 null
2026-03-20 Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models Qi Cao et.al. 2603.20161 null
2026-03-20 Reasoning Gets Harder for LLMs Inside A Dialogue Ivan Kartáč et.al. 2603.20133 null
2026-03-20 Revisiting Gene Ontology Knowledge Discovery with Hierarchical Feature Selection and Virtual Study Group of AI Agents Cen Wan et.al. 2603.20132 null
2026-03-20 Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models Wenjing Hong et.al. 2603.20122 null
2026-03-20 Current LLMs still cannot ‘talk much’ about grammar modules: Evidence from syntax Mohammed Q. Shormani et.al. 2603.20114 null
2026-03-20 The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$ -Calculus Amartya Roy et.al. 2603.20105 null
2026-03-20 Pitfalls in Evaluating Interpretability Agents Tal Haklay et.al. 2603.20101 null
2026-03-20 An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models Yuming Feng et.al. 2603.20100 null
2026-03-20 Beyond Accuracy: Towards a Robust Evaluation Methodology for AI Systems for Language Education James Edgell et.al. 2603.20088 null
2026-03-20 Agentic Harness for Real-World Compilers Yingwei Zheng et.al. 2603.20075 null
2026-03-20 Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs Wenjian Zhang et.al. 2603.20046 null
2026-03-20 LoASR-Bench: Evaluating Large Speech Language Models on Low-Resource Automatic Speech Recognition Across Language Families Jianan Chen et.al. 2603.20042 null
2026-03-20 Detached Skip-Links and $R$ -Probe: Decoupling Feature Aggregation from Gradient Propagation for MLLM OCR Ziye Yuan et.al. 2603.20020 null
2026-03-20 ReViSQL: Achieving Human-Level Text-to-SQL Yuxuan Zhu et.al. 2603.20004 null
2026-03-20 An Agentic Approach to Generating XAI-Narratives Yifan He et.al. 2603.20003 null
2026-03-20 MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI Rozain Shakeel et.al. 2603.19993 null
2026-03-20 Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States Yurun Yuan et.al. 2603.19987 null
2026-03-20 HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction Ruicheng Yuan et.al. 2603.19957 null
2026-03-20 Large Language Models and Stock Investing: Is the Human Factor Required? Ricardo Crisostomo et.al. 2603.19944 null
2026-03-20 Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents Luiz C. Borro et.al. 2603.19935 null
2026-03-20 SAGE: Sustainable Agent-Guided Expert-tuning for Culturally Attuned Translation in Low-Resource Southeast Asia Zhixiang Lu et.al. 2603.19931 null
2026-03-20 DALI: LLM-Agent Enhanced Dual-Stream Adaptive Leadership Identification for Group Recommendations Boxun Song et.al. 2603.19909 null
2026-03-20 Utility-Guided Agent Orchestration for Efficient LLM Tool Use Boyan Liu et.al. 2603.19896 null
2026-03-20 What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time Dong Yan et.al. 2603.19880 null
2026-03-20 MedQ-Engine: A Closed-Loop Data Engine for Evolving MLLMs in Medical Image Quality Assessment Jiyao Liu et.al. 2603.19863 null
2026-03-20 IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment Simone Magistri et.al. 2603.19862 null
2026-03-20 Beyond detection: cooperative multi-agent reasoning for rapid onboard EO crisis response Alejandro D. Mousist et.al. 2603.19858 null
2026-03-20 Overreliance on AI in Information-seeking from Video Content Anders Giovanni Møller et.al. 2603.19843 null
2026-03-20 FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization Chiyu Ma et.al. 2603.19835 null
2026-03-20 Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech? Lokesh Kumar et.al. 2603.19831 null
2026-03-19 Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Xianjin Wu et.al. 2603.19235 null
2026-03-19 Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens Yuqing Wang et.al. 2603.19232 null
2026-03-19 FinTradeBench: A Financial Reasoning Benchmark for LLMs Yogesh Agrawal et.al. 2603.19225 null
2026-03-19 Online Learning and Equilibrium Computation with Ranking Feedback Mingyang Liu et.al. 2603.19221 null
2026-03-19 LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs Keda Tao et.al. 2603.19217 null
2026-03-19 Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders Shang-Jui Ray Kuo et.al. 2603.19209 null
2026-03-19 Tinted Frames: Question Framing Blinds Vision-Language Models Wan-Cyuan Fan et.al. 2603.19203 null
2026-03-19 How Auditory Knowledge in LLM Backbones Shapes Audio Language Models: A Holistic Evaluation Ke-Han Lu et.al. 2603.19195 null
2026-03-19 Box Maze: A Process-Control Architecture for Reliable LLM Reasoning Zou Qiang et.al. 2603.19182 null
2026-03-19 Evaluating Counterfactual Strategic Reasoning in Large Language Models Dimitrios Georgousis et.al. 2603.19167 null
2026-03-19 Meanings and Measurements: Multi-Agent Probabilistic Grounding for Vision-Language Navigation Swagat Padhan et.al. 2603.19166 null
2026-03-19 PPI is the Difference Estimator: Recognizing the Survey Sampling Roots of Prediction-Powered Inference Reagan Mozer et.al. 2603.19160 null
2026-03-19 ADAPT: Attention Driven Adaptive Prompt Scheduling and InTerpolating Orthogonal Complements for Rare Concepts Generation Kwanyoung Lee et.al. 2603.19157 null
2026-03-19 VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models Chonghan Liu et.al. 2603.19152 null
2026-03-19 Optimal Splitting of Language Models from Mixtures to Specialized Domains Skyler Seto et.al. 2603.19149 null
2026-03-19 UGID: Unified Graph Isomorphism for Debiasing Large Language Models Zikang Ding et.al. 2603.19144 null
2026-03-19 GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning Yiren Lu et.al. 2603.19137 null
2026-03-19 A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference Yida Zhang et.al. 2603.19133 null
2026-03-19 From Inference Efficiency to Embodied Efficiency: Revisiting Efficiency Metrics for Vision-Language-Action Models Zhuofan Li et.al. 2603.19131 null
2026-03-19 On Optimizing Multimodal Jailbreaks for Spoken Language Models Aravind Krishnan et.al. 2603.19127 null
2026-03-19 MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model Youngwan Lee et.al. 2603.18892 null
2026-03-19 PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment Tianci Luo et.al. 2603.18891 null
2026-03-19 Evaluating LLM-Generated Lessons from the Language Learning Students’ Perspective: A Short Case Study on Duolingo Carlos Rafael Catalan et.al. 2603.18873 null
2026-03-19 DriftGuard: Mitigating Asynchronous Data Drift in Federated Learning Yizhou Han et.al. 2603.18872 null
2026-03-19 Bridging Network Fragmentation: A Semantic-Augmented DRL Framework for UAV-aided VANETs Gaoxiang Cao et.al. 2603.18871 null
2026-03-19 Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders Yana Veitsman et.al. 2603.18863 null
2026-03-19 RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models Xiao Feng et.al. 2603.18859 null
2026-03-19 Motion-o: Trajectory-Grounded Video Reasoning Bishoy Galoaa et.al. 2603.18856 null
2026-03-19 BeamAgent: LLM-Aided MIMO Beamforming with Decoupled Intent Parsing and Alternating Optimization for Joint Site Selection and Precoding Xiucheng Wang et.al. 2603.18855 null
2026-03-19 HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models Xiangyu Bai et.al. 2603.18850 null
2026-03-19 V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors Songjia He et.al. 2603.18811 null
2026-03-19 dTRPO: Trajectory Reduction in Policy Optimization of Diffusion Large Language Models Wenxuan Zhang et.al. 2603.18806 null
2026-03-19 Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation Yuchen Li et.al. 2603.18795 null
2026-03-19 Functional Subspace Watermarking for Large Language Models Zikang Ding et.al. 2603.18793 null
2026-03-19 SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for RGB+Thermal 3D Reconstruction Vsevolod Skorokhodov et.al. 2603.18774 null
2026-03-19 Empathetic Motion Generation for Humanoid Educational Robots via Reasoning-Guided Vision–Language–Motion Diffusion Architecture Fuze Sun et.al. 2603.18771 null
2026-03-19 Implicit Grading Bias in Large Language Models: How Writing Style Affects Automated Assessment Across Math, Programming, and Essay Tasks Rudra Jadhav et.al. 2603.18765 null
2026-03-19 Are complicated loss functions necessary for teaching LLMs to reason? Gabriele Carrino et.al. 2603.18756 null
2026-03-19 Automatic detection of Gen-AI texts: A comparative framework of neural models Cristian Buttaro et.al. 2603.18750 null
2026-03-19 Measuring and Exploiting Confirmation Bias in LLM-Assisted Security Code Review Dimitris Mitropoulos et.al. 2603.18740 null
2026-03-18 Unified Spatio-Temporal Token Scoring for Efficient Video VLMs Jianrui Zhang et.al. 2603.18004 null
2026-03-18 Universal Skeleton Understanding via Differentiable Rendering and MLLMs Ziyi Wang et.al. 2603.18003 null
2026-03-18 Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models Kevin Qu et.al. 2603.18002 null
2026-03-18 The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering Yigit Ekin et.al. 2603.17998 null
2026-03-18 Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding Shuyao Shi et.al. 2603.17980 null
2026-03-18 Beyond Muon: MUD (MomentUm Decorrelation) for Faster Transformer Training Ben S. Southworth et.al. 2603.17970 null
2026-03-18 ConGA: Guidelines for Contextual Gender Annotation. A Framework for Annotating Gender in Machine Translation Argentina Anna Rescigno et.al. 2603.17962 null
2026-03-18 Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures Chiara Manna et.al. 2603.17952 null
2026-03-18 VideoAtlas: Navigating Long-Form Video in Logarithmic Compute Mohamed Eltahir et.al. 2603.17948 null
2026-03-18 Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing Raghavv Goel et.al. 2603.17942 null
2026-03-18 Training Diffusion Language Models for Black-Box Optimization Zipeng Sun et.al. 2603.17919 null
2026-03-18 Only relative ranks matter in weight-clustered large language models Borja Aizpurua et.al. 2603.17917 null
2026-03-18 IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia Priyaranjan Pattnayak et.al. 2603.17915 null
2026-03-18 Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages Yue Zhao et.al. 2603.17912 null
2026-03-18 Differential Privacy in Generative AI Agents: Analysis and Optimal Tradeoffs Ya-Ting Yang et.al. 2603.17902 null
2026-03-18 RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference Arpit Singh Gautam et.al. 2603.17891 null
2026-03-18 AI-Assisted Goal Setting Improves Goal Progress Through Social Accountability Michel Schimpf et.al. 2603.17887 null
2026-03-18 DebugLM: Learning Traceable Training Data Provenance for LLMs Wenjie Jacky Mo et.al. 2603.17884 null
2026-03-18 Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval Md. Asraful Haque et.al. 2603.17872 null
2026-03-18 ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models Zhou Fang et.al. 2603.17850 null
2026-03-18 Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions Madhav S. Baidya et.al. 2603.17522 null
2026-03-18 PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation Jianjian Yin et.al. 2603.17520 null
2026-03-18 Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality Mengyu Bu et.al. 2603.17512 null
2026-03-18 Interpreting Context-Aware Human Preferences for Multi-Objective Robot Navigation Tharun Sethuraman et.al. 2603.17510 null
2026-03-18 Inducing Epistemological Humility in Large Language Models: A Targeted SFT Approach to Reducing Hallucination Cem Uluoglakci et.al. 2603.17504 null
2026-03-18 Learning When to Attend: Conditional Memory Access for Long-Context LLMs Sakshi Choudhary et.al. 2603.17484 null
2026-03-18 Humans and transformer LMs: Abstraction drives language learning Jasper Jian et.al. 2603.17475 null
2026-03-18 Efficient Soft Actor-Critic with LLM-Based Action-Level Guidance for Continuous Control Hao Ma et.al. 2603.17468 null
2026-03-18 VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation Junyoung Kim et.al. 2603.17450 null
2026-03-18 TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL Tingcheng Bian et.al. 2603.17449 null
2026-03-18 Large Language Models as a Semantic Interface and Ethical Mediator in Neuro-Digital Ecosystems: Conceptual Foundations and a Regulatory Imperative Alexander V. Shenderuk-Zhidkov et.al. 2603.17444 null
2026-03-18 AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement Siqi Pei et.al. 2603.17441 null
2026-03-18 Baguan-TS: A Sequence-Native In-Context Learning Model for Time Series Forecasting with Covariates Linxiao Yang et.al. 2603.17439 null
2026-03-18 ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression Ruibo Fan et.al. 2603.17435 null
2026-03-18 Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare Saikat Maiti et.al. 2603.17419 null
2026-03-18 Is Your LLM-as-a-Recommender Agent Trustable? LLMs’ Recommendation is Easily Hacked by Biases (Preferences) Zichen Tang et.al. 2603.17417 null
2026-03-18 Harnessing the Power of Foundation Models for Accurate Material Classification Qingran Lin et.al. 2603.17390 null
2026-03-18 Efficient Exploration at Scale Seyed Mohammad Asghari et.al. 2603.17378 null
2026-03-18 Shot-Aware Frame Sampling for Video Understanding Mengyu Zhao et.al. 2603.17374 null
2026-03-18 SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems Rima Hazra et.al. 2603.17373 null
2026-03-17 Efficient Reasoning on the Edge Yelysei Bondarenko et.al. 2603.16867 null
2026-03-17 Chronos: Temporal-Aware Conversational Agents with Structured Event Retrieval for Long-Term Memory Sahil Sen et.al. 2603.16862 null
2026-03-17 MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation Abhay Deshpande et.al. 2603.16861 null
2026-03-17 DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models Emily Yue-Ting Jia et.al. 2603.16860 null
2026-03-17 SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Tianyu Xie et.al. 2603.16859 null
2026-03-17 Online Experiential Learning for Language Models Tianzhu Ye et.al. 2603.16856 null
2026-03-17 Internalizing Agency from Reflective Experience Rui Ge et.al. 2603.16843 null
2026-03-17 Prompt Programming for Cultural Bias and Alignment of Large Language Models Maksim Eren et.al. 2603.16827 null
2026-03-17 Surg $Σ$ : A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence Zhitao Zeng et.al. 2603.16822 null
2026-03-17 Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights Yi Chen et.al. 2603.16817 null
2026-03-17 InCoder-32B: Code Foundation Model for Industrial Scenarios Jian Yang et.al. 2603.16790 null
2026-03-17 IOSVLM: A 3D Vision-Language Model for Unified Dental Diagnosis from Intraoral Scans Huimin Xiong et.al. 2603.16781 null
2026-03-17 SOMP: Scalable Gradient Inversion for Large Language Models via Subspace-Guided Orthogonal Matching Pursuit Yibo Li et.al. 2603.16761 null
2026-03-17 TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities Victoria Graf et.al. 2603.16759 null
2026-03-17 Probing Cultural Signals in Large Language Models through Author Profiling Valentin Lafargue et.al. 2603.16749 null
2026-03-17 SpecMoE: Spectral Mixture-of-Experts Foundation Model for Cross-Species EEG Decoding D. Darankoum et.al. 2603.16739 null
2026-03-17 MedCL-Bench: Benchmarking stability-efficiency trade-offs and scaling in biomedical continual learning Min Zeng et.al. 2603.16738 null
2026-03-17 Retrieving Counterfactuals Improves Visual In-Context Learning Guangzhi Xiong et.al. 2603.16737 null
2026-03-17 Differential Harm Propensity in Personalized LLM Agents: The Curious Case of Mental Health Disclosure Caglar Yildirim et.al. 2603.16734 null
2026-03-17 IQuest-Coder-V1 Technical Report Jian Yang et.al. 2603.16733 null
2026-03-17 Visual Distraction Undermines Moral Reasoning in Vision-Language Models Xinyi Yang et.al. 2603.16445 null
2026-03-17 Capability-Guided Compression: Toward Interpretability-Aware Budget Allocation for Large Language Models Rishaank Gupta et.al. 2603.16440 null
2026-03-17 VQKV: High-Fidelity and High-Ratio Cache Compression via Vector-Quantization Yixuan Wang et.al. 2603.16435 null
2026-03-17 From Natural Language to Executable Option Strategies via Large Language Models Haochen Luo et.al. 2603.16434 null
2026-03-17 EngGPT2: Sovereign, Efficient and Open Intelligence G. Ciarfaglia et.al. 2603.16430 null
2026-03-17 An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU Ruijia Yang et.al. 2603.16428 null
2026-03-17 Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences Quan Cheng et.al. 2603.16417 null
2026-03-17 Trained Persistent Memory for Frozen Encoder–Decoder LLMs: Six Architectural Methods Hong Jeong et.al. 2603.16413 null
2026-03-17 RECOVER: Robust Entity Correction via agentic Orchestration of hypothesis Variants for Evidence-based Recovery Abhishek Kumar et.al. 2603.16411 null
2026-03-17 PlotTwist: A Creative Plot Generation Framework with Small Language Models Abhinav Thorat et.al. 2603.16410 null
2026-03-17 Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic Finnur Ágúst Ingimundarson et.al. 2603.16406 null
2026-03-17 Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models Deng Liu et.al. 2603.16382 null
2026-03-17 InViC: Intent-aware Visual Cues for Medical Visual Question Answering Zhisong Wang et.al. 2603.16372 null
2026-03-17 Beyond Grading Accuracy: Exploring Alignment of TAs and LLMs Matthijs Jansen op de Haar et.al. 2603.16357 null
2026-03-17 Toward Experimentation-as-a-Service in 5G/6G: The Plaza6G Prototype for AI-Assisted Trials Sergio Barrachina-Muñoz et.al. 2603.16356 null
2026-03-17 Detecting Sentiment Steering Attacks on RAG-enabled Large Language Models Isha Andrade et.al. 2603.16342 null
2026-03-17 Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits Jia Qing Yap et.al. 2603.16335 null
2026-03-17 Decoding the Critique Mechanism in Large Reasoning Models Hoang Phan et.al. 2603.16331 null
2026-03-17 An Interpretable Machine Learning Framework for Non-Small Cell Lung Cancer Drug Response Analysis Ann Rachel et.al. 2603.16330 null
2026-03-17 A Human-Centred Architecture for Large Language Models-Cognitive Assistants in Manufacturing within Quality Management Systems Marcos Galdino et.al. 2603.16325 null
2026-03-16 Mixture-of-Depths Attention Lianghui Zhu et.al. 2603.15619 null
2026-03-16 HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification Erik Y. Wang et.al. 2603.15617 null
2026-03-16 Mechanistic Origin of Moral Indifference in Language Models Lingyu Li et.al. 2603.15615 null
2026-03-16 From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation Yibin Liu et.al. 2603.15600 null
2026-03-16 OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data Yuwen Du et.al. 2603.15594 null
2026-03-16 Effective Distillation to Hybrid xLSTM Architectures Lukas Hauzenberger et.al. 2603.15590 null
2026-03-16 LEXI: Lossless Exponent Coding for Efficient Inter-Chiplet Communication in Hybrid LLMs Miao Sun et.al. 2603.15589 null
2026-03-16 Mamba-3: Improved Sequence Modeling using State Space Principles Aakash Lahoti et.al. 2603.15569 null
2026-03-16 The PokeAgent Challenge: Competitive and Long-Context Learning at Scale Seth Karten et.al. 2603.15563 null
2026-03-16 Anatomy of a Lie: A Multi-Stage Diagnostic Framework for Tracing Hallucinations in Vision-Language Models Lexiang Xiong et.al. 2603.15557 null
2026-03-16 Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation Yanick Zengaffinen et.al. 2603.15547 null
2026-03-16 InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems Shaojie Shi et.al. 2603.15542 null
2026-03-16 Bridging Local and Global Knowledge: Cascaded Mixture-of-Experts Learning for Near-Shortest Path Routing Yung-Fu Chen et.al. 2603.15541 null
2026-03-16 QiboAgent: a practitioner’s guideline to open source assistants for Quantum Computing code development Lorenzo Esposito et.al. 2603.15538 null
2026-03-16 DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages Alish Kanani et.al. 2603.15530 null
2026-03-16 Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph Zhenheng Tang et.al. 2603.15527 null
2026-03-16 Beyond the Covariance Trap: Unlocking Generalization in Same-Subject Knowledge Editing for Large Language Models Xiyu Liu et.al. 2603.15518 null
2026-03-16 ViX-Ray: A Vietnamese Chest X-Ray Dataset for Vision-Language Models Duy Vu Minh Nguyen et.al. 2603.15513 null
2026-03-16 Not All Invariants Are Equal: Curating Training Data to Accelerate Program Verification with SLMs Ido Pinto et.al. 2603.15510 null
2026-03-16 TabKD: Tabular Knowledge Distillation through Interaction Diversity of Learned Feature Bins Shovon Niverd Pereira et.al. 2603.15481 null
2026-03-14 Improving Visual Reasoning with Iterative Evidence Refinement Zeru Shi et.al. 2603.14117 null
2026-03-14 OasisSimp: An Open-source Asian-English Sentence Simplification Dataset Hannah Liu et.al. 2603.14111 null
2026-03-14 SVD Contextual Sparsity Predictors for Fast LLM Inference Georgii Serbin et.al. 2603.14110 null
2026-03-14 Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors Mark Rofin et.al. 2603.14087 null
2026-03-14 CMHL: Contrastive Multi-Head Learning for Emotionally Consistent Text Classification Menna Elgabry et.al. 2603.14078 null
2026-03-14 Demand-Driven Context: A Methodology for Building Enterprise Knowledge Bases Through Agent Failure Raj Navakoti et.al. 2603.14057 null
2026-03-14 LegacyTranslate: LLM-based Multi-Agent Method for Legacy Code Translation Zahra Moti et.al. 2603.14054 null
2026-03-14 A Theory of Appropriateness That Accounts for Norms of Rationality Joel Z. Leibo et.al. 2603.14050 null
2026-03-14 The Reasoning Bottleneck in Graph-RAG: Structured Prompting and Context Compression for Multi-Hop QA Yasaman Zarinkia et.al. 2603.14045 null
2026-03-14 GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models Zhijie Wang et.al. 2603.14041 null
2026-03-14 SemEval-2026 Task 6: CLARITY – Unmasking Political Question Evasions Konstantinos Thomas et.al. 2603.14027 null
2026-03-14 LLM-Guided Safe Reinforcement Learning for Energy System Topology Reconfiguration Zongyan Zhang et.al. 2603.14018 null
2026-03-14 Multi-Grained Vision-Language Alignment for Domain Generalized Person Re-Identification Jiachen Li et.al. 2603.14012 null
2026-03-14 Measuring Weather Effects and Link Quality Dynamics in LEO Satellite Networks Clemens Lottermoser et.al. 2603.14008 null
2026-03-14 Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning Halimat Afolabi et.al. 2603.13988 null
2026-03-14 Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models Haitao Jiang et.al. 2603.13985 null
2026-03-14 When Visual Privacy Protection Meets Multimodal Large Language Models Xiaofei Hui et.al. 2603.13978 null
2026-03-14 FLUX: Data Worth Training On Gowtham et.al. 2603.13972 null
2026-03-14 EviAgent: Evidence-Driven Agent for Radiology Report Generation Tuoshi Qi et.al. 2603.13956 null
2026-03-14 LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement Chih-Ning Chen et.al. 2603.13952 null
2026-03-13 Visual-ERM: Reward Modeling for Visual Equivalence Ziyu Liu et.al. 2603.13224 null
2026-03-13 A Generative Model of Conspicuous Consumption and Status Signaling Logan Cross et.al. 2603.13220 null
2026-03-13 MoEKD: Mixture-of-Experts Knowledge Distillation for Robust and High-Performing Compressed Code Models Md. Abdul Awal et.al. 2603.13213 null
2026-03-13 Neuron-Aware Data Selection In Instruction Tuning For Large Language Models Xin Chen et.al. 2603.13201 null
2026-03-13 Navig-AI-tion: Navigation by Contextual AI and Spatial Audio Mathias N. Lystbæk et.al. 2603.13200 null
2026-03-13 From Experiments to Expertise: Scientific Knowledge Consolidation for AI-Driven Computational Research Haonan Huang et.al. 2603.13191 null
2026-03-13 LLM Constitutional Multi-Agent Governance J. de Curtò et.al. 2603.13189 null
2026-03-13 Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos Rohith Peddi et.al. 2603.13185 null
2026-03-13 Semantic Invariance in Agentic AI I. de Zarzà et.al. 2603.13173 null
2026-03-13 ESG-Bench: Benchmarking Long-Context ESG Reports for Hallucination Mitigation Siqi Sun et.al. 2603.13154 null
2026-03-13 Developing the PsyCogMetrics AI Lab to Evaluate Large Language Models and Advance Cognitive Science – A Three-Cycle Action Design Science Study Zhiye Jin et.al. 2603.13126 null
2026-03-13 Geometry-Guided Camera Motion Understanding in VideoLLMs Haoan Feng et.al. 2603.13119 null
2026-03-13 AgentRM: An OS-Inspired Resource Manager for LLM Agent Systems Jianshu She et.al. 2603.13110 null
2026-03-13 Evaluating VLMs’ Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences Wenxi Wu et.al. 2603.13100 null
2026-03-13 SldprtNet: A Large-Scale Multimodal Dataset for CAD Generation in Language-Driven 3D Design Ruogu Li et.al. 2603.13098 null
2026-03-13 Breaking the Tuning Barrier: Zero-Hyperparameters Yield Multi-Corner Analysis Via Learned Priors Wei W. Xing et.al. 2603.13092 null
2026-03-13 Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence Seunghwan Bang et.al. 2603.13091 null
2026-03-13 Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach Elena Ryumina et.al. 2603.13056 null
2026-03-13 Topo-R1: Detecting Topological Anomalies via Vision-Language Models Meilong Xu et.al. 2603.13054 null
2026-03-13 Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation Yifeng Liu et.al. 2603.13045 null
2026-03-13 Before and After ChatGPT: Revisiting AI-Based Dialogue Systems for Emotional Support Daeun Lee et.al. 2603.13043 null
2026-03-13 ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models Yanpeng Zhao et.al. 2603.13033 null
2026-03-13 ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning Bangjun Xiao et.al. 2603.13019 null
2026-03-13 A Closed-Form Solution for Debiasing Vision-Language Models with Utility Guarantees Across Modalities and Tasks Tangzheng Lian et.al. 2603.12998 null
2026-03-13 Test-Time Attention Purification for Backdoored Large Vision Language Models Zhifang Zhang et.al. 2603.12989 null
2026-03-13 Delta1 with LLM: symbolic and neural integration for credible and explainable reasoning Yang Xu et.al. 2603.12953 null
2026-03-13 RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics Yuzhi Huang et.al. 2603.12939 null
2026-03-13 MotionAnymesh: Physics-Grounded Articulation for Simulation-Ready Digital Twins WenBo Xu et.al. 2603.12936 null
2026-03-13 Can Fairness Be Prompted? Prompt-Based Debiasing Strategies in High-Stakes Recommendations Mihaela Rotar et.al. 2603.12935 null
2026-03-12 MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning Haozhan Shen et.al. 2603.12266 null
2026-03-12 Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously Yiran Guan et.al. 2603.12262 null
2026-03-12 Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing Baifeng Shi et.al. 2603.12254 null
2026-03-12 EndoCoT: Scaling Endogenous Chain-of-Thought Reasoning in Diffusion Models Xuanlang Dai et.al. 2603.12252 null
2026-03-12 Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models Samy Jelassi et.al. 2603.12248 null
2026-03-12 Separable neural architectures as a primitive for unified predictive and generative intelligence Reza T. Batley et.al. 2603.12244 null
2026-03-12 SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation Jun Luo et.al. 2603.12238 null
2026-03-12 Language Model Teams as Distributed Systems Elizabeth Mieczkowski et.al. 2603.12229 null
2026-03-12 Sparking Scientific Creativity via LLM-Driven Interdisciplinary Inspiration Priyanka Kargupta et.al. 2603.12226 null
2026-03-12 A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition Jiajun Sun et.al. 2603.12221 null
2026-03-12 ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models Yingxin Lai et.al. 2603.12208 null
2026-03-12 CLASP: Defending Hybrid Large Language Models Against Hidden State Poisoning Attacks Alexandre Le Mercier et.al. 2603.12206 null
2026-03-12 IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse Yushi Bai et.al. 2603.12201 null
2026-03-12 Long-Context Encoder Models for Polish Language Understanding Sławomir Dadas et.al. 2603.12191 null
2026-03-12 BehaviorVLM: Unified Finetuning-Free Behavioral Understanding with Vision-Language Reasoning Jingyang Ke et.al. 2603.12176 null
2026-03-12 LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning Haiying Xu et.al. 2603.12166 null
2026-03-12 LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation Feiyu Duan et.al. 2603.12152 null
2026-03-12 IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL Zhoujun Cheng et.al. 2603.12151 null
2026-03-12 Linking Perception, Confidence and Accuracy in MLLMs Yuetian Du et.al. 2603.12149 null
2026-03-12 EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next Ye Pan et.al. 2603.12147 null
2026-03-12 Hoi3DGen: Generating High-Quality Human-Object-Interactions in 3D Agniv Sharma et.al. 2603.12126 null
2026-03-12 Cross-Context Review: Improving LLM Output Quality by Separating Production and Review Sessions Tae-Eun Song et.al. 2603.12123 null
2026-03-12 SommBench: Assessing Sommelier Expertise of Language Models William Brach et.al. 2603.12117 null
2026-03-12 On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents Deyu Zou et.al. 2603.12109 null
2026-03-12 EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation Yan Li et.al. 2603.12108 null
2026-03-12 To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times Thomas Hikaru Clark et.al. 2603.12105 null
2026-03-12 Human-Centred LLM Privacy Audits: Findings and Frictions Dimitri Staufer et.al. 2603.12094 null
2026-03-12 Resource-Efficient Iterative LLM-Based NAS with Feedback Memory Xiaojie Gu et.al. 2603.12091 null
2026-03-12 EmbTracker: Traceable Black-box Watermarking for Federated Language Models Haodong Zhao et.al. 2603.12089 null
2026-03-12 Paper Title: LoV3D: Grounding Cognitive Prognosis Reasoning in Longitudinal 3D Brain MRI via Regional Volume Assessments Zhaoyang Jiang et.al. 2603.12071 null
2026-03-12 Continual Learning with Vision-Language Models via Semantic-Geometry Preservation Chiyuan He et.al. 2603.12055 null
2026-03-12 Frequentist Consistency of Prior-Data Fitted Networks for Causal Inference Valentyn Melnychuk et.al. 2603.12037 null
2026-03-12 Cascade: Composing Software-Hardware Attack Gadgets for Adversarial Threat Amplification in Compound AI Systems Sarbartha Banerjee et.al. 2603.12023 null
2026-03-12 CrossEarth-SAR: A SAR-Centric and Billion-Scale Geospatial Foundation Model for Domain Generalizable Semantic Segmentation Ziqi Ye et.al. 2603.12008 null
2026-03-12 BTZSC: A Benchmark for Zero-Shot Text Classification Across Cross-Encoders, Embedding Models, Rerankers and LLMs Ilias Aarab et.al. 2603.11991 null
2026-03-12 LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories Qianpu Sun et.al. 2603.11987 null
2026-03-12 HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios Jiayue Pu et.al. 2603.11975 null
2026-03-12 CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading Pranav Raikote et.al. 2603.11957 null
2026-03-12 PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents Minjia Wang et.al. 2603.11955 null
2026-03-12 Learning Transferable Sensor Models via Language-Informed Pretraining Yuliang Chen et.al. 2603.11950 null
2026-03-11 Instruction set for the representation of graphs Ezequiel Lopez-Rubio et.al. 2603.11039 null
2026-03-11 LLMGreenRec: LLM-Based Multi-Agent Recommender System for Sustainable E-Commerce Hao N. Nguyen et.al. 2603.11025 null
2026-03-11 Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style Marvin Limpijankit et.al. 2603.11024 null
2026-03-11 Leech Lattice Vector Quantization for Efficient LLM Compression Tycho F. A. van der Ouderaa et.al. 2603.11021 null
2026-03-11 A Systematic Study of Pseudo-Relevance Feedback with LLMs Nour Jedidi et.al. 2603.11008 null
2026-03-11 The Discrete Charm of the MLP: Binary Routing of Continuous Signals in Transformer Feed-Forward Layers Peter Balogh et.al. 2603.10985 null
2026-03-11 GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations Boyuan Chen et.al. 2603.10978 null
2026-03-11 TOSSS: a CVE-based Software Security Benchmark for Large Language Models Marc Damie et.al. 2603.10969 null
2026-03-11 LLM2Vec-Gen: Generative Embeddings from Large Language Models Parishad BehnamGhader et.al. 2603.10913 null
2026-03-11 When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS Anupam Purwar et.al. 2603.10904 null
2026-03-11 LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation Jinwoo Ahn et.al. 2603.10899 null
2026-03-11 A Hybrid Knowledge-Grounded Framework for Safety and Traceability in Prescription Verification Yichi Zhu et.al. 2603.10891 null
2026-03-11 Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models Yixiu Mao et.al. 2603.10887 null
2026-03-11 From Images to Words: Efficient Cross-Modal Knowledge Distillation to Language Models from Black-box Teachers Ayan Sengupta et.al. 2603.10877 null
2026-03-11 Beyond Sequential Distance: Inter-Modal Distance Invariant Position Encoding Lin Chen et.al. 2603.10863 null
2026-03-11 OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs Yujie Liao et.al. 2603.10862 null
2026-03-11 Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis Yujie Zheng et.al. 2603.10846 null
2026-03-11 PivotAttack: Rethinking the Search Trajectory in Hard-Label Text Attacks via Pivot Words Yuzhi Liang et.al. 2603.10842 null
2026-03-11 Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation Thomas Thebaud et.al. 2603.10827 null
2026-03-11 ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement Learning Xiaofeng Lin et.al. 2603.10823 null
2026-03-11 HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation Hongji Yang et.al. 2603.10814 null
2026-03-11 Nurture-First Agent Development: Building Domain-Expert AI Agents Through Conversational Knowledge Crystallization Linghao Zhang et.al. 2603.10808 null
2026-03-11 Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services Fabrizio Dimino et.al. 2603.10807 null
2026-03-11 Semantic Satellite Communications for Synchronized Audiovisual Reconstruction Fangyu Liu et.al. 2603.10791 null
2026-03-11 Taking Shortcuts for Categorical VQA Using Super Neurons Pierre Musacchio et.al. 2603.10781 null
2026-03-11 Large Language Models as Annotators for Machine Translation Quality Estimation Sidi Wang et.al. 2603.10775 null
2026-03-11 Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness Zhipeng Yang et.al. 2603.10771 null
2026-03-11 mAceReason-Math: A Dataset of High-Quality Multilingual Math Problems Ready For RLVR Konstantin Dobler et.al. 2603.10767 null
2026-03-10 From Data Statistics to Feature Geometry: How Correlations Shape Superposition Lucas Prieto et.al. 2603.09972 null
2026-03-10 Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision People Jazmin Collins et.al. 2603.09964 null
2026-03-10 BEACON: Language-Conditioned Navigation Affordance Prediction under Occlusion Xinyu Gao et.al. 2603.09961 null
2026-03-10 Think Before You Lie: How Reasoning Improves Honesty Ann Yuan et.al. 2603.09957 null
2026-03-10 Towards a Neural Debugger for Python Maximilian Beck et.al. 2603.09951 null
2026-03-10 PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs Jinyue Li et.al. 2603.09943 null
2026-03-10 Model Merging in the Era of Large Language Models: Methods, Applications, and Future Directions Mingyang Song et.al. 2603.09938 null
2026-03-10 Fine-grained Motion Retrieval via Joint-Angle Motion Images and Token-Patch Late Interaction Yao Zhang et.al. 2603.09930 null
2026-03-10 WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition Shan Ning et.al. 2603.09921 null
2026-03-10 MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems Yunhang Qian et.al. 2603.09909 null
2026-03-10 Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports Yuchen Yang et.al. 2603.09896 null
2026-03-10 MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning Yiyang Lu et.al. 2603.09892 null
2026-03-10 Influencing LLM Multi-Agent Dialogue via Policy-Parameterized Prompts Hongbo Bo et.al. 2603.09890 null
2026-03-10 Benchmarking Political Persuasion Risks Across Frontier Large Language Models Zhongren Chen et.al. 2603.09884 null
2026-03-10 Do What I Say: A Spoken Prompt Dataset for Instruction-Following Maike Züfle et.al. 2603.09881 null
2026-03-10 InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing Changyao Tian et.al. 2603.09877 null
2026-03-10 N-gram-like Language Models Predict Reading Time Best James A. Michaelov et.al. 2603.09872 null
2026-03-10 GAST: Gradient-aligned Sparse Tuning of Large Language Models with Data-layer Selection Kai Yao et.al. 2603.09865 null
2026-03-10 SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases Laya Iyer et.al. 2603.09853 null
2026-03-10 RecThinker: An Agentic Framework for Tool-Augmented Reasoning in Recommendation Haobo Zhang et.al. 2603.09843 null
2026-03-10 Understanding the Interplay between LLMs’ Utilisation of Parametric and Contextual Knowledge: A keynote at ECIR 2025 Isabelle Augenstein et.al. 2603.09654 null
2026-03-10 MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants Zuhao Zhang et.al. 2603.09652 null
2026-03-10 MM-tau-p $^2$ : Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings Anupam Purwar et.al. 2603.09643 null
2026-03-10 Tracking Cancer Through Text: Longitudinal Extraction From Radiology Reports Using Open-Source Large Language Models Luc Builtjes et.al. 2603.09638 null
2026-03-10 X-GS: An Extensible Open Framework Unifying 3DGS Architectures with Downstream Multimodal Models Yueen Ma et.al. 2603.09632 null
2026-03-10 Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models Dehua Tao et.al. 2603.09627 null
2026-03-10 Grounding Synthetic Data Generation With Vision and Language Models Ümit Mert Çağlar et.al. 2603.09625 null
2026-03-10 Surgical Repair of Collapsed Attention Heads in ALiBi Transformers Palmer Schallon et.al. 2603.09616 null
2026-03-10 Nonparametric Variational Differential Privacy via Embedding Parameter Clipping Dina El Zein et.al. 2603.09583 null
2026-03-10 More than the Sum: Panorama-Language Models for Adverse Omni-Scenes Weijia Fan et.al. 2603.09573 null
2026-03-10 ALARM: Audio-Language Alignment for Reasoning Models Petr Grinberg et.al. 2603.09556 null
2026-03-10 GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision Lang Sun et.al. 2603.09551 null
2026-03-10 Compartmentalization-Aware Automated Program Repair Jia Hu et.al. 2603.09544 null
2026-03-10 Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization Ming Nie et.al. 2603.09538 null
2026-03-10 Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective Ninghao Wan et.al. 2603.09536 null
2026-03-10 Enhancing Debunking Effectiveness through LLM-based Personality Adaptation Pietro Dell’Oglio et.al. 2603.09533 null
2026-03-10 You Didn’t Have to Say It like That: Subliminal Learning from Faithful Paraphrases Isaia Gisler et.al. 2603.09517 null
2026-03-10 Probing the Reliability of Driving VLMs: From Inconsistent Responses to Grounded Temporal Reasoning Chun-Peng Chang et.al. 2603.09512 null
2026-03-10 EmbC-Test: How to Speed Up Embedded Software Testing Using LLMs and RAG Maximilian Harnot et.al. 2603.09497 null
2026-03-10 Evolving Prompt Adaptation for Vision-Language Models Enming Zhang et.al. 2603.09493 null
2026-03-09 FVG-PT: Adaptive Foreground View-Guided Prompt Tuning for Vision-Language Models Haoyang Li et.al. 2603.08708 null
2026-03-09 Agentic Critical Training Weize Liu et.al. 2603.08706 null
2026-03-09 Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines Akshay Gulati et.al. 2603.08704 null
2026-03-09 Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio Phillip Long et.al. 2603.08683 null
2026-03-09 Exp-Force: Experience-Conditioned Pre-Grasp Force Selection with Vision-Language Models Siqi Shang et.al. 2603.08668 null
2026-03-09 CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation Haodong Li et.al. 2603.08652 null
2026-03-09 PostTrainBench: Can LLM Agents Automate LLM Post-Training? Ben Rank et.al. 2603.08640 null
2026-03-09 UNBOX: Unveiling Black-box visual models with Natural-language Simone Carnemolla et.al. 2603.08639 null
2026-03-09 Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations Jiangye Yuan et.al. 2603.08592 null
2026-03-09 MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation Yutong Shen et.al. 2603.08572 null
2026-03-09 RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback Xiaoying Zhang et.al. 2603.08561 null
2026-03-09 The Neural Compass: Probabilistic Relative Feature Fields for Robotic Search Gabriele Somaschini et.al. 2603.08544 null
2026-03-09 SecAgent: Efficient Mobile GUI Agent with Semantic Context Yiping Xie et.al. 2603.08533 null
2026-03-09 SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement Yi Chen et.al. 2603.08520 null
2026-03-09 AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models Xiaoquan Sun et.al. 2603.08519 null
2026-03-09 Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA Ummar Abbas et.al. 2603.08501 null
2026-03-09 Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models Heng Zhou et.al. 2603.08497 null
2026-03-09 Efficient Credal Prediction through Decalibration Paul Hofman et.al. 2603.08495 null
2026-03-09 Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework Yutong Hu et.al. 2603.08491 null
2026-03-09 Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images Qishun Yang et.al. 2603.08486 null
2026-03-08 AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions Jihyoung Jang et.al. 2603.07394 null
2026-03-08 Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams Jiyeon Kim et.al. 2603.07392 null
2026-03-08 Deterministic Fuzzy Triage for Legal Compliance Classification and Evidence Retrieval Rian Atri et.al. 2603.07390 null
2026-03-07 SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions Saroj Mishra et.al. 2603.07379 null
2026-03-07 Multi-Agentic AI for Conflict-Aware rApp Policy Orchestration in Open RAN Haiyuan Li et.al. 2603.07375 null
2026-03-07 Position: LLMs Must Use Functor-Based and RAG-Driven Bias Mitigation for Fairness Ravi Ranjan et.al. 2603.07368 null
2026-03-07 RILEC: Detection and Generation of L1 Russian Interference Errors in English Learner Texts Darya Kharlamova et.al. 2603.07366 null
2026-03-07 Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes Mohammed Alnemari et.al. 2603.07365 null
2026-03-07 The Yerkes-Dodson Curve for AI Agents: Emergent Cooperation Under Environmental Pressure in Multi-Agent LLM Simulations Ivan Pasichnyk et.al. 2603.07360 null
2026-03-07 How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection Nouran Khallaf et.al. 2603.07346 null
2026-03-07 VisualScratchpad: Inference-time Visual Concepts Analysis in Vision Language Models Hyesu Lim et.al. 2603.07335 null
2026-03-07 The Third Ambition: Artificial Intelligence and the Science of Human Behavior W. Russell Neuman et.al. 2603.07329 null
2026-03-07 FinSheet-Bench: From Simple Lookups to Complex Reasoning, Where LLMs Break on Financial Spreadsheets Jan Ravnik et.al. 2603.07316 null
2026-03-07 Data-Driven Hints in Intelligent Tutoring Systems Sutapa Dey Tithi et.al. 2603.07311 null
2026-03-07 Seeing the Reasoning: How LLM Rationales Influence User Trust and Decision-Making in Factual Verification Tasks Xin Sun et.al. 2603.07306 null
2026-03-07 Tursio for Credit Unions: Powering Structured Data Search with Automated Context Graph Shivani Tripathi et.al. 2603.07304 null
2026-03-07 MAviS: A Multimodal Conversational Assistant For Avian Species Yevheniia Kryklyvets et.al. 2603.07294 null
2026-03-07 LLM-FK: Multi-Agent LLM Reasoning for Foreign Key Detection in Large-Scale Complex Databases Zijian Tang et.al. 2603.07278 null
2026-03-07 From Passive Consumption to Active Interaction: Exploring Interactive LLM Scaffolding to Support Learning Engagement Zixin Chen et.al. 2603.07277 null
2026-03-07 How to Steal Reasoning Without Reasoning Traces Tingwei Zhang et.al. 2603.07267 null
2026-03-06 Multimodal Large Language Models as Image Classifiers Nikita Kisel et.al. 2603.06578 null
2026-03-06 Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion Lijiang Li et.al. 2603.06577 null
2026-03-06 BEVLM: Distilling Semantic Knowledge from LLMs into Bird’s-Eye View Representations Thomas Monninger et.al. 2603.06576 null
2026-03-06 SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning Alejandra Perez et.al. 2603.06570 null
2026-03-06 Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders Boqiang Zhang et.al. 2603.06569 null
2026-03-06 EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking Fangrui Zhu et.al. 2603.06561 null
2026-03-06 RAMoEA-QA: Hierarchical Specialization for Robust Respiratory Audio Question Answering Gaia A. Bertolino et.al. 2603.06542 null
2026-03-06 Speak in Context: Multilingual ASR with Speech Context Alignment via Contrastive Learning Yuchen Zhang et.al. 2603.06505 null
2026-03-06 Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing Anmol Gulati et.al. 2603.06503 null
2026-03-06 COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamics Kartik Sharma et.al. 2603.06495 null
2026-03-06 PONTE: Personalized Orchestration for Natural Language Trustworthy Explanations Vittoria Vineis et.al. 2603.06485 null
2026-03-06 A Mixture-of-Experts Framework for Practical Hybrid-Quantum Models in Credit Card Fraud Detection Rodrigo Chaves et.al. 2603.06473 null
2026-03-06 Do Foundation Models Know Geometry? Probing Frozen Features for Continuous Physical Measurement Yakov Pyotr Shkolnikov et.al. 2603.06459 null
2026-03-06 CaTok: Taming Mean Flows for One-Dimensional Causal Image Tokenization Yitong Chen et.al. 2603.06449 null
2026-03-06 Abductive Reasoning with Syllogistic Forms in Large Language Models Hirohiko Abe et.al. 2603.06428 null
2026-03-06 From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring Minh Hoang Nguyen et.al. 2603.06424 null
2026-03-06 Before You Hand Over the Wheel: Evaluating LLMs for Security Incident Analysis Sourov Jajodia et.al. 2603.06422 null
2026-03-06 Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason’s Selection Task Hirohiko Abe et.al. 2603.06416 null
2026-03-06 Adapter-Augmented Bandits for Online Multi-Constrained Multi-Modal Inference Scheduling Xianzhi Zhang et.al. 2603.06403 null
2026-03-06 Talk Freely, Execute Strictly: Schema-Gated Agentic AI for Flexible and Reproducible Scientific Workflows Joel Strickland et.al. 2603.06394 null
2026-03-06 MoEMambaMIL: Structure-Aware Selective State Space Modeling for Whole-Slide Image Analysis Dongqing Xie et.al. 2603.06378 null
2026-03-06 OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis Yuxuan Fan et.al. 2603.06366 null
2026-03-06 ESAA-Security: An Event-Sourced, Verifiable Architecture for Agent-Assisted Security Audits of AI-Generated Code Elzo Brito dos Santos Filho et.al. 2603.06365 null
2026-03-06 A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management Yang Liu et.al. 2603.06358 null
2026-03-06 MoEless: Efficient MoE LLM Serving via Serverless Computing Hanfei Yu et.al. 2603.06350 null
2026-03-06 Transparent AI for Mathematics: Transformer-Based Large Language Models for Mathematical Entity Relationship Extraction with XAI Tanjim Taharat Aurpa et.al. 2603.06348 null
2026-03-06 K-MaT: Knowledge-Anchored Manifold Transport for Cross-Modal Prompt Learning in Medical Imaging Jiajun Zeng et.al. 2603.06340 null
2026-03-05 POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation Zeju Qiu et.al. 2603.05500 null
2026-03-05 The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks Shangwen Sun et.al. 2603.05498 null
2026-03-05 Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation Helena Casademunt et.al. 2603.05494 null
2026-03-05 NL2GDS: LLM-aided interface for Open Source Chip Design Max Eland et.al. 2603.05489 null
2026-03-05 Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought Siddharth Boppana et.al. 2603.05488 null
2026-03-05 Observing and Controlling Features in Vision-Language-Action Models Hugo Buurmeijer et.al. 2603.05487 null
2026-03-05 Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval Artem Vazhentsev et.al. 2603.05471 null
2026-03-05 HALP: Detecting Hallucinations in Vision-Language Models without Generating a Single Token Sai Akhil Kogilathota et.al. 2603.05465 null
2026-03-05 Beyond Scattered Acceptance: Fast and Coherent Inference for DLMs via Longest Stable Prefixes Pengxiang Li et.al. 2603.05454 null
2026-03-05 FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling Ted Zadouri et.al. 2603.05451 null
2026-03-05 Distributed Partial Information Puzzles: Examining Common Ground Construction Under Epistemic Asymmetry Yifan Zhu et.al. 2603.05450 null
2026-03-05 Ensembling Language Models with Sequential Monte Carlo Robin Shing Moon Chan et.al. 2603.05432 null
2026-03-05 An Exploration-Analysis-Disambiguation Reasoning Framework for Word Sense Disambiguation with Low-Parameter LLMs Deshan Sumanathilaka et.al. 2603.05400 null
2026-03-05 Legal interpretation and AI: from expert systems to argumentation and LLMs Václav Janeček et.al. 2603.05392 null
2026-03-05 ORMOT: A Dataset and Framework for Omnidirectional Referring Multi-Object Tracking Sijia Chen et.al. 2603.05384 null
2026-03-05 Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection Junchuan Zhao et.al. 2603.05373 null
2026-03-05 Progressive Residual Warmup for Language Model Pretraining Tianhao Chen et.al. 2603.05369 null
2026-03-05 DiSCTT: Consensus-Guided Self-Curriculum for Efficient Test-Time Adaptation in Reasoning Mohammad Mahdi Moradi et.al. 2603.05357 null
2026-03-05 PersianPunc: A Large-Scale Dataset and BERT-Based Approach for Persian Punctuation Restoration Mohammad Javad Ranjbar Kalahroodi et.al. 2603.05314 null
2026-03-05 Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution Qiao Jin et.al. 2603.05308 null
2026-03-04 A Dual-Helix Governance Approach Towards Reliable Agentic AI for WebGIS Development Boyuan et.al. 2603.04390 null
2026-03-04 TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning Maximilian von Klinski et.al. 2603.04380 null
2026-03-04 Robustness of Agentic AI Systems via Adversarially-Aligned Jacobian Regularization Furkan Mumcu et.al. 2603.04378 null
2026-03-04 LLM-supported 3D Modeling Tool for Radio Radiance Field Reconstruction Chengling Xu et.al. 2603.04368 null
2026-03-04 Efficient Refusal Ablation in LLM through Optimal Transport Geraldin Nanfack et.al. 2603.04355 null
2026-03-04 RANGER: Sparsely-Gated Mixture-of-Experts with Adaptive Retrieval Re-ranking for Pathology Report Generation Yixin Chen et.al. 2603.04348 null
2026-03-04 Underrepresented in Foundation Model Pretraining Data? A One-Shot Probe Chris Vorster et.al. 2603.04346 null
2026-03-04 Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection Dacheng Qi et.al. 2603.04337 null
2026-03-04 Scalable Evaluation of the Realism of Synthetic Environmental Augmentations in Images Damian J. Ruck et.al. 2603.04325 null
2026-03-04 CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential Recommendation Jinfeng Xu et.al. 2603.04320 null
2026-03-04 World Properties without World Models: Recovering Spatial and Temporal Structure from Co-occurrence Statistics in Static Word Embeddings Elan Barenholtz et.al. 2603.04317 null
2026-03-04 Theory Discovery in Social Networks: Automating ERGM Specification with Large Language Models Yidan Sun et.al. 2603.04306 null
2026-03-04 The Company You Keep: How LLMs Respond to Dark Triad Traits Zeyi Lu et.al. 2603.04299 null
2026-03-04 LabelBuddy: An Open Source Music and Audio Language Annotation Tagging Tool Using AI Assistance Ioannis Prokopiou et.al. 2603.04293 null
2026-03-04 Position: Vector Prompt Interfaces Should Be Exposed to Enable Customization of Large Language Models Liangwei Yang et.al. 2603.04292 null
2026-03-04 Causality Elicitation from Large Language Models Takashi Kameyama et.al. 2603.04276 null
2026-03-04 SSR: A Generic Framework for Text-Aided Map Compression for Localization Mohammad Omama et.al. 2603.04272 null
2026-03-04 When AI Fails, What Works? A Data-Driven Taxonomy of Real-World AI Risk Mitigation Strategies Evgenija Popchanovska et.al. 2603.04259 null
2026-03-04 Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory Zhenting Wang et.al. 2603.04257 null
2026-03-04 FeedAIde: Guiding App Users to Submit Rich Feedback Reports by Asking Context-Aware Follow-Up Questions Ali Ebrahimi Pourasad et.al. 2603.04244 null
2026-03-03 Utonia: Toward One Encoder for All Point Clouds Yujia Zhang et.al. 2603.03283 null
2026-03-03 MIBURI: Towards Expressive Interactive Gesture Synthesis M. Hamza Mughal et.al. 2603.03282 null
2026-03-03 Tether: Autonomous Functional Play with Correspondence-Driven Trajectory Warping William Liang et.al. 2603.03278 null
2026-03-03 Beyond Language Modeling: An Exploration of Multimodal Pretraining Shengbang Tong et.al. 2603.03276 null
2026-03-03 Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals Achyutha Menon et.al. 2603.03258 null
2026-03-03 Using Learning Progressions to Guide AI Feedback for Science Learning Xin Xia et.al. 2603.03249 null
2026-03-03 Density-Guided Response Optimization: Community-Grounded Alignment via Implicit Acceptance Signals Patrick Gerard et.al. 2603.03242 null
2026-03-03 UniG2U-Bench: Do Unified Models Advance Multimodal Understanding? Zimo Wen et.al. 2603.03241 null
2026-03-03 Conversational Learning Diagnosis via Reasoning Multi-Turn Interactive Learning Fangzhou Yao et.al. 2603.03236 null
2026-03-03 AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework Zihang Zeng et.al. 2603.03233 null
2026-03-03 Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use Aradhye Agarwal et.al. 2603.03205 null
2026-03-03 No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models Omer Sela et.al. 2603.03203 null
2026-03-03 Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration? Dadi Guo et.al. 2603.03202 null
2026-03-03 ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments Ziyang Gong et.al. 2603.03198 null
2026-03-03 MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization Ashutosh Chaubey et.al. 2603.03192 null
2026-03-03 Type-Aware Retrieval-Augmented Generation with Dependency Closure for Solver-Executable Industrial Optimization Modeling Y. Zhong et.al. 2603.03180 null
2026-03-03 Saarthi for AGI: Towards Domain-Specific General Intelligence for Formal Verification Aman Kumar et.al. 2603.03175 null
2026-03-03 From Language to Action: Can LLM-Based Agents Be Used for Embodied Robot Cognition? Shinas Shaji et.al. 2603.03148 null
2026-03-03 Agentic AI-based Coverage Closure for Formal Verification Sivaram Pothireddypalli et.al. 2603.03147 null
2026-03-03 APRES: An Agentic Paper Revision and Evaluation System Bingchen Zhao et.al. 2603.03142 null
2026-03-02 Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Training Valentin Lacombe et.al. 2603.02208 null
2026-03-02 Frontier Models Can Take Actions at Low Probabilities Alex Serrano et.al. 2603.02202 null
2026-03-02 Symbol-Equivariant Recurrent Reasoning Models Richard Freinschlag et.al. 2603.02193 null
2026-03-02 Multi-Head Low-Rank Attention Songtao Liu et.al. 2603.02188 null
2026-03-02 How Small Can 6G Reason? Scaling Tiny Language Models for AI-Native Networks Mohamed Amine Ferrag et.al. 2603.02156 null
2026-03-02 Zero- and Few-Shot Named-Entity Recognition: Case Study and Dataset in the Crime Domain (CrimeNER) Miguel Lopez-Duran et.al. 2603.02150 null
2026-03-02 LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards Guanzheng Chen et.al. 2603.02146 null
2026-03-02 OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens Yiying Yang et.al. 2603.02138 null
2026-03-02 LLMs as Strategic Actors: Behavioral Alignment, Risk Calibration, and Argumentation Framing in Geopolitical Simulations Veronika Solopova et.al. 2603.02128 null
2026-03-02 Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy Jiahao Huang et.al. 2603.02123 null
2026-03-02 Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning Justin Waugh et.al. 2603.02119 null
2026-03-02 Recursive Models for Long-Horizon Reasoning Chenxiao Yang et.al. 2603.02112 null
2026-03-02 Recursive Think-Answer Process for LLMs and VLMs Byung-Kwan Lee et.al. 2603.02099 null
2026-03-02 OmniRet: Efficient and High-Fidelity Omni Modality Retrieval Chuong Huynh et.al. 2603.02098 null
2026-03-02 ClinConsensus: A Consensus-Based Benchmark for Evaluating Chinese Medical LLMs across Difficulty Levels Xiang Zheng et.al. 2603.02097 null
2026-03-02 Adam Converges Without Any Modification On Update Rules Yushun Zhang et.al. 2603.02092 null
2026-03-02 Learning from Synthetic Data Improves Multi-hop Reasoning Anmol Kabra et.al. 2603.02091 null
2026-03-02 What Exactly do Children Receive in Language Acquisition? A Case Study on CHILDES with Automated Detection of Filler-Gap Dependencies Zhenghao Herbert Zhou et.al. 2603.02082 null
2026-03-02 GenDB: The Next Generation of Query Processing – Synthesized, Not Engineered Jiale Lao et.al. 2603.02081 null
2026-03-02 Trident: Adaptive Scheduling for Heterogeneous Multimodal Data Pipelines Ding Pan et.al. 2603.02075 null
2026-02-27 DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science Fan Shu et.al. 2602.24288 null
2026-02-27 Do LLMs Benefit From Their Own Words? Jenny Y. Huang et.al. 2602.24287 null
2026-02-27 CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation Weinan Dai et.al. 2602.24286 null
2026-02-27 Taming Momentum: Rethinking Optimizer States Through Low-Rank Approximation Zhengbo Wang et.al. 2602.24283 null
2026-02-27 Memory Caching: RNNs with Growing Memory Ali Behrouz et.al. 2602.24281 null
2026-02-27 UXSim: Towards a Hybrid User Search Simulation Saber Zerhoudi et.al. 2602.24241 null
2026-02-27 SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems Jialiang Fan et.al. 2602.24235 null
2026-02-27 Anansi: Scalable Characterization of Message-Based Job Scams Abisheka Pitumpe et.al. 2602.24223 null
2026-02-27 Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume Gregory Kang Ruey Lau et.al. 2602.24195 null
2026-02-27 MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games Jacob Eisenstein et.al. 2602.24188 null
2026-02-27 Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions Saleh Afroogh et.al. 2602.24176 null
2026-02-27 Task-Centric Acceleration of Small-Language Models Dor Tsur et.al. 2602.24174 null
2026-02-27 LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Antoine Peyronnet et.al. 2602.24173 null
2026-02-27 ArgLLM-App: An Interactive System for Argumentative Reasoning with Large Language Models Adam Dejl et.al. 2602.24172 null
2026-02-27 Terminology Rarity Predicts Catastrophic Failure in LLM Translation of Low-Resource Ancient Languages: Evidence from Ancient Greek James L. Zainaldin et.al. 2602.24119 null
2026-02-27 Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification Vikash Singh et.al. 2602.24111 null
2026-02-27 ARGUS: Seeing the Influence of Narrative Features on Persuasion in Argumentative Texts Sara Nabhani et.al. 2602.24109 null
2026-02-27 The Subjectivity of Monoculture Nathanael Jo et.al. 2602.24086 null
2026-02-27 Preference Packing: Efficient Preference Optimization for Large Language Models Jaekyung Cho et.al. 2602.24082 null
2026-02-27 A Novel Hierarchical Multi-Agent System for Payments Using LLMs Joon Kiat Chua et.al. 2602.24068 null
2026-02-26 MediX-R1: Open Ended Medical Reinforcement Learning Sahal Shaji Mullappilly et.al. 2602.23363 null
2026-02-26 SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport Simon Roschmann et.al. 2602.23353 null
2026-02-26 Scale Can’t Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning Amita Kamath et.al. 2602.23351 null
2026-02-26 Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation? Tilemachos Aravanis et.al. 2602.23339 null
2026-02-26 Utilizing LLMs for Industrial Process Automation Salim Fares et.al. 2602.23331 null
2026-02-26 Toward Expert Investment Teams:A Multi-Agent LLM System with Fine-Grained Trading Tasks Kunihiro Miyazaki et.al. 2602.23330 null
2026-02-26 LLM Novice Uplift on Dual-Use, In Silico Biology Tasks Chen Bo Calvin Zhang et.al. 2602.23329 null
2026-02-26 Evaluating Zero-Shot and One-Shot Adaptation of Small Language Models in Leader-Follower Interaction Rafael R. Baptista et.al. 2602.23312 null
2026-02-26 ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding Yiran Guan et.al. 2602.23306 null
2026-02-26 A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversations Soumya Dutta et.al. 2602.23300 null
2026-02-26 PGVMS: A Prompt-Guided Unified Framework for Virtual Multiplex IHC Staining with Pathological Semantic Learning Fuqiang Chen et.al. 2602.23292 null
2026-02-26 CXReasonAgent: Evidence-Grounded Diagnostic Reasoning Agent for Chest X-rays Hyungyung Lee et.al. 2602.23276 null
2026-02-26 Mitigating Legibility Tax with Decoupled Prover-Verifier Games Yegon Kim et.al. 2602.23248 null
2026-02-26 Agency and Architectural Limits: Why Optimization-Based Systems Cannot Be Norm-Responsive Radha Sarma et.al. 2602.23239 null
2026-02-26 Large Multimodal Models as General In-Context Classifiers Marco Garosi et.al. 2602.23229 null
2026-02-26 MovieTeller: Tool-augmented Movie Synopsis with ID Consistent Progressive Abstraction Yizhi Li et.al. 2602.23228 null
2026-02-26 Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding? Pengxiang Li et.al. 2602.23225 null
2026-02-26 STELLAR: Storage Tuning Engine Leveraging LLM Autonomous Reasoning for High Performance Parallel File Systems Chris Egersdoerfer et.al. 2602.23220 null
2026-02-26 Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language Max S. Bennett et.al. 2602.23201 null
2026-02-26 InnerQ: Hardware-aware Tuning-free Quantization of KV Cache for Large Language Models Sayed Mohammadreza Tayaranian Hosseini et.al. 2602.23200 null
2026-02-25 Recovered in Translation: Efficient Pipeline for Automated Translation of Benchmarks and Datasets Hanna Yukhymenko et.al. 2602.22207 null
2026-02-25 SumTablets: A Transliteration Dataset of Sumerian Tablets Cole Simmons et.al. 2602.22200 null
2026-02-25 Improving Parametric Knowledge Access in Reasoning Language Models Melody Ma et.al. 2602.22193 null
2026-02-25 DySCO: Dynamic Attention-Scaling Decoding for Long-Context LMs Xi Ye et.al. 2602.22175 null
2026-02-25 A Taxonomy of Human–MLLM Interaction in Early-Stage Sketch-Based Design Ideation Weiayn Shi et.al. 2602.22171 null
2026-02-25 LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models Minqiu Sun et.al. 2602.22158 null
2026-02-25 Dynamic Personality Adaptation in Large Language Models via State Machines Leon Pielage et.al. 2602.22157 null
2026-02-25 Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual Yining Li et.al. 2602.22146 null
2026-02-25 When AI Writes, Whose Voice Remains? Quantifying Cultural Marker Erasure Across World English Varieties in Large Language Models Satyam Kumar Navneet et.al. 2602.22145 null
2026-02-25 NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors Lingfeng Ren et.al. 2602.22144 null
2026-02-25 WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs Yulin Zhang et.al. 2602.22142 null
2026-02-25 SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents Patrick Tser Jern Kon et.al. 2602.22124 null
2026-02-25 GeoDiv: Framework For Measuring Geographical Diversity In Text-To-Image Models Abhipsa Basu et.al. 2602.22120 null
2026-02-25 Brain3D: Brain Report Automation via Inflated Vision Transformers in 3D Mariano Barone et.al. 2602.22098 null
2026-02-25 Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference Bo-Wei Chen et.al. 2602.22090 null
2026-02-25 ViSTAR: Virtual Skill Training with Augmented Reality with 3D Avatars and LLM coaching agent Chunggi Lee et.al. 2602.22077 null
2026-02-25 Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models Christian Nickel et.al. 2602.22072 null
2026-02-25 Language Models Exhibit Inconsistent Biases Towards Algorithmic Agents and Human Experts Jessica Y. Bo et.al. 2602.22070 null
2026-02-25 NESTOR: A Nested MOE-based Neural Operator for Large-Scale PDE Pre-Training Dengdi Sun et.al. 2602.22059 null
2026-02-25 RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking Yanqiu Yu et.al. 2602.22033 null
2026-02-24 On Data Engineering for Scaling LLM Terminal Capabilities Renjie Pi et.al. 2602.21193 null
2026-02-24 Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training Anas Barakat et.al. 2602.21189 null
2026-02-24 Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning Haoyi Jiang et.al. 2602.21186 null
2026-02-24 The Diffusion Duality, Chapter II: $Ψ$ -Samplers and Efficient Curriculum Justin Deschenaux et.al. 2602.21185 null
2026-02-24 Seeing Through Words: Controlling Visual Retrieval Quality with Language Models Jianglin Lu et.al. 2602.21175 null
2026-02-24 ActionReasoning: Robot Action Reasoning in 3D Space with LLM for Robotic Brick Stacking Guangming Wang et.al. 2602.21161 null
2026-02-24 SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards Dengjia Zhang et.al. 2602.21158 null
2026-02-24 HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning Quanxin Shou et.al. 2602.21157 null
2026-02-24 Scaling State-Space Models on Multiple GPUs with Tensor Parallelism Anurag Dutt et.al. 2602.21144 null
2026-02-24 A Benchmark for Deep Information Synthesis Debjit Paul et.al. 2602.21143 null
2026-02-24 LUMEN: Longitudinal Multi-Modal Radiology Model for Prognosis and Diagnosis Zhifan Jiang et.al. 2602.21142 null
2026-02-24 UDVideoQA: A Traffic Video Question Answering Dataset for Multi-Object Spatio-Temporal Reasoning in Urban Dynamics Joseph Raj Vishal et.al. 2602.21137 null
2026-02-24 SparkMe: Adaptive Semi-Structured Interviewing for Qualitative Insight Discovery David Anugraha et.al. 2602.21136 null
2026-02-24 “Are You Sure?”: An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems Xinfeng Li et.al. 2602.21127 null
2026-02-24 Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning Sanket Badhe et.al. 2602.21103 null
2026-02-24 Turning Semantics into Topology: LLM-Driven Attribute Augmentation for Collaborative Filtering Junjie Meng et.al. 2602.21099 null
2026-02-24 Can Interest-Bearing Positions Solve the Long-Horizon Problem in Prediction Markets? Caleb Maresca et.al. 2602.21091 null
2026-02-24 Beyond the Star Rating: A Scalable Framework for Aspect-Based Sentiment Analysis Using LLMs and Text Classification Vishal Patil et.al. 2602.21082 null
2026-02-24 Scaling Vision Transformers: Evaluating DeepSpeed for Image-Centric Workloads Huy Trinh et.al. 2602.21081 null
2026-02-24 An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems Anna Martin-Boyle et.al. 2602.21059 null
2026-02-23 Do Large Language Models Understand Data Visualization Rules? Martin Sinnona et.al. 2602.20137 null
2026-02-23 KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration Mohammad Amanlou et.al. 2602.20135 null
2026-02-23 AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization Mert Cemri et.al. 2602.20133 null
2026-02-23 To Reason or Not to: Selective Chain-of-Thought in Medical Question Answering Zaifu Zhan et.al. 2602.20130 null
2026-02-23 Adaptation to Intrinsic Dependence in Diffusion Language Models Yunxiao Zhao et.al. 2602.20126 null
2026-02-23 NanoKnow: How to Know What Your Language Model Knows Lingwei Gu et.al. 2602.20122 null
2026-02-23 NovaPlan: Zero-Shot Long-Horizon Manipulation via Closed-Loop Video Language Planning Jiahui Fu et.al. 2602.20119 null
2026-02-23 ReSyn: Autonomously Scaling Synthetic Environments for Reasoning Models Andre He et.al. 2602.20117 null
2026-02-23 BarrierSteer: LLM Safety via Learning Barrier Steering Thanh Q. Tran et.al. 2602.20102 null
2026-02-23 CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching Yuzhe Wang et.al. 2602.20094 null
2026-02-23 BabyLM Turns 4: Call for Papers for the 2026 BabyLM Workshop Leshem Choshen et.al. 2602.20092 null
2026-02-23 How Retrieved Context Shapes Internal Representations in RAG Samuel Yeh et.al. 2602.20091 null
2026-02-23 StructXLIP: Enhancing Vision-language Models with Multimodal Structural Cues Zanxi Ruan et.al. 2602.20089 null
2026-02-23 Do Large Language Models Understand Data Visualization Principles? Martin Sinnona et.al. 2602.20084 null
2026-02-23 HeatPrompt: Zero-Shot Vision-Language Modeling of Urban Heat Demand from Satellite Images Kundan Thota et.al. 2602.20066 null
2026-02-23 Multilingual Large Language Models do not comprehend all natural languages to equal degrees Natalia Moskvina et.al. 2602.20065 null
2026-02-23 The LLMbda Calculus: AI Agents, Conversations, and Information Flow Zac Garby et.al. 2602.20064 null
2026-02-23 Can You Tell It’s AI? Human Perception of Synthetic Voices in Vishing Scenarios Zoha Hayat Bhatti et.al. 2602.20061 null
2026-02-23 Entropy in Large Language Models Marco Scharringhausen et.al. 2602.20052 null
2026-02-23 Position: General Alignment Has Hit a Ceiling; Edge Alignment Must Be Taken Seriously Han Bao et.al. 2602.20042 null
2026-02-20 Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory Vatsal Agarwal et.al. 2602.18434 null
2026-02-20 VIRAASAT: Traversing Novel Paths for Indian Cultural Reasoning Harshul Raj Surana et.al. 2602.18429 null
2026-02-20 CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation Xia Su et.al. 2602.18424 null
2026-02-20 SPQ: An Ensemble Technique for Large Language Model Compression Jiamin Yao et.al. 2602.18420 null
2026-02-20 AI-Wrapped: Participatory, Privacy-Preserving Measurement of Longitudinal LLM Use In-the-Wild Cathy Mengying Fang et.al. 2602.18415 null
2026-02-20 Zero-shot Interactive Perception Venkatesh Sripada et.al. 2602.18374 null
2026-02-20 “How Do I …?”: Procedural Questions Predominate Student-LLM Chatbot Conversations Alexandra Neagu et.al. 2602.18372 null
2026-02-20 Quantum Maximum Likelihood Prediction via Hilbert Space Embeddings Sreejith Sreekumar et.al. 2602.18364 null
2026-02-20 Qualitative Coding Analysis through Open-Source Large Language Models: A User Study and Design Recommendations Tung T. Ngo et.al. 2602.18352 null
2026-02-20 Validating Political Position Predictions of Arguments Jordan Robinson et.al. 2602.18351 null
2026-02-20 Vichara: Appellate Judgment Prediction and Explanation for the Indian Judicial System Pavithra PM Nair et.al. 2602.18346 null
2026-02-20 On the “Induction Bias” in Sequence Models M. Reza Ebrahimi et.al. 2602.18333 null
2026-02-20 VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean Yutong Xin et.al. 2602.18307 null
2026-02-20 On the Semantic and Syntactic Information Encoded in Proto-Tokens for One-Step Text Reconstruction Ivan Bondarenko et.al. 2602.18301 null
2026-02-20 Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory Usman Anwar et.al. 2602.18297 null
2026-02-20 Context-Aware Mapping of 2D Drawing Annotations to 3D CAD Features Using LLM-Assisted Reasoning for Manufacturing Automation Muhammad Tayyab Khana et.al. 2602.18296 null
2026-02-20 Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplers Xiaotong Ji et.al. 2602.18292 null
2026-02-20 Simplifying Outcomes of Language Model Component Analyses with ELIA Aaron Louis Eidt et.al. 2602.18262 null
2026-02-20 Dual-Tree LLM-Enhanced Negative Sampling for Implicit Collaborative Filtering Jiayi Wu et.al. 2602.18249 null
2026-02-20 Thinking by Subtraction: Confidence-Driven Contrastive Decoding for LLM Reasoning Lexiang Tang et.al. 2602.18232 null
2026-02-19 Sink-Aware Pruning for Diffusion Language Models Aidar Myrzakhan et.al. 2602.17664 null
2026-02-19 What Language is This? Ask Your Tokenizer Clara Meister et.al. 2602.17655 null
2026-02-19 Differences in Typological Alignment in Language Models’ Treatment of Differential Argument Marking Iskar Deng et.al. 2602.17653 null
2026-02-19 Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting Xiaohan Zhao et.al. 2602.17645 null
2026-02-19 When to Trust the Cheap Check: Weak and Strong Verification for Reasoning Shayan Kiyani et.al. 2602.17633 null
2026-02-19 Catastrophic Forgetting Resilient One-Shot Incremental Federated Learning Obaidullah Zaland et.al. 2602.17625 null
2026-02-19 Unmasking the Factual-Conceptual Gap in Persian Language Models Alireza Sakhaeirad et.al. 2602.17623 null
2026-02-19 Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs Luke Huang et.al. 2602.17616 null
2026-02-19 Towards Anytime-Valid Statistical Watermarking Baihe Huang et.al. 2602.17608 null
2026-02-19 AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games Lance Ying et.al. 2602.17594 null
2026-02-19 Modeling Distinct Human Interaction in Web Agents Faria Huq et.al. 2602.17588 null
2026-02-19 ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment Hongjue Zhao et.al. 2602.17560 null
2026-02-19 RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward Qiucheng Wu et.al. 2602.17558 null
2026-02-19 GraphThinker: Reinforcing Video Reasoning with Event Graph Thinking Zixu Cheng et.al. 2602.17555 null
2026-02-19 A Theoretical Framework for Modular Learning of Robust Generative Models Corinna Cortes et.al. 2602.17554 null
2026-02-19 MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning Xiaoliang Fu et.al. 2602.17550 null
2026-02-19 Learning to Stay Safe: Adaptive Regularization Against Safety Degradation during Fine-Tuning Jyotin Goel et.al. 2602.17546 null
2026-02-19 Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability Shashank Aggarwal et.al. 2602.17544 null
2026-02-19 Using LLMs for Knowledge Component-level Correctness Labeling in Open-ended Coding Problems Zhangqi Duan et.al. 2602.17542 null
2026-02-19 LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs Behzad Bozorgtabar et.al. 2602.17535 null
2026-02-18 Reinforced Fast Weights with Next-Sequence Prediction Hee Seung Hwang et.al. 2602.16704 null
2026-02-18 Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology Shen Zhou Hong et.al. 2602.16703 null
2026-02-18 Saliency-Aware Multi-Route Thinking: Revisiting Vision-Language Reasoning Mingjia Shi et.al. 2602.16702 null
2026-02-18 Causality is Key for Interpretability Claims to Generalise Shruti Joshi et.al. 2602.16698 null
2026-02-18 Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens Potsawee Manakul et.al. 2602.16687 null
2026-02-18 SPARC: Scenario Planning and Reasoning for Automated C Unit Test Generation Jaid Monwar Chowdhury et.al. 2602.16671 null
2026-02-18 Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment Yuyan Bu et.al. 2602.16660 null
2026-02-18 Agent Skill Framework: Perspectives on the Potential of Small Language Models in Industrial Environments Yangjie Xu et.al. 2602.16653 null
2026-02-18 Retrieval Augmented Generation of Literature-derived Polymer Knowledge: The Example of a Biodegradable Polymer Expert System Sonakshi Gupta et.al. 2602.16650 null
2026-02-18 Quecto-V1: Empirical Analysis of 8-bit Quantized Small Language Models for On-Device Legal Retrieval Subrit Dikshit et.al. 2602.16640 null
2026-02-18 AREG: Adversarial Resource Extraction Game for Evaluating Persuasion and Resistance in Large Language Models Adib Sakhawat et.al. 2602.16639 null
2026-02-18 Who can we trust? LLM-as-a-jury for Comparative Assessment Mengjie Qian et.al. 2602.16610 null
2026-02-18 CitiLink-Summ: Summarization of Discussion Subjects in European Portuguese Municipal Meeting Minutes Miguel Marques et.al. 2602.16607 null
2026-02-18 FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving Chia-chi Hsieh et.al. 2602.16603 null
2026-02-18 A Contrastive Learning Framework Empowered by Attention-based Feature Adaptation for Street-View Image Classification Qi You et.al. 2602.16590 null
2026-02-18 Why Thinking Hurts? Diagnosing and Rectifying the Reasoning Shift in Foundation Recommender Models Luankang Zhang et.al. 2602.16587 null
2026-02-18 Creating a digital poet Vered Tohar et.al. 2602.16578 null
2026-02-18 Automated Extraction of Mechanical Constitutive Models from Scientific Literature using Large Language Models: Applications in Cultural Heritage Conservation Rui Hu et.al. 2602.16551 null
2026-02-18 Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents Doron Shavit et.al. 2602.16520 null
2026-02-18 Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification Taja Kuzman Pungeršek et.al. 2602.16516 null
2026-02-17 Operationalising the Superficial Alignment Hypothesis via Task Complexity Tomás Vergara-Browne et.al. 2602.15829 null
2026-02-17 CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM Editing Zarif Ikram et.al. 2602.15823 null
2026-02-17 VideoSketcher: Video Models Prior Enable Versatile Sequential Sketch Generation Hui Ren et.al. 2602.15819 null
2026-02-17 FAST-EQA: Efficient Embodied Question Answering with Global and Local Region Relevancy Haochen Zhang et.al. 2602.15813 null
2026-02-17 Decision Quality Evaluation Framework at Pinterest Yuqi Tian et.al. 2602.15809 null
2026-02-17 The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety Max Springer et.al. 2602.15799 null
2026-02-17 Enhancing Building Semantics Preservation in AI Model Training with Large Language Model Encodings Suhyung Jang et.al. 2602.15791 null
2026-02-17 This human study did not involve human subjects: Validating LLM simulations as behavioral evidence Jessica Hullman et.al. 2602.15785 null
2026-02-17 Neural Scaling Laws for Boosted Jet Tagging Matthias Vigl et.al. 2602.15781 null
2026-02-17 ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution Yahia Alqurnawi et.al. 2602.15769 null
2026-02-17 TAC: Timestamped Audio Captioning Sonal Kumar et.al. 2602.15766 null
2026-02-17 GLM-5: from Vibe Coding to Agentic Engineering GLM-5 Team et.al. 2602.15763 null
2026-02-17 A Differential Fuzzing-Based Evaluation of Functional Equivalence in LLM-Generated Code Refactorings Simantika Bhattacharjee Dristi et.al. 2602.15761 null
2026-02-17 ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models Manav Nitin Kapadnis et.al. 2602.15758 null
2026-02-17 Under-resourced studies of under-resourced languages: lemmatization and POS-tagging with LLM annotators for historical Armenian, Georgian, Greek and Syriac Chahan Vidal-Gorène et.al. 2602.15753 null
2026-02-17 Causal Effect Estimation with Latent Textual Treatments Omri Feldman et.al. 2602.15730 null
2026-02-17 Recursive Concept Evolution for Compositional Reasoning in Large Language Models Sarim Chaudhry et.al. 2602.15725 null
2026-02-17 Learning to Retrieve Navigable Candidates for Efficient Vision-and-Language Navigation Shutian Gu et.al. 2602.15724 null
2026-02-17 Rethinking Metrics for Lexical Semantic Change Detection Roksana Goworek et.al. 2602.15716 null
2026-02-17 Proactive Conversational Assistant for a Procedural Manual Task based on Audio and IMU Rehana Mahfuz et.al. 2602.15707 null
2026-02-16 Symmetry in language statistics shapes the geometry of model representations Dhruva Karkada et.al. 2602.15029 null
2026-02-16 Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization Shangding Gu et.al. 2602.15028 null
2026-02-16 Scaling Beyond Masked Diffusion Language Models Subham Sekhar Sahoo et.al. 2602.15014 null
2026-02-16 Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation Ruoxi Liu et.al. 2602.15013 null
2026-02-16 BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames Max Sobol Mark et.al. 2602.15010 null
2026-02-16 Learning User Interests via Reasoning and Distillation for Cross-Domain News Recommendation Mengdan Zhu et.al. 2602.15005 null
2026-02-16 ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery Ayush Shrivastava et.al. 2602.14989 null
2026-02-16 DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI En Yu et.al. 2602.14974 null
2026-02-16 Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance System Kawin Mayilvaghanan et.al. 2602.14970 null
2026-02-16 Activation-Space Uncertainty Quantification for Pretrained Networks Richard Bergna et.al. 2602.14934 null
2026-02-16 MAC-AMP: A Closed-Loop Multi-Agent Collaboration System for Multi-Objective Antimicrobial Peptide Design Gen Zhou et.al. 2602.14926 null
2026-02-16 The Potential of CoT for Reasoning: A Closer Look at Trace Dynamics Gregor Bachmann et.al. 2602.14903 null
2026-02-16 Algorithmic Simplification of Neural Networks with Mosaic-of-Motifs Pedram Bakhtiarifard et.al. 2602.14896 null
2026-02-16 Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution Matthew Kowal et.al. 2602.14869 null
2026-02-16 Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning Ilia Mahrooghi et.al. 2602.14868 null
2026-02-16 The Well-Tempered Classifier: Some Elementary Properties of Temperature Scaling Pierre-Alexandre Mattei et.al. 2602.14862 null
2026-02-16 World Models for Policy Refinement in StarCraft II Yixin Zhang et.al. 2602.14857 null
2026-02-16 RF-GPT: Teaching AI to See the Wireless World Hang Zou et.al. 2602.14833 null
2026-02-16 Testimole-Conversational: A 30-Billion-Word Italian Discussion Board Corpus (1996-2024) for Language Modeling and Sociolinguistic Research Matteo Rinaldi et.al. 2602.14819 null
2026-02-16 Learning State-Tracking from Code Using Linear RNNs Julien Siems et.al. 2602.14814 null
2026-02-13 Semantic Chunking and the Entropy of Natural Language Weishun Zhong et.al. 2602.13194 null
2026-02-13 Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control William Chen et.al. 2602.13193 null
2026-02-13 CoPE-VideoLM: Codec Primitives For Efficient Video Language Models Sayan Deb Sarkar et.al. 2602.13191 null
2026-02-13 Asynchronous Verified Semantic Caching for Tiered LLM Architectures Asmit Kumar Singh et.al. 2602.13165 null
2026-02-13 In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach Yiran Gao et.al. 2602.13156 null
2026-02-13 Peaceful Anarcho-Accelerationism: Decentralized Full Automation for a Society of Universal Care Eduardo C. Garrido-Merchán et.al. 2602.13154 null
2026-02-13 Quantization-Robust LLM Unlearning via Low-Rank Adaptation João Vitor Boer Abitante et.al. 2602.13151 null
2026-02-13 SCOPE: Selective Conformal Optimized Pairwise LLM Judging Sher Badshah et.al. 2602.13110 null
2026-02-13 Consistency of Large Reasoning Models Under Multi-Turn Attacks Yubo Li et.al. 2602.13093 null
2026-02-13 Exploring a New Competency Modeling Process with Large Language Models Silin Du et.al. 2602.13084 null
2026-02-13 Agentic AI for Robot Control: Flexible but still Fragile Oscar Lima et.al. 2602.13081 null
2026-02-13 LCSB: Layer-Cyclic Selective Backpropagation for Memory-Efficient On-Device LLM Fine-Tuning Juneyoung Park et.al. 2602.13073 null
2026-02-13 Bus-Conditioned Zero-Shot Trajectory Generation via Task Arithmetic Shuai Liu et.al. 2602.13071 null
2026-02-13 Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning Juneyoung Park et.al. 2602.13069 null
2026-02-13 GPTZero: Robust Detection of LLM-Generated Texts George Alexandru Adam et.al. 2602.13042 null
2026-02-13 Implicit-Scale 3D Reconstruction for Multi-Food Volume Estimation from Monocular Images Yuhao Chen et.al. 2602.13041 null
2026-02-13 Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL Yixiao Zhou et.al. 2602.13035 null
2026-02-13 Buy versus Build an LLM: A Decision Framework for Governments Jiahao Lu et.al. 2602.13033 null
2026-02-13 Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis Runzhou Liu et.al. 2602.13028 null
2026-02-13 Know More, Know Clearer: A Meta-Cognitive Framework for Knowledge Augmentation in Large Language Models Hao Chen et.al. 2602.12996 null
2026-02-12 Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment Jacky Kwok et.al. 2602.12281 null
2026-02-12 UniT: Unified Multimodal Chain-of-Thought Test-time Scaling Leon Liangyu Chen et.al. 2602.12279 null
2026-02-12 AttentionRetriever: Attention Layers are Secretly Long Document Retrievers David Jiahao Fu et.al. 2602.12278 null
2026-02-12 On-Policy Context Distillation for Language Models Tianzhu Ye et.al. 2602.12275 null
2026-02-12 T3D: Few-Step Diffusion Language Models via Trajectory Self-Distillation with Direct Discriminative Optimization Tunyu Zhang et.al. 2602.12262 null
2026-02-12 Think like a Scientist: Physics-guided LLM Agent for Equation Discovery Jianke Yang et.al. 2602.12259 null
2026-02-12 Automated Test Suite Enhancement Using Large Language Models with Few-shot Prompting Alex Chudic et.al. 2602.12256 null
2026-02-12 Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks Zhihong Liu et.al. 2602.12244 null
2026-02-12 Olmix: A Framework for Data Mixing Throughout LM Development Mayee F. Chen et.al. 2602.12237 null
2026-02-12 Detecting Overflow in Compressed Token Representations for Retrieval-Augmented Generation Julia Belikova et.al. 2602.12235 null
2026-02-12 VIRENA: Virtual Arena for Research, Education, and Democratic Innovation Emma Hoes et.al. 2602.12207 null
2026-02-12 ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images Mathieu Sibue et.al. 2602.12203 null
2026-02-12 Visual Reasoning Benchmark: Evaluating Multimodal LLMs on Classroom-Authentic Visual Problems from Primary Education Mohamed Huti et.al. 2602.12196 null
2026-02-12 Query-focused and Memory-aware Reranker for Long Context Processing Yuqing Li et.al. 2602.12192 null
2026-02-12 Unknown Attack Detection in IoT Networks using Large Language Models: A Robust, Data-efficient Approach Shan Ali et.al. 2602.12183 null
2026-02-12 How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics Yurong Chen et.al. 2602.12180 null
2026-02-12 Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation Bowei He et.al. 2602.12172 null
2026-02-12 Sci-CoE: Co-evolving Scientific Reasoning LLMs via Geometric Consensus with Sparse Supervision Xiaohan He et.al. 2602.12164 null
2026-02-12 3DGSNav: Enhancing Vision-Language Model Reasoning for Object Navigation via Active 3D Gaussian Splatting Wancai Zheng et.al. 2602.12159 null
2026-02-12 SafeNeuron: Neuron-Level Safety Alignment for Large Language Models Zhaoxin Wang et.al. 2602.12158 null
2026-02-11 Diffusion-Pretrained Dense and Contextual Embeddings Sedigheh Eslami et.al. 2602.11151 null
2026-02-11 Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning Dawid J. Kopiczko et.al. 2602.11149 null
2026-02-11 Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling Gongye Liu et.al. 2602.11146 null
2026-02-11 TabICLv2: A better, faster, scalable, and open tabular foundation model Jingang Qu et.al. 2602.11139 null
2026-02-11 Weight Decay Improves Language Model Plasticity Tessa Han et.al. 2602.11137 null
2026-02-11 Just on Time: Token-Level Early Stopping for Diffusion Language Models Zahar Kohut et.al. 2602.11133 null
2026-02-11 TEGRA: Text Encoding With Graph and Retrieval Augmentation for Misinformation Detection Géraud Faye et.al. 2602.11106 null
2026-02-11 Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away Soumya Suvra Ghosal et.al. 2602.11096 null
2026-02-11 Can Large Language Models Make Everyone Happy? Usman Naseem et.al. 2602.11091 null
2026-02-11 DataChef: Cooking Up Optimal Data Recipes for LLM Adaptation via Reinforcement Learning Yicheng Chen et.al. 2602.11089 null
2026-02-11 Vulnerabilities in Partial TEE-Shielded LLM Inference with Precomputed Noise Abhishek Saini et.al. 2602.11088 null
2026-02-11 SteuerLLM: Local specialized large language model for German tax law analysis Sebastian Wind et.al. 2602.11081 null
2026-02-11 In-the-Wild Model Organisms: Mitigating Undesirable Emergent Behaviors in Production LLM Post-Training via Data Attribution Frank Xiao et.al. 2602.11079 null
2026-02-11 Chatting with Images for Introspective Visual Thinking Junfei Wu et.al. 2602.11073 null
2026-02-11 Conversational Behavior Modeling Foundation Model With Multi-Level Perception Dingkun Zhou et.al. 2602.11065 null
2026-02-11 Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language Models Xinyu Yuan et.al. 2602.11057 null
2026-02-11 Embedding Inversion via Conditional Masked Diffusion Language Models Han Xiao et.al. 2602.11047 null
2026-02-11 Language Model Inversion through End-to-End Differentiation Kevin Yandoka Denamganaï et.al. 2602.11044 null
2026-02-11 Chain-of-Look Spatial Reasoning for Dense Surgical Instrument Counting Rishikesh Bhyri et.al. 2602.11024 null
2026-02-11 Information Abstraction for Data Transmission Networks based on Large Language Models Haoyuan Zhu et.al. 2602.11022 null
2026-02-10 Biases in the Blind Spot: Detecting What LLMs Fail to Mention Iván Arcuschin et.al. 2602.10117 null
2026-02-10 ST4VLA: Spatially Guided Training for Vision-Language-Action Models Jinhui Ye et.al. 2602.10109 null
2026-02-10 Quantum-Audit: Evaluating the Reasoning Limits of LLMs on Quantum Computing Mohamed Afane et.al. 2602.10092 null
2026-02-10 Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning Zhaoyang Wang et.al. 2602.10090 null
2026-02-10 CAPID: Context-Aware PII Detection for Question-Answering Systems Mariia Ponomarenko et.al. 2602.10074 null
2026-02-10 Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability Aaditya Vikram Prasad et.al. 2602.10067 null
2026-02-10 WildCat: Near-Linear Attention in Theory and Practice Tobias Schröder et.al. 2602.10056 null
2026-02-10 Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization Xinchen Han et.al. 2602.10048 null
2026-02-10 Fake-HR1: Rethinking reasoning of vision language model for synthetic image detection Changjiang Jiang et.al. 2602.10042 null
2026-02-10 Decoupled Reasoning with Implicit Fact Tokens (DRIFT): A Dual-Model Framework for Efficient Long-Context Inference Wenxuan Xie et.al. 2602.10021 null
2026-02-10 SCORE: Specificity, Context Utilization, Robustness, and Relevance for Reference-Free LLM Evaluation Homaira Huda Shomee et.al. 2602.10017 null
2026-02-10 Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design Bojian Hou et.al. 2602.10016 null
2026-02-10 A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Chenruo Liu et.al. 2602.10014 null
2026-02-10 Discovering High Level Patterns from Simulation Traces Sean Memery et.al. 2602.10009 null
2026-02-10 Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning Shijie Zhang et.al. 2602.10006 null
2026-02-10 ESTAR: Early-Stopping Token-Aware Reasoning For Efficient Inference Junda Wang et.al. 2602.10004 null
2026-02-10 A Unified Assessment of the Poverty of the Stimulus Argument for Neural Language Models Xiulin Yang et.al. 2602.09992 null
2026-02-10 RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation Hao Li et.al. 2602.09973 null
2026-02-10 Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning Zixuan Wang et.al. 2602.09972 null
2026-02-10 Trustworthy Agentic AI Requires Deterministic Architectural Boundaries Manish Bhattarai et.al. 2602.09947 null
2026-02-09 CIC-Trap4Phish: A Unified Multi-Format Dataset for Phishing and Quishing Attachment Detection Fatemeh Nejati et.al. 2602.09015 null
2026-02-09 ANCRe: Adaptive Neural Connection Reassignment for Efficient Depth Scaling Yilang Zhang et.al. 2602.09009 null
2026-02-09 From Obstacles to Etiquette: Robot Social Navigation with VLM-Informed Path Selection Zilin Fang et.al. 2602.09002 null
2026-02-09 DirMoE: Dirichlet-routed Mixture of Experts Amirhossein Vahidi et.al. 2602.09001 null
2026-02-09 iGRPO: Self-Feedback-Driven LLM Reasoning Ali Hatamizadeh et.al. 2602.09000 null
2026-02-09 Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models Yuliang Liu et.al. 2602.08984 null
2026-02-09 A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents Raghu Arghal et.al. 2602.08964 null
2026-02-09 CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute Chen Jin et.al. 2602.08948 null
2026-02-09 Automatic In-Domain Exemplar Construction and LLM-Based Refinement of Multi-LLM Expansions for Query Expansion Minghan Li et.al. 2602.08917 null
2026-02-09 Efficient and Stable Reinforcement Learning for Diffusion Language Models Jiawei Liu et.al. 2602.08905 null
2026-02-09 GSS: Gated Subspace Steering for Selective Memorization Mitigation in LLMs Xuanqi Zhang et.al. 2602.08901 null
2026-02-09 OmniReview: A Large-scale Benchmark and LLM-enhanced Framework for Realistic Reviewer Recommendation Yehua Huang et.al. 2602.08896 null
2026-02-09 AI-based Verbal and Visual Scaffolding in a Serious Game: Effects on Learning and Cognitive Load Caroline Wermann et.al. 2602.08893 null
2026-02-09 Scalable Delphi: Large Language Models for Structured Risk Estimation Tobias Lorenz et.al. 2602.08889 null
2026-02-09 DeepQuali: Initial results of a study on the use of large language models for assessing the quality of user stories Adam Trendowicz et.al. 2602.08887 null
2026-02-09 Is Reasoning Capability Enough for Safety in Long-Context Language Models? Yu Fu et.al. 2602.08874 null
2026-02-09 Whose Name Comes Up? Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation Lisette Espin-Noboa et.al. 2602.08873 null
2026-02-09 Large Language Models for Geolocation Extraction in Humanitarian Crisis Response G. Cafferata et.al. 2602.08872 null
2026-02-09 AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection Junru Zhang et.al. 2602.08868 null
2026-02-09 ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS Bang Xie et.al. 2602.08866 null
2026-02-06 MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images Ankan Deria et.al. 2602.06965 null
2026-02-06 InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning Yuchen Yan et.al. 2602.06960 null
2026-02-06 DAWN: Dependency-Aware Fast Inference for Diffusion LLMs Lizhuo Luo et.al. 2602.06953 null
2026-02-06 Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay Duygu Altinok et.al. 2602.06942 null
2026-02-06 Endogenous Resistance to Activation Steering in Language Models Alex McKenzie et.al. 2602.06941 null
2026-02-06 Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs Samir Abdaljalil et.al. 2602.06920 null
2026-02-06 Directing Space: Rehearsing Architecture as Performer with Explainable AI Pavlos Panagiotidis et.al. 2602.06915 null
2026-02-06 Seeing Beyond Redundancy: Task Complexity’s Role in Vision Token Specialization in VLLMs Darryl Hannan et.al. 2602.06914 null
2026-02-06 TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering Saad Hossain et.al. 2602.06911 null
2026-02-06 Plato’s Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations Chen Chen et.al. 2602.06887 null
2026-02-06 Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization Zitao Song et.al. 2602.06880 null
2026-02-06 NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile Devices Ruchika Chavhan et.al. 2602.06879 null
2026-02-06 TraceCoder: A Trace-Driven Multi-Agent Framework for Automated Debugging of LLM-Generated Code Jiangping Huang et.al. 2602.06875 null
2026-02-06 Uncovering Cross-Objective Interference in Multi-Objective Alignment Yining Lu et.al. 2602.06869 null
2026-02-06 Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing Meng Lou et.al. 2602.06862 null
2026-02-06 AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents Alisia Lupidi et.al. 2602.06855 null
2026-02-06 SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks Mingqian Feng et.al. 2602.06854 null
2026-02-06 The Quantum Sieve Tracer: A Hybrid Framework for Layer-Wise Activation Tracing in Large Language Models Jonathan Pan et.al. 2602.06852 null
2026-02-06 Improved Sampling Schedules for Discrete Diffusion Models Alberto Foresti et.al. 2602.06849 null
2026-02-06 The Representational Geometry of Number Zhimin Hu et.al. 2602.06843 null
2026-02-05 Predicting Camera Pose from Perspective Descriptions for Spatial Reasoning Xuejun Zhang et.al. 2602.06041 null
2026-02-05 SwimBird: Eliciting Switchable Reasoning Mode in Hybrid Autoregressive MLLMs Jintao Tong et.al. 2602.06040 null
2026-02-05 DyTopo: Dynamic Topology Routing for Multi-Agent Reasoning via Semantic Matching Yuxing Lu et.al. 2602.06039 null
2026-02-05 Thinking with Geometry: Active Geometry Integration for Spatial Reasoning Haoyuan Li et.al. 2602.06037 null
2026-02-05 DFlash: Block Diffusion for Flash Speculative Decoding Jian Chen et.al. 2602.06036 null
2026-02-05 V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval Dongyang Chen et.al. 2602.06034 null
2026-02-05 Can vision language models learn intuitive physics from interaction? Luca M. Schulze Buschoff et.al. 2602.06033 null
2026-02-05 AP-OOD: Attention Pooling for Out-of-Distribution Detection Claus Hofmann et.al. 2602.06031 null
2026-02-05 PhysicsAgentABM: Physics-Guided Generative Agent-Based Modeling Kavana Venkatesh et.al. 2602.06030 null
2026-02-05 Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory Haozhen Zhang et.al. 2602.06025 null
2026-02-05 Correctness-Optimized Residual Activation Lens (CORAL): Transferrable and Calibration-Aware Inference-Time Steering Miranda Muqing Miao et.al. 2602.06022 null
2026-02-05 Multi-Token Prediction via Self-Distillation John Kirchenbauer et.al. 2602.06019 null
2026-02-05 A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies Panagiotis Kaliosis et.al. 2602.06015 null
2026-02-05 GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks? Ruihang Li et.al. 2602.06013 null
2026-02-05 AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions Xianyang Liu et.al. 2602.06008 null
2026-02-05 VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation Jie Deng et.al. 2602.05998 null
2026-02-05 DSB: Dynamic Sliding Block Scheduling for Diffusion LLMs Lizhuo Luo et.al. 2602.05992 null
2026-02-05 Layer-wise LoRA fine-tuning: a similarity metric approach Keith Ando Ogawa et.al. 2602.05988 null
2026-02-05 From Human-Human Collaboration to Human-Agent Collaboration: A Vision, Design Philosophy, and an Empirical Framework for Achieving Successful Partnerships Between Humans and LLM Agents Bingsheng Yao et.al. 2602.05987 null
2026-02-05 Inverse Depth Scaling From Most Layers Being Similar Yizhou Liu et.al. 2602.05970 null
2026-02-04 Reinforced Attention Learning Bangzheng Li et.al. 2602.04884 null
2026-02-04 Protein Autoregressive Modeling via Multiscale Structure Generation Yanru Qu et.al. 2602.04883 null
2026-02-04 Rethinking the Trust Region in LLM Reinforcement Learning Penghui Qi et.al. 2602.04879 null
2026-02-04 Multi-layer Cross-Attention is Provably Optimal for Multi-modal In-context Learning Nicholas Barnfield et.al. 2602.04872 null
2026-02-04 Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism Chenwei Cui et.al. 2602.04870 null
2026-02-04 When LLaVA Meets Objects: Token Composition for Vision-Language-Models Soumya Jahagirdar et.al. 2602.04864 null
2026-02-04 Subliminal Effects in Your Data: A General Mechanism via Log-Linearity Ishaq Aden-Ali et.al. 2602.04863 null
2026-02-04 CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation Zhao Tong et.al. 2602.04856 null
2026-02-04 Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say “I Don’t Know” Dhruv Madhwal et.al. 2602.04853 null
2026-02-04 El Agente Estructural: An Artificially Intelligent Molecular Editor Changhyeok Choi et.al. 2602.04849 null
2026-02-04 Fluid Representations in Reasoning Models Dmitrii Kharlapenko et.al. 2602.04843 null
2026-02-04 Horizon-LM: A RAM-Centric Architecture for LLM Training Zhengqing Yuan et.al. 2602.04816 null
2026-02-04 Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents Shubham Vatsal et.al. 2602.04813 null
2026-02-04 OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models Yue Ding et.al. 2602.04804 null
2026-02-04 VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? Qing’an Liu et.al. 2602.04802 null
2026-02-04 LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues Amir Ivry et.al. 2602.04796 null
2026-02-04 Team, Then Trim: An Assembly-Line LLM Framework for High-Quality Tabular Data Generation Congjing Zhang et.al. 2602.04785 null
2026-02-04 NeuroCanvas: VLLM-Powered Robust Seizure Detection by Reformulating Multichannel EEG as Image Yan Chen et.al. 2602.04769 null
2026-02-04 Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation Luis Frentzen Salim et.al. 2602.04764 null
2026-02-04 When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond? Xinyu Zhou et.al. 2602.04755 null
2026-02-03 Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL Erfan Miahi et.al. 2602.03839 null
2026-02-03 Accelerating Scientific Research with Gemini: Case Studies and Common Techniques David P. Woodruff et.al. 2602.03837 null
2026-02-03 They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References Sahil Tripathi et.al. 2602.03822 null
2026-02-03 Fast-Slow Efficient Training for Multimodal Large Language Models via Visual Token Pruning Dingkun Zhang et.al. 2602.03815 null
2026-02-03 Conformal Thinking: Risk Control for Reasoning on a Compute Budget Xi Wang et.al. 2602.03814 null
2026-02-03 Antidistillation Fingerprinting Yixuan Even Xu et.al. 2602.03812 null
2026-02-03 Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation Ziru Chen et.al. 2602.03806 null
2026-02-03 Context Compression via Explicit Information Transmission Jiangnan Ye et.al. 2602.03784 null
2026-02-03 Efficient Estimation of Kernel Surrogate Models for Task Attribution Zhenshuo Zhang et.al. 2602.03783 null
2026-02-03 QVLA: Not All Channels Are Equal in Vision-Language-Action Model’s Quantization Yuhao Xu et.al. 2602.03782 null
2026-02-03 A Scene Graph Backed Approach to Open Set Semantic Mapping Martin Günther et.al. 2602.03781 null
2026-02-03 An Empirical Study of Collective Behaviors and Social Dynamics in Large Language Model Agents Farnoosh Hashemi et.al. 2602.03775 null
2026-02-03 Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL Ian Wu et.al. 2602.03773 null
2026-02-03 UniGeM: Unifying Data Mixing and Selection via Geometric Exploration and Mining Changhao Wang et.al. 2602.03772 null
2026-02-03 Reasoning with Latent Tokens in Diffusion Language Models Andre He et.al. 2602.03769 null
2026-02-03 Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives Owen Dong et.al. 2602.03750 null
2026-02-03 Edge-Optimized Vision-Language Models for Underground Infrastructure Assessment Johny J. Lopez et.al. 2602.03742 null
2026-02-03 RegionReasoner: Region-Grounded Multi-Round Visual Reasoning Wenfang Sun et.al. 2602.03733 null
2026-02-03 Training Multi-Turn Search Agent via Contrastive Dynamic Branch Sampling Yubao Zhao et.al. 2602.03719 null
2026-02-03 SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring Yisen Xu et.al. 2602.03712 null
2026-02-02 Reward-free Alignment for Conflicting Objectives Peter Chen et.al. 2602.02495 null
2026-02-02 MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-Training Dulhan Jayalath et.al. 2602.02494 null
2026-02-02 Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability Xiao Liang et.al. 2602.02477 null
2026-02-02 MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents Haozhen Zhang et.al. 2602.02474 null
2026-02-02 SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning Qifan Yu et.al. 2602.02472 null
2026-02-02 Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge Xutao Ma et.al. 2602.02470 null
2026-02-02 Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts Aiden Yiliu Li et.al. 2602.02468 null
2026-02-02 Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models Noam Steinmetz Yalon et.al. 2602.02467 null
2026-02-02 MentisOculi: Revealing the Limits of Reasoning with Mental Imagery Jana Zeller et.al. 2602.02465 null
2026-02-02 From Directions to Regions: Decomposing Activations in Language Models via Local Geometry Or Shafran et.al. 2602.02464 null
2026-02-02 Abstract Activation Spaces for Content-Invariant Reasoning in Large Language Models Gabriele Maraia et.al. 2602.02462 null
2026-02-02 MetaCLASS: Metacognitive Coaching for Learning with Adaptive Self-regulation Support Naiming Liu et.al. 2602.02457 null
2026-02-02 Relationship-Aware Hierarchical 3D Scene Graph for Task Reasoning Albert Gassol Puigjaner et.al. 2602.02456 null
2026-02-02 Drift-Bench: Diagnosing Cooperative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction Han Bao et.al. 2602.02455 null
2026-02-02 World-Gymnast: Training Robots with Reinforcement Learning in a World Model Ansh Kumar Sharma et.al. 2602.02454 null
2026-02-02 Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling Andong Chen et.al. 2602.02453 null
2026-02-02 Lower bounds for multivariate independence polynomials and their generalisations Joonkyung Lee et.al. 2602.02450 null
2026-02-02 Large Language Models for Mental Health: A Multilingual Evaluation Nishat Raihan et.al. 2602.02440 null
2026-02-02 Embedding Perturbation may Better Reflect the Uncertainty in LLM Reasoning Qihao Wen et.al. 2602.02427 null
2026-02-02 Repurposing Protein Language Models for Latent Flow-Based Fitness Optimization Amaru Caceres Arroyo et.al. 2602.02425 null
2026-01-30 User Prompting Strategies and Prompt Enhancement Methods for Open-Set Object Detection in XR Environments Junfeng Lin et.al. 2601.23281 null
2026-01-30 FOCUS: DLLMs Know How to Tame Their Compute Bound Kaihua Liang et.al. 2601.23278 null
2026-01-30 UPA: Unsupervised Prompt Agent via Tree-Based Search and Selection Siran Peng et.al. 2601.23273 null
2026-01-30 PaperBanana: Automating Academic Illustration for AI Scientists Dawei Zhu et.al. 2601.23265 null
2026-01-30 TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training Ruijie Zhang et.al. 2601.23261 null
2026-01-30 Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models Ye Yu et.al. 2601.23255 null
2026-01-30 GrepRAG: An Empirical Study and Optimization of Grep-Like Retrieval for Code Completion Baoyi Wang et.al. 2601.23254 null
2026-01-30 Training-Free Test-Time Adaptation with Brownian Distance Covariance in Vision-Language Models Yi Zhang et.al. 2601.23253 null
2026-01-30 Structured Over Scale: Learning Spatial Reasoning from Educational Video Bishoy Galoaa et.al. 2601.23251 null
2026-01-30 ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search Tao Yu et.al. 2601.23232 null
2026-01-30 Agile Reinforcement Learning through Separable Neural Architecture Rajib Mostakim et.al. 2601.23225 null
2026-01-30 Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning Xiangyu Zeng et.al. 2601.23224 null
2026-01-30 Are you going to finish that? A Practical Study of the Tokenization Boundary Problem Hao Xu et.al. 2601.23223 null
2026-01-30 Med-Scout: Curing MLLMs’ Geometric Blindness in Medical Perception via Geometry-Aware RL Post-Training Anglin Liu et.al. 2601.23220 null
2026-01-30 High-quality generation of dynamic game content via small language models: A proof of concept Morten I. K. Munk et.al. 2601.23206 null
2026-01-30 TSAQA: Time Series Analysis Question And Answering Benchmark Baoyu Jing et.al. 2601.23204 null
2026-01-30 Large Language Models for Patent Classification: Strengths, Trade-offs, and the Long Tail Effect Lorenzo Emer et.al. 2601.23200 null
2026-01-30 Deep Search with Hierarchical Meta-Cognitive Monitoring Inspired by Cognitive Neuroscience Zhongxiang Sun et.al. 2601.23188 null
2026-01-30 ReGuLaR: Variational Latent Reasoning Guided by Rendered Chain-of-Thought Fanmeng Wang et.al. 2601.23184 null
2026-01-30 FourierSampler: Unlocking Non-Autoregressive Potential in Diffusion Language Models via Frequency-Guided Generation Siyang He et.al. 2601.23182 null
2026-01-29 UEval: A Benchmark for Unified Multimodal Generation Bo Li et.al. 2601.22155 null
2026-01-29 Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions Xiaoxiao Sun et.al. 2601.22150 null
2026-01-29 DynaWeb: Model-Based Reinforcement Learning of Web Agents Hang Ding et.al. 2601.22149 null
2026-01-29 FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale Ajay Patel et.al. 2601.22146 null
2026-01-29 Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers Xin Chen et.al. 2601.22139 null
2026-01-29 Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference Ziming Dong et.al. 2601.22132 null
2026-01-29 World of Workflows: a Benchmark for Bringing World Models to Enterprise Systems Lakshya Gupta et.al. 2601.22130 null
2026-01-29 SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents Yifeng Ding et.al. 2601.22129 null
2026-01-29 The Patient is not a Moving Document: A World Model Training Paradigm for Longitudinal EHR Irsyad Adam et.al. 2601.22128 null
2026-01-29 A Federated and Parameter-Efficient Framework for Large Language Model Training in Medicine Anran Li et.al. 2601.22124 null
2026-01-29 SINA: A Circuit Schematic Image-to-Netlist Generator Using Artificial Intelligence Saoud Aldowaish et.al. 2601.22114 null
2026-01-29 Value-Based Pre-Training with Downstream Feedback Shuqi Ke et.al. 2601.22108 null
2026-01-29 ECO: Quantized Training without Full-Precision Master Weights Mahdi Nikdan et.al. 2601.22101 null
2026-01-29 Latent Adversarial Regularization for Offline Preference Optimization Enyi Jiang et.al. 2601.22083 null
2026-01-29 VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning Yibo Wang et.al. 2601.22069 null
2026-01-29 Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models Wenxuan Huang et.al. 2601.22060 null
2026-01-29 AIRPET: Virtual Positron Emission Tomography J. Renner et.al. 2601.22059 null
2026-01-29 MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources Baorui Ma et.al. 2601.22054 null
2026-01-29 MasalBench: A Benchmark for Contextual and Cross-Cultural Understanding of Persian Proverbs in LLMs Ghazal Kalhor et.al. 2601.22050 null
2026-01-29 On the Paradoxical Interference between Instruction-Following and Task Solving Yunjia Qi et.al. 2601.22047 null
2026-01-28 When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation David Tan et.al. 2601.20858 null
2026-01-28 SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models Sebastiano Monti et.al. 2601.20856 null
2026-01-28 Reward Models Inherit Value Biases from Pretraining Brian Christian et.al. 2601.20838 null
2026-01-28 Open-Vocabulary Functional 3D Human-Scene Interaction Generation Jie Liu et.al. 2601.20835 null
2026-01-28 Linear representations in language models can change dramatically over a conversation Andrew Kyle Lampinen et.al. 2601.20834 null
2026-01-28 Idea2Story: An Automated Pipeline for Transforming Research Concepts into Complete Scientific Narratives Tengyue Xu et.al. 2601.20833 null
2026-01-28 MemCtrl: Using MLLMs as Active Memory Controllers on Embodied Agents Vishnu Sashank Dorbala et.al. 2601.20831 null
2026-01-28 Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning Minwu Kim et.al. 2601.20829 null
2026-01-28 Context-Augmented Code Generation Using Programming Knowledge Graphs Shahd Seddik et.al. 2601.20810 null
2026-01-28 Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning in Few-Shot Relation Extraction Aunabil Chakma et.al. 2601.20803 null
2026-01-28 Reinforcement Learning via Self-Distillation Jonas Hübotter et.al. 2601.20802 null
2026-01-28 Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers Yiran Huang et.al. 2601.20796 null
2026-01-28 Agentic Fog: A Policy-driven Framework for Distributed Intelligence in Fog Computing Saeed Akbar et.al. 2601.20764 null
2026-01-28 Persona Prompting as a Lens on LLM Social Reasoning Jing Yang et.al. 2601.20757 null
2026-01-28 ProfInfer: An eBPF-based Fine-Grained LLM Inference Profiler Bohua Zou et.al. 2601.20755 null
2026-01-28 Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts Elham Aghakhani et.al. 2601.20747 null
2026-01-28 HESTIA: A Hessian-Guided Differentiable Quantization-Aware Training Framework for Extremely Low-Bit LLMs Guoan Wang et.al. 2601.20745 null
2026-01-28 Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification Xin Jin et.al. 2601.20742 null
2026-01-28 QueerGen: How LLMs Reflect Societal Norms on Gender and Sexuality in Sentence Completion Tasks Mae Sosto et.al. 2601.20731 null
2026-01-28 AgentLongBench: A Controllable Long Benchmark For Long-Contexts Agents via Environment Rollouts Shicheng Fang et.al. 2601.20730 null
2026-01-27 Evaluation of Oncotimia: An LLM based system for supporting tumour boards Luis Lorenzo et.al. 2601.19899 null
2026-01-27 Self-Distillation Enables Continual Learning Idan Shenfeld et.al. 2601.19897 null
2026-01-27 Post-LayerNorm Is Back: Stable, ExpressivE, and Deep Chen Chen et.al. 2601.19895 null
2026-01-27 Reflective Translation: Improving Low-Resource Machine Translation via Structured Self-Reflection Nicholas Cheng et.al. 2601.19871 null
2026-01-27 EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning Binzhu Xie et.al. 2601.19850 null
2026-01-27 Identifying and Transferring Reasoning-Critical Neurons: Improving LLM Inference Reliability via Activation Steering Fangan Dong et.al. 2601.19847 null
2026-01-27 HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs Jeanne Malécot et.al. 2601.19839 null
2026-01-27 Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models Jialong Wu et.al. 2601.19834 null
2026-01-27 Neural Neural Scaling Laws Michael Y. Hu et.al. 2601.19831 null
2026-01-27 When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering Mahdi Astaraki et.al. 2601.19827 null
2026-01-27 Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts TrungKhang Tran et.al. 2601.19811 null
2026-01-27 Zero-Shot Stance Detection in the Wild: Dynamic Target Generation and Multi-Target Adaptation Aohua Li et.al. 2601.19802 null
2026-01-27 Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision Zhixiang Wei et.al. 2601.19798 null
2026-01-27 Component-Aware Pruning Framework for Neural Network Controllers via Gradient-Based Importance Estimation Ganesh Sundaram et.al. 2601.19794 null
2026-01-27 Phonological Tokenizer: Prosody-Aware Phonetic Token via Multi-Objective Fine-Tuning with Differentiable K-Means Kentaro Onda et.al. 2601.19781 null
2026-01-27 GAVEL: Towards rule-based safety through activation monitoring Shir Rozenfeld et.al. 2601.19768 null
2026-01-27 Reimagining Social Robots as Recommender Systems: Foundations, Framework, and Applications Jin Huang et.al. 2601.19761 null
2026-01-27 Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues Junchen Fu et.al. 2601.19750 null
2026-01-27 Veri-Sure: A Contract-Aware Multi-Agent Framework with Temporal Tracing and Formal Verification for Correct RTL Code Generation Jiale Liu et.al. 2601.19747 null
2026-01-27 TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token Ditching Runjia Zeng et.al. 2601.19739 null
2026-01-26 ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models Brian Ondov et.al. 2601.18796 null
2026-01-26 MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts Etienne Lanzeray et.al. 2601.18790 null
2026-01-26 Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System Tiffany Wang et.al. 2601.18785 null
2026-01-26 POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration Yuxiao Qu et.al. 2601.18779 null
2026-01-26 PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation Abhishek Divekar et.al. 2601.18777 null
2026-01-26 Dep-Search: Learning Dependency-Aware Reasoning Traces with Persistent Memory Yanming Liu et.al. 2601.18771 null
2026-01-26 Goal-oriented Communication for Fast and Robust Robotic Fault Detection and Recovery Shutong Chen et.al. 2601.18765 null
2026-01-26 Beyond Preferences: Learning Alignment Principles Grounded in Human Reasons and Values Henry Bell et.al. 2601.18760 null
2026-01-26 $α^3$ -SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks Mohamed Amine Ferrag et.al. 2601.18754 null
2026-01-26 HalluGuard: Demystifying Data-Driven and Reasoning-Driven Hallucinations in LLMs Xinyue Zeng et.al. 2601.18753 null
2026-01-26 Why Keep Your Doubts to Yourself? Trading Visual Uncertainties in Multi-Agent Bandit Systems Jusheng Zhang et.al. 2601.18735 null
2026-01-26 Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models Siyan Zhao et.al. 2601.18734 null
2026-01-26 Advances and Innovations in the Multi-Agent Robotic System (MARS) Challenge Li Kang et.al. 2601.18733 null
2026-01-26 One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment Hongru Cai et.al. 2601.18731 null
2026-01-26 Reflect: Transparent Principle-Guided Reasoning for Constitutional Alignment at Scale Henry Bell et.al. 2601.18730 null
2026-01-26 Trustworthy Evaluation of Robotic Manipulation: A New Benchmark and AutoEval Methods Mengyuan Liu et.al. 2601.18723 null
2026-01-26 Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning Lintang Sutawika et.al. 2601.18722 null
2026-01-26 Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs Zhichao Yang et.al. 2601.18706 null
2026-01-26 From Fuzzy to Exact: The Halo Architecture for Infinite-Depth Reasoning via Rational Arithmetic Hansheng Ren et.al. 2601.18702 null
2026-01-26 Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning Olaf Yunus Laitinen Imanov et.al. 2601.18699 null
2026-01-23 A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs Dayal Singh Kalra et.al. 2601.16979 null
2026-01-23 VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents Zirui Wang et.al. 2601.16973 null
2026-01-23 Auto-Regressive Masked Diffusion Models Mahdi Karami et.al. 2601.16971 null
2026-01-23 Empowering Medical Equipment Sustainability in Low-Resource Settings: An AI-Powered Diagnostic and Support Platform for Biomedical Technicians Bernes Lorier Atabonfack et.al. 2601.16967 null
2026-01-23 AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems Mohamed Amine Ferrag et.al. 2601.16964 null
2026-01-23 DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Avinash Maurya et.al. 2601.16956 null
2026-01-23 Strategies for Span Labeling with Large Language Models Danil Semin et.al. 2601.16946 null
2026-01-23 GRIP: Algorithm-Agnostic Machine Unlearning for Mixture-of-Experts via Geometric Router Constraints Andy Zhu et.al. 2601.16905 null
2026-01-23 Evaluating Large Vision-language Models for Surgical Tool Detection Nakul Poudel et.al. 2601.16895 null
2026-01-23 Mixture-of-Models: Unifying Heterogeneous Agents via N-Way Self-Evaluating Deliberation Tims Pecerskis et.al. 2601.16863 null
2026-01-23 Reasoning Promotes Robustness in Theory of Mind Tasks Ian B. de Haan et.al. 2601.16853 null
2026-01-23 Trapped in the past? Disentangling fluid and crystallized intelligence of large language models using chess Leonard S. Pleiss et.al. 2601.16823 null
2026-01-23 Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis Gaurav Negi et.al. 2601.16800 null
2026-01-23 Persuasion Tokens for Editing Factual Knowledge in LLMs Paul Youssef et.al. 2601.16781 null
2026-01-23 PocketDVDNet: Realtime Video Denoising for Real Camera Noise Crispian Morris et.al. 2601.16780 null
2026-01-23 LLM-powered Real-time Patent Citation Recommendation for Financial Technologies Tianang Deng et.al. 2601.16775 null
2026-01-23 Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation Xinyi Wang et.al. 2601.16753 null
2026-01-23 LongCat-Flash-Thinking-2601 Technical Report Meituan LongCat Team et.al. 2601.16725 null
2026-01-23 Supporting Stakeholder Requirements Expression with LLM Revisions: An Empirical Evaluation Michael Mircea et.al. 2601.16699 null
2026-01-23 AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning Suzhong Fu et.al. 2601.16685 null
2026-01-22 Point Bridge: 3D Representations for Cross Domain Policy Learning Siddhant Haldar et.al. 2601.16212 null
2026-01-22 IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance Jongwoo Park et.al. 2601.16207 null
2026-01-22 Provable Robustness in Multimodal Large Language Models via Feature Space Smoothing Song Xia et.al. 2601.16200 null
2026-01-22 PAL*M: Property Attestation for Large Generative Models Prach Chantasantitam et.al. 2601.16199 null
2026-01-22 Structured Hints for Sample-Efficient Lean Theorem Proving Zachary Burton et.al. 2601.16172 null
2026-01-22 Beat-ssl: Capturing Local ECG Morphology through Heartbeat-level Contrastive Learning with Soft Targets Muhammad Ilham Rizqyawan et.al. 2601.16147 null
2026-01-22 Low-altitude Multi-UAV-assisted Data Collection and Semantic Forwarding for Post-Disaster Relief Xiaoya Zheng et.al. 2601.16146 null
2026-01-22 LLM Prompt Evaluation for Educational Applications Langdon Holmes et.al. 2601.16134 null
2026-01-22 Improving Training Efficiency and Reducing Maintenance Costs via Language Specific Model Merging Alphaeus Dmonte et.al. 2601.16127 null
2026-01-22 Multimodal Climate Disinformation Detection: Integrating Vision-Language Models with External Knowledge Sources Marzieh Adeli Shamsabad et.al. 2601.16108 null
2026-01-22 Adapter Fusion for Multilingual Text2Cypher with Linear and Learned Gating Makbule Gulcin Ozsoy et.al. 2601.16097 null
2026-01-22 Controlling Long-Horizon Behavior in Language Model Agents with Explicit State Dynamics Sukesh Subaharan et.al. 2601.16087 null
2026-01-22 DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models Chenyang Li et.al. 2601.16065 null
2026-01-22 DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning Junha Lee et.al. 2601.16046 null
2026-01-22 Grounding Large Language Models in Reaction Knowledge Graphs for Synthesis Retrieval Olga Bunkova et.al. 2601.16038 null
2026-01-22 Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10 Yifan Zhu et.al. 2601.16032 null
2026-01-22 Deja Vu in Plots: Leveraging Cross-Session Evidence with Retrieval-Augmented LLMs for Live Streaming Risk Assessment Yiran Qiao et.al. 2601.16027 null
2026-01-22 Timbre-Aware LLM-based Direct Speech-to-Speech Translation Extendable to Multiple Language Pairs Lalaram Arya et.al. 2601.16023 null
2026-01-22 Mecellem Models: Turkish Models Trained from Scratch and Continually Pre-trained for the Legal Domain Özgür Uğur et.al. 2601.16018 null
2026-01-22 PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models Chak-Wing Mak et.al. 2601.16007 null
2026-01-21 Towards Understanding Best Practices for Quantization of Vision-Language Models Gautom Das et.al. 2601.15287 null
2026-01-21 Iterative Refinement Improves Compositional Image Generation Shantanu Jaiswal et.al. 2601.15286 null
2026-01-21 MolecularIQ: Characterizing Chemical Reasoning Capabilities Through Symbolic Verification on Molecular Graphs Christoph Bartmann et.al. 2601.15279 null
2026-01-21 Robust Fake News Detection using Large Language Models under Adversarial Sentiment Attacks Sahar Tahmasebi et.al. 2601.15277 null
2026-01-21 Lightweight LLMs for Network Attack Detection in IoT Networks Piyumi Bhagya Sudasinghe et.al. 2601.15269 null
2026-01-21 Evaluation of Large Language Models in Legal Applications: Challenges, Methods, and Future Directions Yiran Hu et.al. 2601.15267 null
2026-01-21 The Effect of Scripts and Formats on LLM Numeracy Varshini Reddy et.al. 2601.15251 null
2026-01-21 Multi-context principal component analysis Kexin Wang et.al. 2601.15239 null
2026-01-21 Metadata Conditioned Large Language Models for Localization Anjishnu Mukherjee et.al. 2601.15236 null
2026-01-21 When Agents Fail: A Comprehensive Study of Bugs in LLM Agents with Automated Labeling Niful Islam et.al. 2601.15232 null
2026-01-21 PROGRESSLM: Towards Progress Reasoning in Vision-Language Models Jianshu Zhang et.al. 2601.15224 null
2026-01-21 Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models Anmol Goel et.al. 2601.15220 null
2026-01-21 Deaf and Hard of Hearing Access to Intelligent Personal Assistants: Comparison of Voice-Based Options with an LLM-Powered Touch Interface Paige S. DeVries et.al. 2601.15209 null
2026-01-21 Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback Stephan Wallraven et.al. 2601.15188 null
2026-01-21 Supporting Humans in Evaluating AI Summaries of Legal Depositions Naghmeh Farzi et.al. 2601.15182 null
2026-01-21 The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models Zanlin Ni et.al. 2601.15165 null
2026-01-21 Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems Yinzhu Chen et.al. 2601.15161 null
2026-01-21 Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning Yuval Kansal et.al. 2601.15160 null
2026-01-21 Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data Yuval Ran-Milo et.al. 2601.15158 null
2026-01-21 How to Build AI Agents by Augmenting LLMs with Codified Human Expert Domain Knowledge? A Software Engineering Framework Choro Ulan uulu et.al. 2601.15153 null
2026-01-20 LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR Said Taghadouini et.al. 2601.14251 null
2026-01-20 Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment Yuming Yang et.al. 2601.14249 null
2026-01-20 Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow Haocheng Xi et.al. 2601.14243 null
2026-01-20 Attention-Based Offline Reinforcement Learning and Clustering for Interpretable Sepsis Treatment Punit Kumar et.al. 2601.14228 null
2026-01-20 Transformer Architectures for Respiratory Sound Analysis and Multimodal Diagnosis Theodore Aptekarev et.al. 2601.14227 null
2026-01-20 HALT: Hallucination Assessment via Latent Testing Rohan Bhatnagar et.al. 2601.14210 null
2026-01-20 InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning Matthew Y. R. Yang et.al. 2601.14209 null
2026-01-20 Toward Efficient Agents: Memory, Tool learning, and Planning Xiaofang Yang et.al. 2601.14192 null
2026-01-20 IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models Liang Shi et.al. 2601.14188 null
2026-01-20 ReSearch: A Multi-Stage Machine Learning Framework for Earth Science Data Discovery Youran Sun et.al. 2601.14176 null
2026-01-20 Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum Víctor Yeste et.al. 2601.14172 null
2026-01-20 Domain-Adaptation through Synthetic Data: Fine-Tuning Large Language Models for German Law Ali Hamza Bashir et.al. 2601.14160 null
2026-01-20 LLM Augmented Intervenable Multimodal Adaptor for Post-operative Complication Prediction in Lung Cancer Surgery Shubham Pandey et.al. 2601.14154 null
2026-01-20 Lost in the Prompt Order: Revealing the Limitations of Causal Attention in Language Models Hyunjong Ok et.al. 2601.14152 null
2026-01-20 The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization Meng Li et.al. 2601.14148 null
2026-01-20 CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems Tong Xie et.al. 2601.14140 null
2026-01-20 TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers Bin Yu et.al. 2601.14133 null
2026-01-20 The Side Effects of Being Smart: Safety Risks in MLLMs’ Multi-Image Reasoning Renmiao Chen et.al. 2601.14127 null
2026-01-20 Style Transfer as Bias Mitigation: Diffusion Models for Synthetic Mental Health Text for Arabic Saad Mankarious et.al. 2601.14124 null
2026-01-20 NewsRECON: News article REtrieval for image CONtextualization Jonathan Tonglet et.al. 2601.14121 null
2026-01-16 Do explanations generalize across large reasoning models? Koyena Pal et.al. 2601.11517 null
2026-01-16 Building Production-Ready Probes For Gemini János Kramár et.al. 2601.11516 null
2026-01-16 ShapeR: Robust Conditional 3D Shape Generation from Casual Captures Yawar Siddiqui et.al. 2601.11514 null
2026-01-16 Health Facility Location in Ethiopia: Leveraging LLMs to Integrate Expert Knowledge into Algorithmic Planning Yohai Trabelsi et.al. 2601.11479 null
2026-01-16 MHA2MLA-VLM: Enabling DeepSeek’s Economical Multi-Head Latent Attention across Vision-Language Models Xiaoran Fan et.al. 2601.11464 null
2026-01-16 Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation Xin Sun et.al. 2601.11443 null
2026-01-16 Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps Xiangjun Gao et.al. 2601.11442 null
2026-01-16 Hierarchical Orthogonal Residual Spread for Precise Massive Editing in Large Language Models Xiaojie Gu et.al. 2601.11441 null
2026-01-16 The unreasonable effectiveness of pattern matching Gary Lupyan et.al. 2601.11432 null
2026-01-16 Relational Linearity is a Predictor of Hallucinations Yuetian Lu et.al. 2601.11429 null
2026-01-16 ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models Linqing Zhong et.al. 2601.11404 null
2026-01-16 Understanding Help Seeking for Digital Privacy, Safety, and Security Kurt Thomas et.al. 2601.11398 null
2026-01-16 Evaluating LLM Behavior in Hiring: Implicit Weights, Fairness Across Groups, and Alignment with Human Preferences Morgane Hoffmann et.al. 2601.11379 null
2026-01-16 RITA: A Tool for Automated Requirements Classification and Specification from Online User Feedback Manjeshwar Aniruddh Mallya et.al. 2601.11362 null
2026-01-16 Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding Wenhui Tan et.al. 2601.11359 null
2026-01-16 AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems Weiyi Wang et.al. 2601.11354 null
2026-01-16 How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting Parker Seegmiller et.al. 2601.11344 null
2026-01-16 Unlocking the Potentials of Retrieval-Augmented Generation for Diffusion Language Models Chuanyue Yu et.al. 2601.11342 null
2026-01-16 Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models Guoming Ling et.al. 2601.11340 null
2026-01-16 Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming Sama Hadhoud et.al. 2601.11332 null
2026-01-15 Alterbute: Editing Intrinsic Attributes of Objects in Images Tal Reiss et.al. 2601.10714 null
2026-01-15 MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching Changle Qu et.al. 2601.10712 null
2026-01-15 From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion Cheng Chen et.al. 2601.10710 null
2026-01-15 Grounding Agent Memory in Contextual Intent Ruozhen Yang et.al. 2601.10702 null
2026-01-15 On the origin of neural scaling laws: from random graphs to natural language Maissam Barkeshli et.al. 2601.10684 null
2026-01-15 Structure and Diversity Aware Context Bubble Construction for Enterprise Retrieval Augmented Systems Amir Khurshid et.al. 2601.10681 null
2026-01-15 Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models Zirui Ren et.al. 2601.10679 null
2026-01-15 Single-Stage Huffman Encoder for ML Compression Aditya Agrawal et.al. 2601.10673 null
2026-01-15 Detecting Winning Arguments with Large Language Models and Persuasion Strategies Tiziano Labruna et.al. 2601.10660 null
2026-01-15 PACEvolve: Enabling Long-Horizon Progress-Aware Consistent Evolution Minghao Yan et.al. 2601.10657 null
2026-01-15 Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs Yuxi Xia et.al. 2601.10645 null
2026-01-15 Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding Christopher Clark et.al. 2601.10611 null
2026-01-15 iTIMO: An LLM-empowered Synthesis Dataset for Travel Itinerary Modification Zhuoxuan Huang et.al. 2601.10609 null
2026-01-15 Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay Hao Wang et.al. 2601.10589 null
2026-01-15 From Single to Multi-Agent Reasoning: Advancing GeneGPT for Genomics QA Kimia Abedini et.al. 2601.10581 null
2026-01-15 Form and Meaning in Intrinsic Multilingual Evaluations Wessel Poelman et.al. 2601.10580 null
2026-01-15 Generative AI collective behavior needs an interactionist paradigm Laura Ferrarotti et.al. 2601.10567 null
2026-01-15 Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure Luxuan Fu et.al. 2601.10551 null
2026-01-15 Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing Yinzhi Zhao et.al. 2601.10543 null
2026-01-15 SVII-3D: Advancing Roadside Infrastructure Inventory with Decimeter-level 3D Localization and Comprehension from Sparse Street Imagery Chong Liu et.al. 2601.10535 null
2026-01-14 Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning Chi-Pin Huang et.al. 2601.09708 null
2026-01-14 Value-Aware Numerical Representations for Transformer Language Models Andreea Dutulescu et.al. 2601.09706 null
2026-01-14 ShortCoder: Knowledge-Augmented Syntax Optimization for Token-Efficient Code Generation Sicong Liu et.al. 2601.09703 null
2026-01-14 How well LLM-based test generation techniques perform with newer LLM versions? Michael Konstantinou et.al. 2601.09695 null
2026-01-14 LLMs can Compress LLMs: Adaptive Pruning by Agents Sai Varun Kodathala et.al. 2601.09694 null
2026-01-14 Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection Tianyi Niu et.al. 2601.09692 null
2026-01-14 Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection Ziyu Yang et.al. 2601.09684 null
2026-01-14 Automating Supply Chain Disruption Monitoring via an Agentic AI Approach Sara AlMahri et.al. 2601.09680 null
2026-01-14 Self-Supervised Animal Identification for Long Videos Xuyang Fang et.al. 2601.09663 null
2026-01-14 LiteEmbed: Adapting CLIP to Rare Classes Aishwarya Agarwal et.al. 2601.09661 null
2026-01-14 Image2Garment: Simulation-ready Garment Generation from a Single Image Selim Emir Can et.al. 2601.09658 null
2026-01-14 Exploring Fine-Tuning for Tabular Foundation Models Aditya Tanna et.al. 2601.09654 null
2026-01-14 LLMs Got Rhythm? Hybrid Phonological Filtering for Greek Poetry Rhyme Detection and Generation Stergios Chatzikyriakidis et.al. 2601.09631 null
2026-01-14 From Prompt to Protocol: Fast Charging Batteries with Large Language Models Ge Lei et.al. 2601.09626 null
2026-01-14 The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multi-Step Malware Ben Nassi et.al. 2601.09625 null
2026-01-14 Toward Understanding Unlearning Difficulty: A Mechanistic Perspective and Circuit-Guided Difficulty Metric Jiali Cheng et.al. 2601.09624 null
2026-01-14 CogRail: Benchmarking VLMs in Cognitive Intrusion Perception for Intelligent Railway Transportation Systems Yonglin Tian et.al. 2601.09613 null
2026-01-14 DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing Qian Cao et.al. 2601.09609 null
2026-01-14 OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding Sheng-Yu Huang et.al. 2601.09575 null
2026-01-14 Dialogue Telemetry: Turn-Level Instrumentation for Autonomous Information Gathering Dimitris Panagopoulos et.al. 2601.09570 null
2026-01-13 Modeling LLM Agent Reviewer Dynamics in Elo-Ranked Review System Hsiang-Wei Huang et.al. 2601.08829 null
2026-01-13 Reasoning Matters for 3D Visual Grounding Hsiang-Wei Huang et.al. 2601.08811 null
2026-01-13 Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge Yao Tang et.al. 2601.08808 null
2026-01-13 MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm Bowen Zhou et.al. 2601.08800 null
2026-01-13 Uncovering Political Bias in Large Language Models using Parliamentary Voting Records Jieying Chen et.al. 2601.08785 null
2026-01-13 LWM-Spectro: A Foundation Model for Wireless Baseband Signal Spectrograms Namhyun Kim et.al. 2601.08780 null
2026-01-13 Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling Yang Cai et.al. 2601.08777 null
2026-01-13 Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs Zhiyuan Hu et.al. 2601.08763 null
2026-01-13 M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding Juntao Jiang et.al. 2601.08758 null
2026-01-13 Inferring Latent Intentions: Attributional Natural Language Inference in LLM Agents Xin Quan et.al. 2601.08742 null
2026-01-13 From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet Understanding Anmol Gulati et.al. 2601.08741 null
2026-01-13 PrivGemo: Privacy-Preserving Dual-Tower Graph Retrieval for Empowering LLM Reasoning with Memory Augmentation Xingyu Tan et.al. 2601.08739 null
2026-01-13 TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback Prithwish Jana et.al. 2601.08734 null
2026-01-13 RAGShaper: Eliciting Sophisticated Agentic RAG Skills via Automated Data Synthesis Zhengwei Tao et.al. 2601.08699 null
2026-01-13 Nationality and Region Prediction from Names: A Comparative Study of Neural Models and Large Language Models Keito Inoshita et.al. 2601.08692 null
2026-01-13 LLMs in Code Vulnerability Analysis: A Proof of Concept Shaznin Sultana et.al. 2601.08691 null
2026-01-13 QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models Zhaolu Kang et.al. 2601.08689 null
2026-01-13 Advancing ESG Intelligence: An Expert-level Agent and Comprehensive Benchmark for Sustainable Finance Yilei Zhao et.al. 2601.08676 null
2026-01-13 Why AI Alignment Failure Is Structural: Learned Human Interaction Structures and AGI as an Endogenous Evolutionary Shock Didier Sornette et.al. 2601.08673 null
2026-01-13 Analyzing Bias in False Refusal Behavior of Large Language Models for Hate Speech Detoxification Kyuri Im et.al. 2601.08668 null
2026-01-12 SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations Mohammed Himayath Ali et.al. 2601.07835 null
2026-01-12 Reference Games as a Testbed for the Alignment of Model Uncertainty and Clarification Requests Manar Ali et.al. 2601.07820 null
2026-01-12 More Images, More Problems? A Controlled Analysis of VLM Failure Modes Anurag Das et.al. 2601.07812 null
2026-01-12 The Confidence Trap: Gender Bias and Predictive Certainty in LLMs Ahmed Sabir et.al. 2601.07806 null
2026-01-12 Learning Through Dialogue: Unpacking the Dynamics of Human-LLM Conversations on Political Issues Shaz Furniturewala et.al. 2601.07796 null
2026-01-12 Vision-Language Model for Accurate Crater Detection Patrick Bauer et.al. 2601.07795 null
2026-01-12 Kinship Data Benchmark for Multi-hop Reasoning Tianda Sun et.al. 2601.07794 null
2026-01-12 Benchmarking Small Language Models and Small Reasoning Language Models on System Log Severity Classification Yahya Masri et.al. 2601.07790 null
2026-01-12 “TODO: Fix the Mess Gemini Created”: Towards Understanding GenAI-Induced Self-Admitted Technical Debt Abdullah Al Mujahid et.al. 2601.07786 null
2026-01-12 Enhancing Self-Correction in Large Language Models through Multi-Perspective Reflection Mariana Costa et.al. 2601.07780 null
2026-01-12 OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent Bowen Yang et.al. 2601.07779 null
2026-01-12 Are LLM Decisions Faithful to Verbal Confidence? Jiawei Wang et.al. 2601.07767 null
2026-01-12 Contrastive Learning with Narrative Twins for Modeling Story Salience Igor Sterner et.al. 2601.07765 null
2026-01-12 Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding Yanxiang Huang et.al. 2601.07761 null
2026-01-12 Structure First, Reason Next: Enhancing a Large Language Model using Knowledge Graph for Numerical Reasoning in Financial Documents Aryan Mishra et.al. 2601.07754 null
2026-01-12 Evaluating the encoding competence of visual language models using uncommon actions Chen Ling et.al. 2601.07737 null
2026-01-12 Is Agentic RAG worth it? An experimental comparison of RAG approaches Pietro Ferrazzi et.al. 2601.07711 null
2026-01-12 Emotional Support Evaluation Framework via Controllable and Diverse Seeker Simulator Chaewon Heo et.al. 2601.07698 null
2026-01-12 Exploring the Meta-level Reasoning of Large Language Models via a Tool-based Multi-hop Tabular Question Answering Task Nick Ferguson et.al. 2601.07696 null
2026-01-12 Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model Siwen Jiao et.al. 2601.07695 null
2026-01-09 AdaFuse: Adaptive Ensemble Decoding with Test-Time Scaling for LLMs Chengming Cui et.al. 2601.06022 null
2026-01-09 Don’t Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks Elias Lumer et.al. 2601.06007 null
2026-01-09 Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models Bang Zeng et.al. 2601.06006 null
2026-01-09 The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning Qiguang Chen et.al. 2601.06002 null
2026-01-09 Open-Vocabulary 3D Instruction Ambiguity Detection Jiayu Ding et.al. 2601.05991 null
2026-01-09 CyberGFM: Graph Foundation Models for Lateral Movement Detection in Enterprise Networks Isaiah J. King et.al. 2601.05988 null
2026-01-09 Context-Aware Decoding for Faithful Vision-Language Generation Mehrdad Fazli et.al. 2601.05939 null
2026-01-09 Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency Haoming Xu et.al. 2601.05905 null
2026-01-09 Can AI mediation improve democratic deliberation? Michael Henry Tessler et.al. 2601.05904 null
2026-01-09 HAPS: Hierarchical LLM Routing with Joint Architecture and Parameter Search Zihang Tian et.al. 2601.05903 null
2026-01-09 TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents Dawei Wang et.al. 2601.05899 null
2026-01-09 StackPlanner: A Centralized Hierarchical Multi-Agent System with Task-Experience Memory Management Ruizhe Zhang et.al. 2601.05890 null
2026-01-09 An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift Constantinos Karouzos et.al. 2601.05882 null
2026-01-09 Gender Bias in LLMs: Preliminary Evidence from Shared Parenting Scenario in Czech Family Law Jakub Harasta et.al. 2601.05879 null
2026-01-09 iReasoner: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models Meghana Sunil et.al. 2601.05877 null
2026-01-09 Continual-learning for Modelling Low-Resource Languages from Large Language Models Santosh Srinath K et.al. 2601.05874 null
2026-01-09 IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck Huilin Deng et.al. 2601.05870 null
2026-01-09 CLewR: Curriculum Learning with Restarts for Machine Translation Preference Learning Alexandra Dragomir et.al. 2601.05858 null
2026-01-09 Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs Sandeep Mishra et.al. 2601.05851 null
2026-01-09 Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation Molly Kennedy et.al. 2601.05835 null
2026-01-08 LaST $_{0}$ : Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model Zhuoyang Liu et.al. 2601.05248 null
2026-01-08 GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Shih-Yang Liu et.al. 2601.05242 null
2026-01-08 Robust Reasoning as a Symmetry-Protected Topological Phase Ilmo Sung et.al. 2601.05240 null
2026-01-08 Measuring and Fostering Peace through Machine Learning and Artificial Intelligence P. Gilda et.al. 2601.05232 null
2026-01-08 Internal Representations as Indicators of Hallucinations in Agent Tool Selection Kait Healy et.al. 2601.05214 null
2026-01-08 MoE3D: A Mixture-of-Experts Module for 3D Reconstruction Zichen Wang et.al. 2601.05208 null
2026-01-08 Stock Market Price Prediction using Neural Prophet with Deep Neural Network Navin Chhibber et.al. 2601.05202 null
2026-01-08 Mechanisms of Prompt-Induced Hallucination in Vision-Language Models William Rudman et.al. 2601.05201 null
2026-01-08 LELA: an LLM-based Entity Linking Approach with Zero-Shot Domain Adaptation Samy Haffoudhi et.al. 2601.05192 null
2026-01-08 Cutting AI Research Costs: How Task-Aware Compression Makes Large Language Model Agents Affordable Zuhair Ahmed Khan Taha et.al. 2601.05191 null
2026-01-08 SimuAgent: An LLM-Based Simulink Modeling Assistant Enhanced with Reinforcement Learning Yanchang Liang et.al. 2601.05187 null
2026-01-08 Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop Yaxuan Wang et.al. 2601.05184 null
2026-01-08 VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice Shuming Liu et.al. 2601.05175 null
2026-01-08 FaST: Efficient and Effective Long-Horizon Forecasting for Large-Scale Spatial-Temporal Graphs via Mixture-of-Experts Yiji Zhao et.al. 2601.05174 null
2026-01-08 CoV: Chain-of-View Prompting for Spatial Reasoning Haoyu Zhao et.al. 2601.05172 null
2026-01-08 Reverse-engineering NLI: A study of the meta-inferential properties of Natural Language Inference Rasmus Blanck et.al. 2601.05170 null
2026-01-08 RelayLLM: Efficient Reasoning via Collaborative Decoding Chengsong Huang et.al. 2601.05167 null
2026-01-08 GenAI-DrawIO-Creator: A Framework for Automated Diagram Generation Jinze Yu et.al. 2601.05162 null
2026-01-08 Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering Shuliang Liu et.al. 2601.05159 null
2026-01-08 Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Large Reasoning Models Shuliang Liu et.al. 2601.05144 null
2026-01-07 Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactions Abhishek Rath et.al. 2601.04170 null
2026-01-07 Scanner-Induced Domain Shifts Undermine the Robustness of Pathology Foundation Models Erik Thiringer et.al. 2601.04163 null
2026-01-07 All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection Yuechen Jiang et.al. 2601.04160 null
2026-01-07 FLEx: Language Modeling with Few-shot Language Explanations Adar Avsian et.al. 2601.04157 null
2026-01-07 Diffusion-DRF: Differentiable Reward Flow for Video Diffusion Fine-Tuning Yifan Wang et.al. 2601.04153 null
2026-01-07 LLMberjack: Guided Trimming of Debate Trees for Multi-Party Conversation Creation Leonardo Bottona et.al. 2601.04135 null
2026-01-07 ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models Nikhil Anand et.al. 2601.04131 null
2026-01-07 GeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement Learning Wenshuai Li et.al. 2601.04118 null
2026-01-07 Layer-wise Positional Bias in Short-Context Language Modeling Maryam Rahimi et.al. 2601.04098 null
2026-01-07 KDCM: Reducing Hallucination in LLMs through Explicit Reasoning Structures Jinbo Hao et.al. 2601.04086 null
2026-01-07 Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts Zhihao Zhu et.al. 2601.04073 null
2026-01-07 Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion Yuanfeng Xu et.al. 2601.04056 null
2026-01-07 Modular Prompt Optimization: Optimizing Structured Prompts with Section-Local Textual Gradients Prith Sharma et.al. 2601.04055 null
2026-01-07 When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life Xinyue Lou et.al. 2601.04043 null
2026-01-07 Analyzing and Improving Cross-lingual Knowledge Transfer for Machine Translation David Stap et.al. 2601.04036 null
2026-01-07 HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense Siyuan Li et.al. 2601.04034 null
2026-01-07 Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model Yuan Wang et.al. 2601.04033 null
2026-01-07 SpeakerSleuth: Evaluating Large Audio-Language Models as Judges for Multi-turn Speaker Consistency Jonggeun Lee et.al. 2601.04029 null
2026-01-07 Simulated Students in Tutoring Dialogues: Substance or Illusion? Alexander Scarlatos et.al. 2601.04025 null
2026-01-07 A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Qi Wu et.al. 2601.03992 null
2026-01-06 NavAI: A Generalizable LLM Framework for Navigation Tasks in Virtual Reality Environments Xue Qin et.al. 2601.03251 null
2026-01-06 SLIM: Stealthy Low-Coverage Black-Box Watermarking via Latent-Space Confusion Zones Hengyu Wu et.al. 2601.03242 null
2026-01-06 PET-TURTLE: Deep Unsupervised Support Vector Machines for Imbalanced Data Clusters Javier Salazar Cavazos et.al. 2601.03237 null
2026-01-06 MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents Dongming Jiang et.al. 2601.03236 null
2026-01-06 Multi-RADS Synthetic Radiology Report Dataset and Head-to-Head Benchmarking of 41 Open-Weight and Proprietary Language Models Kartik Bose et.al. 2601.03232 null
2026-01-06 The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization Ruixing Zhang et.al. 2601.03227 null
2026-01-06 MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics Xinghe Chen et.al. 2601.03217 null
2026-01-06 Fine-tuning Small Language Models as Efficient Enterprise Search Relevance Labelers Yue Kang et.al. 2601.03211 null
2026-01-06 UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward Yile Liu et.al. 2601.03205 null
2026-01-06 DIP: Dynamic In-Context Planner For Diffusion Language Models Yang Li et.al. 2601.03199 null
2026-01-06 Empowering Reliable Visual-Centric Instruction Following in MLLMs Weilei He et.al. 2601.03198 null
2026-01-06 Sparse Knowledge Distillation: A Mathematical Framework for Probability-Domain Temperature Scaling and Multi-Stage Compression Aaron R. Flouro et.al. 2601.03195 null
2026-01-06 X-MuTeST: A Multilingual Benchmark for Explainable Hate Speech Detection and A Novel LLM-consulted Explanation Framework Mohammad Zia Ur Rehman et.al. 2601.03194 null
2026-01-06 MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory Shengtao Zhang et.al. 2601.03192 null
2026-01-06 AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation Anees Ur Rehman Hashmi et.al. 2601.03191 null
2026-01-06 Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning Naixin Zhai et.al. 2601.03190 null
2026-01-06 Decentralized Autoregressive Generation Stepan Maschan et.al. 2601.03184 null
2026-01-06 DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code Generation Jiajun jiao et.al. 2601.03178 null
2026-01-06 WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning Yu Xinmiao et.al. 2601.03164 null
2026-01-06 Decoupling the Effect of Chain-of-Thought Reasoning: A Human Label Variation Perspective Beiduo Chen et.al. 2601.03154 null
2026-01-05 Heterogeneous Low-Bandwidth Pre-Training of LLMs Yazan Obeidi et.al. 2601.02360 null
2026-01-05 VINO: A Unified Visual Generator with Interleaved OmniModal Context Junyi Chen et.al. 2601.02358 null
2026-01-05 Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes Jing Tan et.al. 2601.02356 null
2026-01-05 Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling Falcon LLM Team et.al. 2601.02346 null
2026-01-05 Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling Berk Atil et.al. 2601.02337 null
2026-01-05 Estimating Text Temperature Nikolay Mikhaylovskiy et.al. 2601.02320 null
2026-01-05 DatBench: Discriminative, Faithful, and Efficient VLM Evaluations Siddharth Joshi et.al. 2601.02316 null
2026-01-05 Project Ariadne: A Structural Causal Framework for Auditing Faithfulness in LLM Agents Sourena Khanzadeh et.al. 2601.02314 null
2026-01-05 Placement Semantics for Distributed Deep Learning: A Systematic Framework for Analyzing Parallelism Strategies Deep Pankajbhai Mehta et.al. 2601.02311 null
2026-01-05 Power-of-Two Quantization-Aware-Training (PoT-QAT) in Large Language Models (LLMs) Mahmoud Elgenedy et.al. 2601.02298 null
2026-01-05 CD4LM: Consistency Distillation and aDaptive Decoding for Diffusion Language Models Yihao Liang et.al. 2601.02236 null
2026-01-05 ELLA: Efficient Lifelong Learning for Adapters in Large Language Models Shristi Das Biswas et.al. 2601.02232 null
2026-01-05 From XAI to Stories: A Factorial Study of LLM-Generated Explanation Quality Fabian Lukassen et.al. 2601.02224 null
2026-01-05 CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents Keyu Wang et.al. 2601.02201 null
2026-01-05 Toward Global Large Language Models in Medicine Rui Yang et.al. 2601.02186 null
2026-01-05 Confidence Estimation for LLMs in Multi-turn Interactions Caiqi Zhang et.al. 2601.02179 null
2026-01-05 Streaming Hallucination Detection in Long Chain-of-Thought Reasoning Haolang Lu et.al. 2601.02170 null
2026-01-05 EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning Chuanrui Hu et.al. 2601.02163 null
2026-01-05 FormationEval, an open multiple-choice benchmark for petroleum geoscience Almaz Ermilov et.al. 2601.02158 null
2026-01-05 BiPrompt: Bilateral Prompt Optimization for Visual and Textual Debiasing in Vision-Language Models Sunny Gupta et.al. 2601.02147 null
2026-01-02 Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning Valentin Noël et.al. 2601.00791 null
2026-01-02 Improving Router Security using BERT John Carter et.al. 2601.00783 null
2026-01-02 Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection Akanksha Chuchra et.al. 2601.00777 null
2026-01-02 Memory Bank Compression for Continual Adaptation of Large Language Models Thomas Katraouras et.al. 2601.00756 null
2026-01-02 The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving Max Ruiz Luyten et.al. 2601.00747 null
2026-01-02 Materials Informatics: Emergence To Autonomous Discovery In The Age Of AI Turab Lookman et.al. 2601.00742 null
2026-01-02 Exploring the Performance of Large Language Models on Subjective Span Identification Tasks Alphaeus Dmonte et.al. 2601.00736 null
2026-01-02 Grading Handwritten Engineering Exams with Multimodal Large Language Models Janez Perš et.al. 2601.00730 null
2026-01-02 Detecting Performance Degradation under Data Shift in Pathology Vision-Language Model Hao Guan et.al. 2601.00716 null
2026-01-02 A Vision-and-Knowledge Enhanced Large Language Model for Generalizable Pedestrian Crossing Behavior Inference Qingwen Pu et.al. 2601.00694 null
2026-01-02 Human-like AI-based Auto-Field-in-Field Whole-Brain Radiotherapy Treatment Planning With Conversation Large Language Model Feedback Adnan Jafar et.al. 2601.00685 null
2026-01-02 Sigmoid Head for Quality Estimation under Language Ambiguity Tu Anh Dinh et.al. 2601.00680 null
2026-01-02 QSLM: A Performance- and Memory-aware Quantization Framework with Tiered Search Strategy for Spike-driven Language Models Rachmad Vidya Wicaksana Putra et.al. 2601.00679 null
2026-01-02 IRPO: Scaling the Bradley-Terry Model via Reinforcement Learning Haonan Song et.al. 2601.00677 null
2026-01-02 RoboReward: General-Purpose Vision-Language Reward Models for Robotics Tony Lee et.al. 2601.00675 null
2026-01-02 Fast-weight Product Key Memory Tianyu Zhao et.al. 2601.00671 null
2026-01-02 CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models Neeraj Anand et.al. 2601.00659 null
2026-01-02 Physio-DPO: Aligning Large Language Models with the Protein Energy Landscape to Eliminate Structural Hallucinations QiWei Meng et.al. 2601.00647 null
2026-01-02 FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding Yuchen Li et.al. 2601.00644 null
2026-01-02 Probabilistic Guarantees for Reducing Contextual Hallucinations in LLMs Nils Rautenberg et.al. 2601.00641 null
2025-12-31 Scaling Open-Ended Reasoning to Predict the Future Nikhil Chandak et.al. 2512.25070 null
2025-12-31 Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search Rohit Dwivedula et.al. 2512.25065 null
2025-12-31 Many Minds from One Model: Bayesian Transformers for Population Intelligence Diji Yang et.al. 2512.25063 null
2025-12-31 Context-aware LLM-based AI Agents for Human-centered Energy Management Systems in Smart Buildings Tianzhi He et.al. 2512.25055 null
2025-12-31 Modeling Language as a Sequence of Thoughts Nasim Borazjanizadeh et.al. 2512.25026 null
2025-12-31 ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning Timo Kaufmann et.al. 2512.25023 null
2025-12-31 MAMA-Memeia! Multi-Aspect Multi-Agent Collaboration for Depressive Symptoms Identification in Memes Siddhant Agarwal et.al. 2512.25015 null
2025-12-31 Diffusion Language Models are Provably Optimal Parallel Samplers Haozhe Jiang et.al. 2512.25014 null
2025-12-31 Efficiently Estimating Data Efficiency for Language Model Fine-tuning Gyung Hyun Je et.al. 2512.24991 null
2025-12-31 PhysTalk: Language-driven Real-time Physics in 3D Gaussian Scenes Luca Collorone et.al. 2512.24986 null
2025-12-31 DarkEQA: Benchmarking Vision-Language Models for Embodied Question Answering in Low-Light Indoor Environments Yohan Park et.al. 2512.24985 null
2025-12-31 Evaluating the Impact of Compression Techniques on the Robustness of CNNs under Natural Corruptions Itallo Patrick Castro Alves Da Silva et.al. 2512.24971 null
2025-12-31 Large language models and the entropy of English Colin Scheibner et.al. 2512.24969 null
2025-12-31 The Impact of LLMs on Online News Consumption and Production Hangcheng Zhao et.al. 2512.24968 null
2025-12-31 AMAP Agentic Planning Technical Report Yulan Hu et.al. 2512.24957 null
2025-12-31 CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement Wentao Zhang et.al. 2512.24947 null
2025-12-31 RAIR: A Rule-Aware Benchmark Uniting Challenging Long-Tail and Visual Salience Subset for E-commerce Relevance Assessment Chenji Lu et.al. 2512.24943 null
2025-12-31 Iterative Deployment Improves Planning Skills in LLMs Augusto B. Corrêa et.al. 2512.24940 null
2025-12-31 Vibe Coding, Interface Flattening Hongrui Jin et.al. 2512.24939 null
2025-12-31 Adaptive Dependency-aware Prompt Optimization Framework for Multi-Step LLM Pipeline Minjun Zhao et.al. 2512.24933 null
2025-12-29 Training AI Co-Scientists Using Rubric Rewards Shashwat Goel et.al. 2512.23707 null
2025-12-29 Eliciting Behaviors in Multi-Turn Conversations Jing Huang et.al. 2512.23701 null
2025-12-29 Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans Sky CH-Wang et.al. 2512.23693 null
2025-12-29 PROFASR-BENCH: A Benchmark for Context-Conditioned ASR in High-Stakes Professional Speech Deepak Babu Piskala et.al. 2512.23686 null
2025-12-29 Multilingual Hidden Prompt Injection Attacks on LLM-Based Academic Reviewing Panagiotis Theocharopoulos et.al. 2512.23684 null
2025-12-29 Web World Models Jichen Feng et.al. 2512.23676 null
2025-12-29 End-to-End Test-Time Training for Long Context Arnuv Tandon et.al. 2512.23675 null
2025-12-29 Less is more: Probabilistic reduction is best explained by small-scale predictability measures Cassandra L. Jacobs et.al. 2512.23659 null
2025-12-29 OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding Keda Tao et.al. 2512.23646 null
2025-12-29 BOAD: Discovering Hierarchical Software Engineering Agents via Bandit Optimization Iris Xu et.al. 2512.23631 null
2025-12-29 Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing Yuwen Li et.al. 2512.23611 null
2025-12-29 The Big Three in Marriage Talk: LLM-Assisted Analysis of Moral Ethics and Sentiment on Weibo and Xiaohongshu Frank Tian-Fang Ye et.al. 2512.23609 null
2025-12-29 Divergent-Convergent Thinking in Large Language Models for Creative Problem Generation Manh Hung Nguyen et.al. 2512.23601 null
2025-12-29 Same or Not? Enhancing Visual Perception in Vision-Language Models Damiano Marsili et.al. 2512.23592 null
2025-12-29 Can AI Recognize Its Own Reflection? Self-Detection Performance of LLMs in Computing Education Christopher Burger et.al. 2512.23587 null
2025-12-29 Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models Yu-Xiang Lin et.al. 2512.23578 null
2025-12-29 LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation Ethan Chern et.al. 2512.23576 null
2025-12-29 Instruction-Following Evaluation of Large Vision-Language Models Daiki Shiono et.al. 2512.23572 null
2025-12-29 ThinkGen: Generalized Thinking for Visual Generation Siyu Jiao et.al. 2512.23568 null
2025-12-29 RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature Hanzheng Li et.al. 2512.23565 null
2025-12-26 See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning Shuoshuo Zhang et.al. 2512.22120 null
2025-12-26 Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis Duygu Altinok et.al. 2512.22100 null
2025-12-26 Unifying Learning Dynamics and Generalization in Transformers Scaling Law Chiwun Yang et.al. 2512.22088 null
2025-12-26 Context as a Tool: Context Management for Long-Horizon SWE-Agents Shukai Liu et.al. 2512.22087 null
2025-12-26 Agent-based simulation of online social networks and disinformation Alejandro Buitrago López et.al. 2512.22082 null
2025-12-26 Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling Hannah Atmer et.al. 2512.22066 null
2025-12-26 FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion Zhuoran Zhu et.al. 2512.22036 null
2025-12-26 Context-Aware Intelligent Chatbot Framework Leveraging Mobile Sensing Ziyan Zhang et.al. 2512.22032 null
2025-12-26 iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception Sarthak Mehrotra et.al. 2512.22009 null
2025-12-26 DuaDeep-SeqAffinity: Dual-Stream Deep Learning Framework for Sequence-Only Antigen-Antibody Affinity Prediction Aicha Boutorh et.al. 2512.22007 null
2025-12-26 Look Closer! An Adversarial Parametric Editing Framework for Hallucination Mitigation in VLMs Jiayu Hu et.al. 2512.21999 null
2025-12-26 LVLM-Aided Alignment of Task-Specific Vision Models Alexander Koebler et.al. 2512.21985 null
2025-12-26 Perceive and Calibrate: Analyzing and Enhancing Robustness of Medical Multi-Modal Large Language Models Dunyuan XU et.al. 2512.21964 null
2025-12-26 Broken Words, Broken Performance: Effect of Tokenization on Performance of LLMs Sachin Pawar et.al. 2512.21933 null
2025-12-26 SWE-RM: Execution-free Feedback For Software Engineering Agents KaShun Shum et.al. 2512.21919 null
2025-12-26 Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model Nathan Kallus et.al. 2512.21917 null
2025-12-26 Exploring the Heterogeneity of Tabular Data: A Diversity-aware Data Generator via LLMs Yafeng Tang et.al. 2512.21915 null
2025-12-26 GQ-VAE: A gated quantized VAE for learning variable length tokens Theo Datta et.al. 2512.21913 null
2025-12-26 Accelerate Speculative Decoding with Sparse Computation in Verification Jikai Wang et.al. 2512.21911 null
2025-12-26 Explainable Statute Prediction via Attention-based Model and LLM Prompting Sachin Pawar et.al. 2512.21902 null
2025-12-24 Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models Li-Zhong Szu-Tu et.al. 2512.21337 null
2025-12-24 Streaming Video Instruction Tuning Jiaer Xia et.al. 2512.21334 null
2025-12-24 C2LLM Technical Report: A New Frontier in Code Retrieval via Adaptive Cross-Attention Pooling Jin Qin et.al. 2512.21332 null
2025-12-24 Your Reasoning Benchmark May Not Test Reasoning: Revealing Perception Bottleneck in Abstract Reasoning Benchmarks Xinhe Wang et.al. 2512.21329 null
2025-12-24 Parallel Token Prediction for Language Models Felix Draxler et.al. 2512.21323 null
2025-12-24 Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Consulting, Data Analyst, and Management Tasks Ali Merali et.al. 2512.21316 null
2025-12-24 A Plan Reuse Mechanism for LLM-Driven Agent Guopeng Li et.al. 2512.21309 null
2025-12-24 Quadrupped-Legged Robot Movement Plan Generation using Large Language Model Muhtadin et.al. 2512.21293 null
2025-12-24 SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance Divij Dudeja et.al. 2512.21280 null
2025-12-24 ReaSeq: Unleashing World Knowledge via Reasoning for Sequential Modeling Chuan Wang et.al. 2512.21257 null
2025-12-24 LookPlanGraph: Embodied Instruction Following Method with VLM Graph Augmentation Anatoly O. Onishchenko et.al. 2512.21243 null
2025-12-24 Assessing the Software Security Comprehension of Large Language Models Mohammed Latif Siddiq et.al. 2512.21238 null
2025-12-24 Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking Yifan Huang et.al. 2512.21236 null
2025-12-24 MiST: Understanding the Role of Mid-Stage Scientific Training in Developing Chemical Reasoning Models Andres M Bran et.al. 2512.21231 null
2025-12-24 Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval Dao Sy Duy Minh et.al. 2512.21221 null
2025-12-24 RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic Le Wang et.al. 2512.21220 null
2025-12-24 Latent Implicit Visual Reasoning Kelvin Li et.al. 2512.21218 null
2025-12-24 SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation Mahi Luthra et.al. 2512.21204 null
2025-12-24 VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs Brigitta Malagurski Törtei et.al. 2512.21194 null
2025-12-24 Encrypted Traffic Detection in Resource Constrained IoT Networks: A Diffusion Model and LLM Integrated Framework Hongjuan Li et.al. 2512.21144 null
2025-12-23 Making Large Language Models Efficient Dense Retrievers Yibin Lei et.al. 2512.20612 null
2025-12-23 MoE-DiffuSeq: Enhancing Long-Document Diffusion Models with Sparse Attention and Mixture of Experts Alexandros Christoforos et.al. 2512.20604 null
2025-12-23 Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs Dhruv Anand et.al. 2512.20595 null
2025-12-23 LightTact: A Visual-Tactile Fingertip Sensor for Deformation-Independent Contact Sensing Changyi Lin et.al. 2512.20591 null
2025-12-23 Automated stereotactic radiosurgery planning using a human-in-the-loop reasoning large language model agent Humza Nusrat et.al. 2512.20586 null
2025-12-23 Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits Amirhosein Ghasemabadi et.al. 2512.20578 null
2025-12-23 Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs Rui Pan et.al. 2512.20573 null
2025-12-23 FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models Kaitong Cai et.al. 2512.20561 null
2025-12-23 Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models Shengchao Zhou et.al. 2512.20557 null
2025-12-23 Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios Mingwei Tang et.al. 2512.20556 null
2025-12-23 LLM-Based Authoring of Agent-Based Narratives through Scene Descriptions Vinayak Regmi et.al. 2512.20550 null
2025-12-23 Benchmarking LLMs for Predictive Applications in the Intensive Care Units Chehak Malhotra et.al. 2512.20520 null
2025-12-23 Coherence in the brain unfolds across separable temporal regimes Davide Stauba et.al. 2512.20481 null
2025-12-23 Laser: Governing Long-Horizon Agentic Search via Structured Protocol and Context Register Shuting Wang et.al. 2512.20458 null
2025-12-23 Chain-of-Anomaly Thoughts with Large Vision-Language Models Pedro Domingos et.al. 2512.20417 null
2025-12-23 ChatGPT: Excellent Paper! Accept It. Editor: Imposter Found! Review Rejected Kanchon Gharami et.al. 2512.20405 null
2025-12-23 BRIDGE: Budget-aware Reasoning via Intermediate Distillation with Guided Examples Xuan-An Le et.al. 2512.20403 null
2025-12-23 CRAFT: Continuous Reasoning and Agentic Feedback Tuning for Multimodal Text-to-Image Generation V. Kovalev et.al. 2512.20362 null
2025-12-23 A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice Yaowei Bai et.al. 2512.20344 null
2025-12-23 MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model Ye Tao et.al. 2512.20339 null
2025-12-22 Scalably Enhancing the Clinical Validity of a Task Benchmark with Physician Oversight Junze Ye et.al. 2512.19691 null
2025-12-22 Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models Zixuan Ye et.al. 2512.19686 null
2025-12-22 From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs Mingrui Wu et.al. 2512.19683 null
2025-12-22 GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators Jiacheng Guo et.al. 2512.19682 null
2025-12-22 Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918) Niclas Griesshaber et.al. 2512.19675 null
2025-12-22 Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies Yuqiao Tan et.al. 2512.19673 null
2025-12-22 Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis Argha Kamal Samanta et.al. 2512.19663 null
2025-12-22 Exploring Zero-Shot ACSA with Unified Meaning Representation in Chain-of-Thought Prompting Filippos Ventirozos et.al. 2512.19651 null
2025-12-22 Exploring the features used for summary evaluation by Human and GPT Zahra Sadeghi et.al. 2512.19620 null
2025-12-22 MapTrace: Scalable Data Generation for Route Tracing on Maps Artemis Panagopoulou et.al. 2512.19609 null
2025-12-22 RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference George Karfakis et.al. 2512.19606 null
2025-12-22 Increasing the Thinking Budget is Not All You Need Ignacio Iacobacci et.al. 2512.19585 null
2025-12-22 The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge Angjelin Hila et.al. 2512.19570 null
2025-12-22 Towards Closed-Loop Embodied Empathy Evolution: Probing LLM-Centric Lifelong Empathic Motion Generation in Unseen Scenarios Jiawen Wang et.al. 2512.19551 null
2025-12-22 Event Extraction in Large Language Model Bobo Li et.al. 2512.19537 null
2025-12-22 CASA: Cross-Attention via Self-Attention for Efficient Vision-Language Fusion Moritz Böhle et.al. 2512.19535 null
2025-12-22 Learning Continuous Solvent Effects from Transient Flow Data: A Graph Neural Network Benchmark on Catechol Rearrangement Hongsheng Xing et.al. 2512.19530 null
2025-12-22 QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models Li Puyin et.al. 2512.19526 null
2025-12-22 Anatomy-R1: Enhancing Anatomy Reasoning in Multimodal Large Language Models via Anatomical Similarity Curriculum and Group Diversity Augmentation Ziyang Song et.al. 2512.19512 null
2025-12-22 Beyond Language Boundaries: Uncovering Programming Language Families for Code Language Models Shangbo Yun et.al. 2512.19509 null
2025-12-19 XAgen: An Explainability Tool for Identifying and Correcting Failures in Multi-Agent Workflows Xinru Wang et.al. 2512.17896 null
2025-12-19 ShareChat: A Dataset of Chatbot Conversations in the Wild Yueru Yan et.al. 2512.17843 null
2025-12-19 ReX-MLE: The Autonomous Agent Benchmark for Medical Imaging Challenges Roshan Kenia et.al. 2512.17838 null
2025-12-19 Structure-Aware Antibody Design with Affinity-Optimized Inverse Folding Xinyan Zhao et.al. 2512.17815 null
2025-12-19 LLM-based Behaviour Driven Development for Hardware Design Rolf Drechsler et.al. 2512.17814 null
2025-12-19 DEER: A Comprehensive and Reliable Benchmark for Deep-Research Expert Reports Janghoon Han et.al. 2512.17776 null
2025-12-19 In Times of Crisis: An Exploratory Study of Media and Political Discourse on YouTube During the 2024 French Elections Vera Sosnovik et.al. 2512.17768 null
2025-12-19 AncientBench: Towards Comprehensive Evaluation on Excavated and Transmitted Chinese Corpora Zhihan Zhou et.al. 2512.17756 null
2025-12-19 When the Gold Standard isn’t Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content Lydia Nishimwe et.al. 2512.17738 null
2025-12-19 AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection Yichen Jiang et.al. 2512.17730 null
2025-12-19 Toward Ethical AI Through Bayesian Uncertainty in Neural Question Answering Riccardo Di Sipio et.al. 2512.17677 null
2025-12-19 Region-Constraint In-Context Generation for Instructional Video Editing Zhongwei Zhang et.al. 2512.17650 null
2025-12-19 Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs Zhaolin Cai et.al. 2512.17640 null
2025-12-19 Linear Personality Probing and Steering in LLMs: A Big Five Study Michel Frising et.al. 2512.17639 null
2025-12-19 Trust-Region Adaptive Policy Optimization Mingyu Su et.al. 2512.17636 null
2025-12-19 Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection Menna Elgabry et.al. 2512.17630 null
2025-12-19 PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology Fengchun Liu et.al. 2512.17621 null
2025-12-19 HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly Detection Zhaolin Cai et.al. 2512.17601 null
2025-12-19 Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing Lingxiao Zhao et.al. 2512.17574 null
2025-12-19 Towards Explainable Conversational AI for Early Diagnosis with Large Language Models Maliha Tabassum et.al. 2512.17559 null
2025-12-18 AdaTooler-V: Adaptive Tool-Use for Images and Videos Chaoyang Wang et.al. 2512.16918 null
2025-12-18 Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning Qihao Liu et.al. 2512.16917 null
2025-12-18 Exploration v.s. Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward Peter Chen et.al. 2512.16912 null
2025-12-18 MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning Yuanchen Ju et.al. 2512.16909 null
2025-12-18 How Good is Post-Hoc Watermarking With Language Model Rephrasing? Pierre Fernandez et.al. 2512.16904 null
2025-12-18 Impacts of Racial Bias in Historical Training Data for News AI Rahul Bhargava et.al. 2512.16901 null
2025-12-18 Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image Yushi Hu et.al. 2512.16899 null
2025-12-18 LinkedOut: Linking World Knowledge Representation Out of Video LLM for Next-Generation Video Recommendation Haichao Zhang et.al. 2512.16891 null
2025-12-18 AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning Tzu-Han Lin et.al. 2512.16883 null
2025-12-18 TOGGLE: Temporal Logic-Guided Large Language Model Compression for Edge Khurram Khalil et.al. 2512.16855 null
2025-12-18 Meta-RL Induces Exploration in Language Agents Yulun Jiang et.al. 2512.16848 null
2025-12-18 LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference Harsh Vardhan Bansal et.al. 2512.16843 null
2025-12-18 Radiology Report Generation with Layer-Wise Anatomical Attention Emmanuel D. Muñiz-De-León et.al. 2512.16841 null
2025-12-18 What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels Aditya Yadavalli et.al. 2512.16832 null
2025-12-18 Tiny Recursive Control: Iterative Reasoning for Efficient Optimal Control Amit Jain et.al. 2512.16824 null
2025-12-18 Toward Systematic Counterfactual Fairness Evaluation of Large Language Models: The CAFFE Framework Alessandra Parziale et.al. 2512.16816 null
2025-12-18 Grammar-Forced Translation of Natural Language to Temporal Logic using LLMs William English et.al. 2512.16814 null
2025-12-18 From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs Shubham Mishra et.al. 2512.16795 null
2025-12-18 PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence Xiaopeng Lin et.al. 2512.16793 null
2025-12-18 Inside Out: Uncovering How Comment Internalization Steers LLMs for Better or Worse Aaron Imani et.al. 2512.16790 null
2025-12-17 DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models Lunbin Zeng et.al. 2512.15713 null
2025-12-17 Dynamic Rebatching for Efficient Early-Exit Inference with DREX Xuting Liu et.al. 2512.15705 null
2025-12-17 VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression Kyle Sargent et.al. 2512.15701 null
2025-12-17 Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning Yifei Li et.al. 2512.15693 null
2025-12-17 Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning Zhenwen Liang et.al. 2512.15687 null
2025-12-17 Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers Adam Karvonen et.al. 2512.15674 null
2025-12-17 Explaining the Reasoning of Large Language Models Using Attribution Graphs Chase Walker et.al. 2512.15663 null
2025-12-17 Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning Jiaqi Xu et.al. 2512.15662 null
2025-12-17 Characterizing Mamba’s Selective Memory using Auto-Encoders Tamanna Hossain et.al. 2512.15653 null
2025-12-17 VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression? Hongbo Zhao et.al. 2512.15649 null
2025-12-17 IC-Effect: Precise and Efficient Video Effects Editing via In-Context Learning Yuanhang Li et.al. 2512.15635 null
2025-12-17 How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness Darshita Rathore et.al. 2512.15634 null
2025-12-17 Evaluating Metrics for Safety with LLM-as-Judges Kester Clegg et.al. 2512.15617 null
2025-12-17 Behavior Tokens Speak Louder: Disentangled Explainable Recommendation with Behavior Vocabulary Xinshun Feng et.al. 2512.15614 null
2025-12-17 Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction Mathieu Blondel et.al. 2512.15605 null
2025-12-17 You Never Know a Person, You Only Know Their Defenses: Detecting Levels of Psychological Defense Mechanisms in Supportive Conversations Hongbin Na et.al. 2512.15601 null
2025-12-17 Corrective Diffusion Language Models Shuibai Zhang et.al. 2512.15596 null
2025-12-17 Bolmo: Byteifying the Next Generation of Language Models Benjamin Minixhofer et.al. 2512.15586 null
2025-12-17 Evaluating Large Language Models in Scientific Discovery Zhangde Song et.al. 2512.15567 null
2025-12-17 GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models Bozhou Li et.al. 2512.15560 null
2025-12-16 TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs Jun Zhang et.al. 2512.14698 null
2025-12-16 Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization Yen-Ju Lu et.al. 2512.14687 null
2025-12-16 Fast and Accurate Causal Parallel Decoding using Jacobi Forcing Lanxiang Hu et.al. 2512.14681 null
2025-12-16 EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models Zechen Bai et.al. 2512.14666 null
2025-12-16 Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models Chiyue Wei et.al. 2512.14661 null
2025-12-16 Adapting Speech Language Model to Singing Voice Synthesis Yiwen Zhao et.al. 2512.14657 null
2025-12-16 TiME: Tiny Monolingual Encoders for Efficient NLP Pipelines David Schulmeister et.al. 2512.14645 null
2025-12-16 Beyond Text-to-SQL: Autonomous Research-Driven Database Exploration with DAR Ostap Vykhopen et.al. 2512.14622 null
2025-12-16 WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling Wenqiang Sun et.al. 2512.14614 null
2025-12-16 PerProb: Indirectly Evaluating Memorization in Large Language Models Yihan Liao et.al. 2512.14600 null
2025-12-16 LLM-driven Knowledge Enhancement for Multimodal Cancer Survival Prediction Chenyu Zhao et.al. 2512.14594 null
2025-12-16 Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer Adarsha Shrestha et.al. 2512.14585 null
2025-12-16 Pairwise Comparison for Bias Identification and Quantification Fabian Haak et.al. 2512.14565 null
2025-12-16 Polypersona: Persona-Grounded LLM for Synthetic Survey Responses Tejaswani Dash et.al. 2512.14562 null
2025-12-16 Agreement Between Large Language Models and Human Raters in Essay Scoring: A Research Synthesis Hongli Li et.al. 2512.14561 null
2025-12-16 VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models Nguyen Tien Dong et.al. 2512.14554 null
2025-12-16 Dual Language Models: Balancing Training Efficiency and Overfitting Resilience David Samuel et.al. 2512.14549 null
2025-12-16 VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse Ying Nie et.al. 2512.14531 null
2025-12-16 RecGPT-V2 Technical Report Chao Yi et.al. 2512.14503 null
2025-12-16 C-ing Clearly: Enhanced Binary Code Explanations using C code Teodor Poncu et.al. 2512.14500 null
2025-12-15 Beyond surface form: A pipeline for semantic analysis in Alzheimer’s Disease detection from spontaneous speech Dylan Phelps et.al. 2512.13685 null
2025-12-15 AgentIAD: Tool-Augmented Single-Agent for Industrial Anomaly Detection Junwen Miao et.al. 2512.13671 null
2025-12-15 A Scientific Reasoning Model for Organic Synthesis Procedure Generation Guoqing Liu et.al. 2512.13668 null
2025-12-15 RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics Enshen Zhou et.al. 2512.13660 null
2025-12-15 Embedding-Based Rankings of Educational Resources based on Learning Outcome Alignment: Benchmarking, Expert Validation, and Learner Performance Mohammadreza Molavi et.al. 2512.13658 null
2025-12-15 Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation Richard J. Young et.al. 2512.13655 null
2025-12-15 Large-Language Memorization During the Classification of United States Supreme Court Cases John E. Ortega et.al. 2512.13654 null
2025-12-15 MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning Haoyu Fu et.al. 2512.13636 null
2025-12-15 StutterFuse: Mitigating Modality Collapse in Stuttering Detection with Jaccard-Weighted Metric Learning and Gated Fusion Guransh Singh et.al. 2512.13632 null
2025-12-15 Temporal Tokenization Strategies for Event Sequence Modeling with Large Language Models Zefang Liu et.al. 2512.13618 null
2025-12-15 Do-Undo: Generating and Reversing Physical Actions in Vision-Language Models Shweta Mahajan et.al. 2512.13609 null
2025-12-15 Textual Gradients are a Flawed Metaphor for Automatic Prompt Optimization Daniel Melcer et.al. 2512.13598 null
2025-12-15 ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding Jia-Nan Li et.al. 2512.13586 null
2025-12-15 Reproducing and Dissecting Denoising Language Models for Speech Recognition Dorian Koch et.al. 2512.13576 null
2025-12-15 MMhops-R1: Multimodal Multi-hop Reasoning Tao Zhang et.al. 2512.13573 null
2025-12-15 Janus: Disaggregating Attention and Experts for Scalable MoE Inference Zhexiang Zhang et.al. 2512.13525 null
2025-12-15 Fine-tuned LLM-based Code Migration Framework Oleg Grynets et.al. 2512.13515 null
2025-12-15 MedCEG: Reinforcing Verifiable Medical Reasoning with Critical Evidence Graph Linjie Mu et.al. 2512.13510 null
2025-12-15 SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block Skipping Yu-Chen Lu et.al. 2512.13494 null
2025-12-15 From Zipf’s Law to Neural Scaling through Heaps’ Law and Hilberg’s Hypothesis Łukasz Dębowski et.al. 2512.13491 null
2025-12-12 Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective Etienne Boursier et.al. 2512.11784 null
2025-12-12 Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously Andrew Adiletta et.al. 2512.11783 null
2025-12-12 Multiscale Causal Geometric Deep Learning for Modeling Brain Structure Chengzhi Xia et.al. 2512.11738 null
2025-12-12 Speculative Decoding Speed-of-Light: Optimal Lower Bounds via Branching Random Walks Sergey Pankratov et.al. 2512.11718 null
2025-12-12 Evaluating Cooperative Resilience in Multiagent Systems: A Comparison Between Humans and LLMs Manuela Chacon-Chamorro et.al. 2512.11689 null
2025-12-12 Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing Xu Zhang et.al. 2512.11680 null
2025-12-12 Bridging Streaming Continual Learning via In-Context Large Tabular Models Afonso Lourenço et.al. 2512.11668 null
2025-12-12 From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews Brenda Nogueira et.al. 2512.11661 null
2025-12-12 Architecting Large Action Models for Human-in-the-Loop Intelligent Robots Kanisorn Sangchai et.al. 2512.11620 null
2025-12-12 LLM tools in the prediction of the stability of perovskite solar cells S. Frenkel et.al. 2512.11615 null
2025-12-12 Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols Björn Deiseroth et.al. 2512.11614 null
2025-12-12 AI Benchmark Democratization and Carpentry Gregor von Laszewski et.al. 2512.11588 null
2025-12-12 In-Context Learning for Seismic Data Processing Fabian Fuchs et.al. 2512.11575 null
2025-12-12 Visualizing token importance for black-box language models Paulius Rauba et.al. 2512.11573 null
2025-12-12 Extending a Parliamentary Corpus with MPs’ Tweets: Automatic Annotation and Evaluation Using MultiParTweet Mevlüt Bagci et.al. 2512.11567 null
2025-12-12 Say it or AI it: Evaluating Hands-Free Text Correction in Virtual Reality Ziming Li et.al. 2512.11564 null
2025-12-12 DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry Zhenyang Cai et.al. 2512.11558 null
2025-12-12 PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration Yifan Zhang et.al. 2512.11550 null
2025-12-12 AI-MASLD Metabolic Dysfunction and Information Steatosis of Large Language Models in Unstructured Clinical Narratives Yuan Shen et.al. 2512.11544 null
2025-12-12 HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning Yiqing Yang et.al. 2512.11534 null
2025-12-11 Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving Jiawei Yang et.al. 2512.10947 null
2025-12-11 VL-JEPA: Joint Embedding Predictive Architecture for Vision-language Delong Chen et.al. 2512.10942 null
2025-12-11 BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models Shengao Wang et.al. 2512.10932 null
2025-12-11 Asynchronous Reasoning: Training-Free Interactive Thinking LLMs George Yakushev et.al. 2512.10931 null
2025-12-11 FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos Yulu Gan et.al. 2512.10927 null
2025-12-11 SparseSwaps: Tractable LLM Pruning Mask Refinement at Scale Max Zimmer et.al. 2512.10922 null
2025-12-11 Multi-Granular Node Pruning for Circuit Discovery Muhammad Umair Haider et.al. 2512.10903 null
2025-12-11 LLMs Can Assist with Proposal Selection at Large User Facilities Lijie Ding et.al. 2512.10895 null
2025-12-11 DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance Peiying Zhang et.al. 2512.10894 null
2025-12-11 PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction Brandon Smock et.al. 2512.10888 null
2025-12-11 Computational emotion analysis with multimodal LLMs: Current evidence on an emerging methodological opportunity Hauke Licht et.al. 2512.10882 null
2025-12-11 Guided Transfer Learning for Discrete Diffusion Models Julian Kleutgens et.al. 2512.10877 null
2025-12-11 From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models Zongzhao Li et.al. 2512.10867 null
2025-12-11 MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence Jingli Lin et.al. 2512.10863 null
2025-12-11 Scaling Behavior of Discrete Diffusion Language Models Dimitri von Rütte et.al. 2512.10858 null
2025-12-11 Large Language Models for Superconductor Discovery Suman Itani et.al. 2512.10847 null
2025-12-11 LabelFusion: Learning to Fuse LLMs and Transformer Classifiers for Robust Text Classification Michael Schlee et.al. 2512.10793 null
2025-12-11 The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality Aileen Cheng et.al. 2512.10791 null
2025-12-11 Natural Language Interface for Firewall Configuration F. Taghiyev et.al. 2512.10789 null
2025-12-11 Developing and Evaluating a Large Language Model-Based Automated Feedback System Grounded in Evidence-Centered Design for Supporting Physics Problem Solving Holger Maus et.al. 2512.10785 null
2025-12-10 ReViSE: Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning Xinyu Liu et.al. 2512.09924 null
2025-12-10 Supervised learning pays attention Erin Craig et.al. 2512.09912 null
2025-12-10 Efficient Continual Learning in Neural Machine Translation: A Low-Rank Adaptation Approach Salvador Carrión et.al. 2512.09910 null
2025-12-10 VisualActBench: Can VLMs See and Act like a Human? Daoan Zhang et.al. 2512.09907 null
2025-12-10 SCOPE: Language Models as One-Time Teacher for Hierarchical Planning in Text Environments Haoye Lu et.al. 2512.09897 null
2025-12-10 Exploring Protein Language Model Architecture-Induced Biases for Antibody Comprehension Mengren et.al. 2512.09894 null
2025-12-10 Provably Learning from Modern Language Models via Low Logit Rank Noah Golowich et.al. 2512.09892 null
2025-12-10 HPM-KD: Hierarchical Progressive Multi-Teacher Framework for Knowledge Distillation and Efficient Model Compression Gustavo Coelho Haase et.al. 2512.09886 null
2025-12-10 Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs Pius Horn et.al. 2512.09874 null
2025-12-10 FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning Khurram Khalil et.al. 2512.09872 null
2025-12-10 MedForget: Hierarchy-Aware Multimodal Unlearning Testbed for Medical AI Fengli Wu et.al. 2512.09867 null
2025-12-10 UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving Hao Lu et.al. 2512.09864 null
2025-12-10 Mitigating Social Bias in English and Urdu Language Models Using PRM-Guided Candidate Selection and Sequential Refinement Muneeb Ur Raheem Khan et.al. 2512.09854 null
2025-12-10 ChronusOmni: Improving Time Awareness of Omni Large Language Models Yijing Chen et.al. 2512.09841 null
2025-12-10 LLMs in Interpreting Legal Documents Simone Corbo et.al. 2512.09830 null
2025-12-10 RIFT: A Scalable Methodology for LLM Accelerator Fault Assessment using Reinforcement Learning Khurram Khalil et.al. 2512.09829 null
2025-12-10 DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation Zhizhong Wang et.al. 2512.09814 null
2025-12-10 M3Net: A Multi-Metric Mixture of Experts Network Digital Twin with Graph Neural Networks Blessed Guda et.al. 2512.09797 null
2025-12-10 DeepSeek’s WEIRD Behavior: The cultural alignment of Large Language Models and the effects of prompt language and cultural prompting James Luther et.al. 2512.09772 null
2025-12-10 Defining Cost Function of Steganography with Large Language Models Hanzhou Wu et.al. 2512.09769 null
2025-12-09 Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs Angela van Sprang et.al. 2512.08923 null
2025-12-09 Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration Jin Hyeon Kim et.al. 2512.08922 null
2025-12-09 Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training Jakub Krajewski et.al. 2512.08894 null
2025-12-09 Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders Guangzhi Xiong et.al. 2512.08892 null
2025-12-09 No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers Damiano Marsili et.al. 2512.08889 null
2025-12-09 AI Didn’t Start the Fire: Examining the Stack Exchange Moderator and Contributor Strike Yiwei Wu et.al. 2512.08884 null
2025-12-09 SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing Aysim Toker et.al. 2512.08881 null
2025-12-09 When Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation Joshua Ward et.al. 2512.08875 null
2025-12-09 SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA Jing Zhang et.al. 2512.08867 null
2025-12-09 Tri-Bench: Stress-Testing VLM Reliability on Spatial Reasoning under Camera Tilt and Object Interference Amit Bendkhale et.al. 2512.08860 null
2025-12-09 InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models Hongyuan Tao et.al. 2512.08829 null
2025-12-09 Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning Yi Zhang et.al. 2512.08820 null
2025-12-09 Ask, Answer, and Detect: Role-Playing LLMs for Personality Detection with Question-Conditioned Mixture-of-Experts Yifan Lyu et.al. 2512.08814 null
2025-12-09 PrivTune: Efficient and Privacy-Preserving Fine-Tuning of Large Language Models via Device-Cloud Collaboration Yi Liu et.al. 2512.08809 null
2025-12-09 Can TabPFN Compete with GNNs for Node Classification via Graph Tabularization? Jeongwhan Choi et.al. 2512.08798 null
2025-12-09 A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs Mahmoud Srewa et.al. 2512.08786 null
2025-12-09 Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages David Samuel et.al. 2512.08777 null
2025-12-09 De novo generation of functional terpene synthases using TpsGPT Hamsini Ramanathan et.al. 2512.08772 null
2025-12-09 A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows Eranga Bandara et.al. 2512.08769 null
2025-12-09 Financial News Summarization: Can extractive methods still offer a true alternative to LLMs? Nicolas Reche et.al. 2512.08764 null
2025-12-08 Relational Visual Similarity Thao Nguyen et.al. 2512.07833 null
2025-12-08 Do Generalisation Results Generalise? Matteo Boglioni et.al. 2512.07832 null
2025-12-08 Provable Long-Range Benefits of Next-Token Prediction Xinyuan Cao et.al. 2512.07818 null
2025-12-08 Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach Hua Yang et.al. 2512.07814 null
2025-12-08 LLM Use for Mental Health: Crowdsourcing Users’ Sentiment-based Perspectives and Values from Social Discussions Lingyao Li et.al. 2512.07797 null
2025-12-08 Large Causal Models from Large Language Models Sridhar Mahadevan et.al. 2512.07796 null
2025-12-08 ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Nearchos Potamitis et.al. 2512.07795 null
2025-12-08 Automating High Energy Physics Data Analysis with LLM-Powered Agents Eli Gendreau-Distler et.al. 2512.07785 null
2025-12-08 On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models Charlie Zhang et.al. 2512.07783 null
2025-12-08 GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory Jiaxu Liu et.al. 2512.07782 null
2025-12-08 Distribution Matching Variational AutoEncoder Sen Ye et.al. 2512.07778 null
2025-12-08 Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives? Karin de Langis et.al. 2512.07777 null
2025-12-08 RL-MTJail: Reinforcement Learning for Automated Black-Box Multi-Turn Jailbreaking of Large Language Models Xiqiao Xiong et.al. 2512.07761 null
2025-12-08 SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery Meng Cao et.al. 2512.07733 null
2025-12-08 SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination Sangha Park et.al. 2512.07730 null
2025-12-08 Privacy Practices of Browser Agents Alisha Ukani et.al. 2512.07725 null
2025-12-08 In-Context and Few-Shots Learning for Forecasting Time Series Data based on Large Language Models Saroj Gopali et.al. 2512.07705 null
2025-12-08 HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs Sujoy Nath et.al. 2512.07687 null
2025-12-08 When Large Language Models Do Not Work: Online Incivility Prediction through Graph Neural Networks Zihan Chen et.al. 2512.07684 null
2025-12-08 Depth-Wise Activation Steering for Honest Language Models Gracjan Góral et.al. 2512.07667 null
2025-12-05 Enhancing Retrieval-Augmented Generation with Entity Linking for Educational Platforms Francesco Granata et.al. 2512.05967 null
2025-12-05 EditThinker: Unlocking Iterative Reasoning for Any Image Editor Hongyu Li et.al. 2512.05965 null
2025-12-05 Training-Time Action Conditioning for Efficient Real-Time Chunking Kevin Black et.al. 2512.05964 null
2025-12-05 M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG David Anugraha et.al. 2512.05959 null
2025-12-05 MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution Sara Patel et.al. 2512.05958 null
2025-12-05 SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models Haowen Liu et.al. 2512.05955 null
2025-12-05 SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code Shima Imani et.al. 2512.05954 null
2025-12-05 Trusted AI Agents in the Cloud Teofil Bodea et.al. 2512.05951 null
2025-12-05 TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models Shima Imani et.al. 2512.05943 null
2025-12-05 Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding Zhiyuan Jiang et.al. 2512.05941 null
2025-12-05 Adsorption energies are necessary but not sufficient to identify good catalysts Shahana Chatterjee et.al. 2512.05938 null
2025-12-05 Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech Xuanru Zhou et.al. 2512.05933 null
2025-12-05 Physically-Based Simulation of Automotive LiDAR L. Dudzik et.al. 2512.05932 null
2025-12-05 PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation Shima Imani et.al. 2512.05930 null
2025-12-05 LLM Harms: A Taxonomy and Discussion Kevin Chen et.al. 2512.05929 null
2025-12-05 KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity Damien Lesens et.al. 2512.05916 null
2025-12-05 From Text to Returns: Using Large Language Models for Mutual Fund Portfolio Optimization and Risk-Adjusted Allocation Abrar Hossain Mufakir Qamar Ansari Haziq Jeelani Monia Digra Fayeq Jeelani Syed et.al. 2512.05907 null
2025-12-05 SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations Wenhao Yan et.al. 2512.05905 null
2025-12-05 Euclid Quick Data Release (Q1). From simulations to sky: Advancing machine-learning lens detection with real Euclid data Euclid Collaboration et.al. 2512.05899 null
2025-12-05 A Machine Learning Framework for Predicting Glass-Forming Ability in Ternary Alloy Systems Fatemeh Mahmoudi et.al. 2512.05895 null
2025-12-05 Bootstrapping Fuzzers for Compilers of Low-Resource Language Dialects Using Language Models Sairam Vaidya et.al. 2512.05887 null
2025-12-05 InstructMPC: A Human-LLM-in-the-Loop Framework for Context-Aware Power Grid Control Ruixiang Wu et.al. 2512.05876 null
2025-12-05 Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework Tasnimul Hassan et.al. 2512.05863 null
2025-12-05 VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack Shiji Zhao et.al. 2512.05853 null
2025-12-05 Using Large Language Models to Create Personalized Networks From Therapy Sessions Clarissa W. Ong et.al. 2512.05836 null
2025-12-05 Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling Saurav Jha et.al. 2512.05809 null
2025-12-05 Mechanistic Interpretability of Antibody Language Models Using SAEs Rebonto Haque et.al. 2512.05794 null
2025-12-04 Value Gradient Guidance for Flow Matching Alignment Zhen Liu et.al. 2512.05116 null
2025-12-04 DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation Dongzhi Jiang et.al. 2512.05112 null
2025-12-04 ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning Shengyuan Ding et.al. 2512.05111 null
2025-12-04 STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models Feng Xu et.al. 2512.05107 null
2025-12-04 NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation Yu Zeng et.al. 2512.05106 null
2025-12-04 Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning Purbesh Mitra et.al. 2512.05105 null
2025-12-04 TV2TV: A Unified Framework for Interleaved Language and Video Generation Xiaochuang Han et.al. 2512.05103 null
2025-12-04 Structured Document Translation via Format Reinforcement Learning Haiyue Song et.al. 2512.05100 null
2025-12-04 From Generated Human Videos to Physically Plausible Robot Trajectories James Ni et.al. 2512.05094 null
2025-12-04 Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark Haobo Yuan et.al. 2512.05091 null
2025-12-04 David vs. Goliath: Can Small Models Win Big with Agentic AI in Hardware Design? Shashwat Shankar et.al. 2512.05073 null
2025-12-04 Multi-LLM Collaboration for Medication Recommendation Huascar Sanchez et.al. 2512.05066 null
2025-12-04 Personalizing Agent Privacy Decisions via Logical Entailment James Flemings et.al. 2512.05065 null
2025-12-04 Debt, Growth, and the Carbon Lock-In Silvia Montagnania et.al. 2512.05063 null
2025-12-04 Arbitrage: Efficient Reasoning via Advantage-Aware Speculation Monishwaran Maheswaran et.al. 2512.05033 null
2025-12-04 HTR-ConvText: Leveraging Convolution and Textual Information for Handwritten Text Recognition Pham Thach Thanh Truc et.al. 2512.05021 null
2025-12-04 Generative Neural Video Compression via Video Diffusion Prior Qi Mao et.al. 2512.05016 null
2025-12-04 Factuality and Transparency Are All RAG Needs! Self-Explaining Contrastive Evidence Re-ranking Francielle Vargas et.al. 2512.05012 null
2025-12-04 Evolutionary Architecture Search through Grammar-Based Sequence Alignment Adri Gómez Martín et.al. 2512.04992 null
2025-12-04 Influence of Object Affordance on Action Language Understanding: Evidence from Dynamic Causal Modeling Analysis Supriya Bordoloi et.al. 2512.04989 null
2025-12-03 Unique Lives, Shared World: Learning from Single-Life Videos Tengda Han et.al. 2512.04085 null
2025-12-03 SimFlow: Simplified and End-to-End Training of Latent Normalizing Flows Qinyu Zhao et.al. 2512.04084 null
2025-12-03 PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design Jiazhe Wei et.al. 2512.04082 null
2025-12-03 SkillFactory: Self-Distillation For Learning Cognitive Behaviors Zayne Sprague et.al. 2512.04072 null
2025-12-03 SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL Siyi Chen et.al. 2512.04069 null
2025-12-03 Eval Factsheets: A Structured Framework for Documenting AI Evaluations Florian Bordes et.al. 2512.04062 null
2025-12-03 Stable Signer: Hierarchical Sign Language Generative Model Sen Fang et.al. 2512.04048 null
2025-12-03 MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking Yizhou Zhao et.al. 2512.04044 null
2025-12-03 RELIC: Interactive Video World Model with Long-Horizon Memory Yicong Hong et.al. 2512.04040 null
2025-12-03 Jina-VLM: Small Multilingual Vision Language Model Andreas Koukounas et.al. 2512.04032 null
2025-12-03 Large Language Models for Limited Noisy Data: A Gravitational Wave Identification Study Yixuan Li et.al. 2512.04031 null
2025-12-03 Teaching Using Immersion - Explaining Magnetism and Eclipses in a Planetarium Dome Patricia H Reiff et.al. 2512.04027 null
2025-12-03 Ultra-lightweight Neural Video Representation Compression Ho Man Kwan et.al. 2512.04019 null
2025-12-03 AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving Ying Wang et.al. 2512.04013 null
2025-12-03 Learning to Comparison-Shop Jie Tang et.al. 2512.04009 null
2025-12-03 Physics-Embedded Gaussian Process for Traffic State Estimation Yanlin Chen et.al. 2512.04004 null
2025-12-03 Training-Free Policy Violation Detection via Activation-Space Whitening in LLMs Oren Rachmil et.al. 2512.03994 null
2025-12-03 DIQ-H: Evaluating Hallucination Persistence in VLMs Under Temporal Visual Degradation Zexin Lin et.al. 2512.03992 null
2025-12-03 Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models Taido Purason et.al. 2512.03989 null
2025-12-03 DirectDrag: High-Fidelity, Mask-Free, Prompt-Free Drag-based Image Editing via Readout-Guided Feature Alignment Sheng-Hao Liao et.al. 2512.03981 null
2025-12-02 CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models Minkyung Kwon et.al. 2512.03045 null
2025-12-02 OneThinker: All-in-one Reasoning Model for Image and Video Kaituo Feng et.al. 2512.03043 null
2025-12-02 ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation Mengchen Zhang et.al. 2512.03036 null
2025-12-02 The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models Saeid Jamshidi et.al. 2512.03026 null
2025-12-02 LORE: A Large Generative Model for Search Relevance Chenji Lu et.al. 2512.03025 null
2025-12-02 TokenPowerBench: Benchmarking the Power Consumption of LLM Inference Chenxu Niu et.al. 2512.03024 null
2025-12-02 Unrolled Networks are Conditional Probability Flows in MRI Reconstruction Kehan Qi et.al. 2512.03020 null
2025-12-02 Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge Hamid Dadkhahi et.al. 2512.03019 null
2025-12-02 Instant Video Models: Universal Adapters for Stabilizing Image-Based Networks Matthew Dutson et.al. 2512.03014 null
2025-12-02 In-Context Sync-LoRA for Portrait Video Editing Sagi Polaczek et.al. 2512.03013 null
2025-12-02 From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars? Dawei Li et.al. 2512.03005 null
2025-12-02 Invasive Context Engineering to Control Large Language Models Thomas Rivasseau et.al. 2512.03001 null
2025-12-02 GraphFusion3D: Dynamic Graph Attention Convolution with Adaptive Cross-Modal Transformer for 3D Object Detection Md Sohag Mia et.al. 2512.02991 null
2025-12-02 Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic Muyu Pan et.al. 2512.02987 null
2025-12-02 ProteinPNet: Prototypical Part Networks for Concept Learning in Spatial Proteomics Louis McConnell et.al. 2512.02983 null
2025-12-02 InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration Zhongyu Yang et.al. 2512.02981 null
2025-12-02 Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities Yuan Xiong et.al. 2512.02973 null
2025-12-02 Lumos: Let there be Language Model System Certification Isha Chaudhary et.al. 2512.02966 null
2025-12-02 The Evolutionary Ecology of Software: Constraints, Innovation, and the AI Disruption Sergi Valverde et.al. 2512.02953 null
2025-12-02 AutoNeural: Co-Designing Vision-Language Models for NPU Inference Wei Chen et.al. 2512.02924 null
2025-12-01 ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation Chenyang Gu et.al. 2512.02013 null
2025-12-01 Improved Mean Flows: On the Challenges of Fastforward Generative Models Zhengyang Geng et.al. 2512.02012 null
2025-12-01 Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling Jack Cook et.al. 2512.02010 null
2025-12-01 AirSim360: A Panoramic Simulation Platform within Drone View Xian Ge et.al. 2512.02009 null
2025-12-01 The Art of Scaling Test-Time Compute for Large Language Models Aradhye Agarwal et.al. 2512.02008 null
2025-12-01 AlignSAE: Concept-Aligned Sparse Autoencoders Minglai Yang et.al. 2512.02004 null
2025-12-01 LLM-Driven Corrective Robot Operation Code Generation with Static Text-Based Simulation Wenhao Wang et.al. 2512.02002 null
2025-12-01 LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess Sai Kolasani et.al. 2512.01992 null
2025-12-01 PAI-Bench: A Comprehensive Benchmark For Physical AI Fengzhe Zhou et.al. 2512.01989 null
2025-12-01 Orientational lineage memory and mechanical ordering during diffusion-limited growth Ilias-Marios Sarris et.al. 2512.01981 null
2025-12-01 Low-Rank Prehab: Preparing Neural Networks for SVD Compression Haoran Qin et.al. 2512.01980 null
2025-12-01 Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback Aiden Yiliu Li et.al. 2512.01979 null
2025-12-01 Consistent Synthetic Sequences Unlock Structural Diversity in Fully Atomistic De Novo Protein Design Danny Reidenbach et.al. 2512.01976 null
2025-12-01 SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning Xu Zhang et.al. 2512.01975 null
2025-12-01 From Atomic to Composite: Reinforcement Learning Enables Generalization in Complementary Reasoning Sitao Cheng et.al. 2512.01970 null
2025-12-01 Learned-Rule-Augmented Large Language Model Evaluators Jie Meng et.al. 2512.01958 null
2025-12-01 KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference Sai Gokhale et.al. 2512.01953 null
2025-12-01 GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment Haoyang He et.al. 2512.01952 null
2025-12-01 Order and shape dependence of mechanical relaxation in proliferating active matter Jonas Isensee et.al. 2512.01950 null
2025-12-01 Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models Zhongyu Yang et.al. 2512.01949 null
2025-11-28 Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models Muhammad Maaz et.al. 2511.23478 null
2025-11-28 Video-CoM: Interactive Video Reasoning via Chain of Manipulations Hanoona Rasheed et.al. 2511.23477 null
2025-11-28 Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction Bao Shu et.al. 2511.23476 null
2025-11-28 ThetaEvolve: Test-time Learning on Open Problems Yiping Wang et.al. 2511.23473 null
2025-11-28 Visual Generation Tuning Jiahao Guo et.al. 2511.23469 null
2025-11-28 The Price of Progress: Algorithmic Efficiency and the Falling Cost of AI Inference Hans Gundlach et.al. 2511.23455 null
2025-11-28 Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent Jianzhe Lin et.al. 2511.23436 null
2025-11-28 The EBLM Project XVIII. 3D Obliquities of Five Low-Mass Eclipsing Binaries Becca Spejcher et.al. 2511.23430 null
2025-11-28 Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model Junshu Tang et.al. 2511.23429 null
2025-11-28 Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities Aayush Garg et.al. 2511.23408 null
2025-11-28 LFM2 Technical Report Alexander Amini et.al. 2511.23404 null
2025-11-28 Quantized-Tinyllava: a new multimodal foundation model enables efficient split learning Jiajun Guo et.al. 2511.23402 null
2025-11-28 MegaChat: A Synthetic Persian Q&A Dataset for High-Quality Sales Chatbot Evaluation Mahdi Rahmani et.al. 2511.23397 null
2025-11-28 Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization Jian Li et.al. 2511.23391 null
2025-11-28 Improving motor imagery decoding methods for an EEG-based mobile brain-computer interface in the context of the 2024 Cybathlon Isabel Whiteley Tscherniak et.al. 2511.23384 null
2025-11-28 Identifying bars in galaxies using machine learning Rajit Shrivastava et.al. 2511.23383 null
2025-11-28 DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline Rui Zhang et.al. 2511.23377 null
2025-11-28 Optimizing Multimodal Language Models through Attention-based Interpretability Alexander Sergeev et.al. 2511.23375 null
2025-11-28 Scaling HuBERT for African Languages: From Base to Large and XL Antoine Caubrière et.al. 2511.23370 null
2025-11-28 Towards Improving Interpretability of Language Model Generation through a Structured Knowledge Discovery Approach Shuqi Liu et.al. 2511.23335 null
2025-11-26 Revisiting Generalization Across Difficulty Levels: It’s Not So Easy Yeganeh Kordi et.al. 2511.21692 null
2025-11-26 TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos Seungjae Lee et.al. 2511.21690 null
2025-11-26 ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration Hongjin Su et.al. 2511.21689 null
2025-11-26 G $^2$ VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning Wenbo Hu et.al. 2511.21688 null
2025-11-26 Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework Dong Wang et.al. 2511.21686 null
2025-11-26 Mean-field Modelling of Moiré Materials: A User’s Guide with Selected Applications to Twisted Bilayer Graphene Yves H. Kwan et.al. 2511.21683 null
2025-11-26 DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving Fengze Yu et.al. 2511.21669 null
2025-11-26 Escaping the Verifier: Learning to Reason via Demonstrations Locke Cai et.al. 2511.21667 null
2025-11-26 Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models Naifu Zhang et.al. 2511.21663 null
2025-11-26 Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following Tianyi Xiong et.al. 2511.21662 null
2025-11-26 Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO Daniel R. Jiang et.al. 2511.21638 null
2025-11-26 Qwen3-VL Technical Report Shuai Bai et.al. 2511.21631 null
2025-11-26 The author is dead, but what if they never lived? A reception experiment on Czech AI- and human-authored poetry Anna Marklová et.al. 2511.21629 null
2025-11-26 TAGFN: A Text-Attributed Graph Dataset for Fake News Detection in the Age of LLMs Kay Liu et.al. 2511.21624 null
2025-11-26 Automated Protein Motif Localization using Concept Activation Vectors in Protein Language Model Embedding Space Ahmad Shamail et.al. 2511.21614 null
2025-11-26 Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining Dongyang Fan et.al. 2511.21613 null
2025-11-26 Auxiliary Metrics Help Decoding Skill Neurons in the Wild Yixiu Zhao et.al. 2511.21610 null
2025-11-26 Beyond Accuracy: An Empirical Study of Uncertainty Estimation in Imputation Zarin Tahia Hossain et.al. 2511.21607 null
2025-11-26 ReSAM: Refine, Requery, and Reinforce: Self-Prompting Point-Supervised Segmentation for Remote Sensing Images M. Naseer Subhani et.al. 2511.21606 null
2025-11-26 Tidal forces around the Letelier-Alencar cloud of strings black hole Marcos V. de S. Silva et.al. 2511.21604 null
2025-11-25 RubricRL: Simple Generalizable Rewards for Text-to-Image Generation Xuelu Feng et.al. 2511.20651 null
2025-11-25 MedROV: Towards Real-Time Open-Vocabulary Detection Across Diverse Medical Imaging Modalities Tooba Tehreem Sheikh et.al. 2511.20650 null
2025-11-25 LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight Yunze Man et.al. 2511.20648 null
2025-11-25 Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization Tahira Kazimi et.al. 2511.20647 null
2025-11-25 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding Xiaoye Wang et.al. 2511.20646 null
2025-11-25 Vision-Language Memory for Spatial Reasoning Zuntao Liu et.al. 2511.20644 null
2025-11-25 Concept-Aware Batch Sampling Improves Language-Image Pretraining Adhiraj Ghosh et.al. 2511.20643 null
2025-11-25 Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition Wei Tang et.al. 2511.20641 null
2025-11-25 Latent Collaboration in Multi-Agent Systems Jiaru Zou et.al. 2511.20639 null
2025-11-25 Reinforcing Action Policies by Prophesying Jiahui Zhang et.al. 2511.20633 null
2025-11-25 MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models Chieh-Yun Chen et.al. 2511.20629 null
2025-11-25 Fighting AI with AI: Leveraging Foundation Models for Assuring AI-Enabled Safety-Critical Systems Anastasia Mavridou et.al. 2511.20627 null
2025-11-25 ROOT: Robust Orthogonalized Optimizer for Neural Network Training Wei He et.al. 2511.20626 null
2025-11-25 Copyright Detection in Large Language Models: An Ethical Approach to Generative AI Development David Szczecina et.al. 2511.20623 null
2025-11-25 DiFR: Inference Verification Despite Nondeterminism Adam Karvonen et.al. 2511.20621 null
2025-11-25 The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment Ziheng Ouyang et.al. 2511.20614 null
2025-11-25 Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning Panayiotis Danassis et.al. 2511.20613 null
2025-11-25 Sparse-to-Field Reconstruction via Stochastic Neural Dynamic Mode Decomposition Yujin Kim et.al. 2511.20612 null
2025-11-25 Adaptive Hopfield Network: Rethinking Similarities in Associative Memory Shurong Wang et.al. 2511.20609 null
2025-11-25 On Evaluating LLM Alignment by Evaluating LLMs as Judges Yixin Liu et.al. 2511.20604 null
2025-11-24 VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection Qiang Wang et.al. 2511.19436 null
2025-11-24 Are Image-to-Video Models Good Zero-Shot Image Editors? Zechuan Zhang et.al. 2511.19435 null
2025-11-24 Mixture of Horizons in Action Chunking Dong Jing et.al. 2511.19433 null
2025-11-24 Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution Dingkang Liang et.al. 2511.19430 null
2025-11-24 Prompt Less, Smile More: MTP with Semantic Engineering in Lieu of Prompt Engineering Jayanaka L. Dantanarayana et.al. 2511.19427 null
2025-11-24 Beyond Protein Language Models: An Agentic LLM Framework for Mechanistic Enzyme Design Bruno Jacob et.al. 2511.19423 null
2025-11-24 SLMFix: Leveraging Small Language Models for Error Fixing with Reinforcement Learning David Jiahao Fu et.al. 2511.19422 null
2025-11-24 Refractive neutrino masses in the solar DM halo: Can the dark-LMA solution be revived? Susobhan Chattopadhyay et.al. 2511.19420 null
2025-11-24 Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens Yiming Qin et.al. 2511.19418 null
2025-11-24 Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration James Y. Huang et.al. 2511.19417 null
2025-11-24 Learning Robust Social Strategies with Large Language Models Dereck Piche et.al. 2511.19405 null
2025-11-24 DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research Rulin Shao et.al. 2511.19399 null
2025-11-24 UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval Maroun Ayli et.al. 2511.19380 null
2025-11-24 LLM-Driven Stationarity-Aware Expert Demonstrations for Multi-Agent Reinforcement Learning in Mobile Systems Tianyang Duan et.al. 2511.19368 null
2025-11-24 An Anatomy Aware Hybrid Deep Learning Framework for Lung Cancer Tumor Stage Classification Saniah Kayenat Chowdhury et.al. 2511.19367 null
2025-11-24 Growing with the Generator: Self-paced GRPO for Video Generation Rui Li et.al. 2511.19356 null
2025-11-24 Leveraging LLMs for reward function design in reinforcement learning control tasks Franklin Cardenoso et.al. 2511.19355 null
2025-11-24 Scalable Parameter-Light Spectral Method for Clustering Short Text Embeddings with a Cohesion-Based Evaluation Metric Nikita Neveditsin et.al. 2511.19350 null
2025-11-24 Revisiting Feedback Models for HyDE Nour Jedidi et.al. 2511.19349 null
2025-11-24 Annotation-Free Class-Incremental Learning Hari Chandana Kuchibhotla et.al. 2511.19344 null
2025-11-21 RynnVLA-002: A Unified Vision-Language-Action and World Model Jun Cen et.al. 2511.17502 null
2025-11-21 Harnessing Data from Clustered LQR Systems: Personalized and Collaborative Policy Optimization Vinay Kanakeri et.al. 2511.17489 null
2025-11-21 Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models Mark Endo et.al. 2511.17487 null
2025-11-21 Counterfactual World Models via Digital Twin-conditioned Video Diffusion Yiqing Shen et.al. 2511.17481 null
2025-11-21 Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition Ayhan Kucukmanisa et.al. 2511.17477 null
2025-11-21 Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards Zhen Wang et.al. 2511.17473 null
2025-11-21 Moving superfluids in the rotating universe Jose Beltrán Jiménez et.al. 2511.17472 null
2025-11-21 PersonaAgent with GraphRAG: Community-Aware Knowledge Graphs for Personalized LLM Siqi Liang et.al. 2511.17467 null
2025-11-21 MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models Yuqi Li et.al. 2511.17448 null
2025-11-21 REMSA: An LLM Agent for Foundation Model Selection in Remote Sensing Binger Chen et.al. 2511.17442 null
2025-11-21 RoboCOIN: An Open-Sourced Bimanual Robotic Data COllection for INtegrated Manipulation Shihan Wu et.al. 2511.17441 null
2025-11-21 SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation Shrikant Kendre et.al. 2511.17432 null
2025-11-21 Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions Guilherme Coelho et.al. 2511.17429 null
2025-11-21 SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding Nikolay Nikolov et.al. 2511.17411 null
2025-11-21 That’s not natural: The Impact of Off-Policy Training Data on Probe Performance Nathalie Kirch et.al. 2511.17408 null
2025-11-21 Beyond Multiple Choice: A Hybrid Framework for Unifying Robust Evaluation and Verifiable Reasoning Training Yesheng Liu et.al. 2511.17405 null
2025-11-21 Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required? Sukwon Yun et.al. 2511.17400 null
2025-11-21 MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment Huangbiao Xu et.al. 2511.17397 null
2025-11-21 MorphSeek: Fine-grained Latent Representation-Level Policy Optimization for Deformable Image Registration Runxun Zhang et.al. 2511.17392 null
2025-11-21 Selective Rotary Position Embedding Sajad Movahedi et.al. 2511.17388 null
2025-11-20 Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation Ziyu Guo et.al. 2511.16671 null
2025-11-20 Learning to Think Fast and Slow for Visual Language Models Chenyu Lin et.al. 2511.16670 null
2025-11-20 Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO Junhao Cheng et.al. 2511.16669 null
2025-11-20 V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models Yang Luo et.al. 2511.16668 null
2025-11-20 Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter Qinghao Hu et.al. 2511.16665 null
2025-11-20 Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs Ali Taghibakhshi et.al. 2511.16664 null
2025-11-20 Cognitive Foundations for Reasoning and Their Manifestation in LLMs Priyanka Kargupta et.al. 2511.16660 null
2025-11-20 Comparison of Text-Based and Image-Based Retrieval in Multimodal Retrieval Augmented Generation Large Language Model Systems Elias Lumer et.al. 2511.16654 null
2025-11-20 Evolution Strategies at the Hyperscale Bidipta Sarkar et.al. 2511.16652 null
2025-11-20 InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy Yang Tian et.al. 2511.16651 null
2025-11-20 Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs Wei-Cheng Tseng et.al. 2511.16639 null
2025-11-20 SurvAgent: Hierarchical CoT-Enhanced Case Banking and Dichotomy-Based Multi-Agent System for Multimodal Survival Prediction Guolin Huang et.al. 2511.16635 null
2025-11-20 MedBayes-Lite: Bayesian Uncertainty Quantification for Safe Clinical Decision Support Elias Hossain et.al. 2511.16625 null
2025-11-20 SAM 3D: 3Dfy Anything in Images SAM 3D Team et.al. 2511.16624 null
2025-11-20 Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization Yi Zhang et.al. 2511.16602 null
2025-11-20 You Only Forward Once: An Efficient Compositional Judging Paradigm Tianlong Zhang et.al. 2511.16600 null
2025-11-20 TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding Boshen Xu et.al. 2511.16595 null
2025-11-20 Formal Abductive Latent Explanations for Prototype-Based Networks Jules Soria et.al. 2511.16588 null
2025-11-20 Integrating Symbolic Natural Language Understanding and Language Models for Word Sense Disambiguation Kexin Zhao et.al. 2511.16577 null
2025-11-20 POMA-3D: The Point Map Way to 3D Scene Understanding Ye Mao et.al. 2511.16567 null
2025-11-19 Think Visually, Reason Textually: Vision-Language Synergy in ARC Beichen Zhang et.al. 2511.15703 null
2025-11-19 Joint Semantic-Channel Coding and Modulation for Token Communications Jingkai Ying et.al. 2511.15699 null
2025-11-19 RescueLens: LLM-Powered Triage and Action on Volunteer Feedback for Food Rescue Naveen Raman et.al. 2511.15698 null
2025-11-19 The Impact of Quantization on Large Reasoning Model Reinforcement Learning Medha Kumar et.al. 2511.15694 null
2025-11-19 MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping Yushi Huang et.al. 2511.15690 null
2025-11-19 Walrus: A Cross-Domain Foundation Model for Continuum Dynamics Michael McCabe et.al. 2511.15684 null
2025-11-19 Quantum-Guided Test Case Minimization for LLM-Based Code Generation Huixiang Zhang et.al. 2511.15665 null
2025-11-19 VisPlay: Self-Evolving Vision-Language Models from Images Yicheng He et.al. 2511.15661 null
2025-11-19 Multi-Stage Residual-Aware Unsupervised Deep Learning Framework for Consistent Ultrasound Strain Elastography Shourov Joarder et.al. 2511.15640 null
2025-11-19 Hierarchical Semantic Tree Anchoring for CLIP-Based Class-Incremental Learning Tao Hu et.al. 2511.15633 null
2025-11-19 The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification Dante Francisco Wasmuht et.al. 2511.15622 null
2025-11-19 Lost in Vagueness: Towards Context-Sensitive Standards for Robustness Assessment under the EU AI Act Roberta Tamponi et.al. 2511.15620 null
2025-11-19 When to Think and When to Look: Uncertainty-Guided Lookback Jing Bi et.al. 2511.15613 null
2025-11-19 SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models Senyu Fei et.al. 2511.15605 null
2025-11-19 Graph Rewriting Language as a Platform for Quantum Diagrammatic Calculi Kayo Tei et.al. 2511.15581 null
2025-11-19 AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning Urjitkumar Patel et.al. 2511.15578 null
2025-11-19 A critical review of pre-post surveys designed to measure student epistemology in undergraduate science courses Kyriaki Chatzikyriakidou et.al. 2511.15575 null
2025-11-19 HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning Qihao Yang et.al. 2511.15574 null
2025-11-19 Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models Vikram K Suresh et.al. 2511.15573 null
2025-11-19 Computer-Use Agents as Judges for Generative User Interface Kevin Qinghong Lin et.al. 2511.15567 null
2025-11-18 ARC Is a Vision Problem! Keya Hu et.al. 2511.14761 null
2025-11-18 UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning Rui Tian et.al. 2511.14760 null
2025-11-18 $π^{*}_{0.6}$ : a VLA That Learns From Experience Ali Amin et.al. 2511.14759 null
2025-11-18 HMC: Learning Heterogeneous Meta-Control for Contact-Rich Loco-Manipulation Lai Wei et.al. 2511.14756 null
2025-11-18 SparseST: Exploiting Data Sparsity in Spatiotemporal Modeling and Prediction Junfeng Wu et.al. 2511.14753 null
2025-11-18 Vision Large Language Models Are Good Noise Handlers in Engagement Analysis Alexander Vedernikov et.al. 2511.14749 null
2025-11-18 Look-Ahead Reasoning on Learning Platforms Haiqing Zhu et.al. 2511.14745 null
2025-11-18 Measuring AI Progress in Drug Discovery: A Reproducible Leaderboard for the Tox21 Challenge Antonia Ebner et.al. 2511.14744 null
2025-11-18 LAUD: Integrating Large Language Models with Active Learning for Unlabeled Data Tzu-Hsuan Chou et.al. 2511.14738 null
2025-11-18 When AI Democratizes Exploitation: LLM-Assisted Strategic Manipulation of Fair Division Algorithms Priyanka Verma et.al. 2511.14722 null
2025-11-18 AdamHD: Decoupled Huber Decay Regularization for Language Model Pre-Training Fu-Ming Guo et.al. 2511.14721 null
2025-11-18 Strategic Innovation Management in the Age of Large Language Models Market Intelligence, Adaptive R&D, and Ethical Governance Raha Aghaei et.al. 2511.14709 null
2025-11-18 Subword Tokenization Strategies for Kurdish Word Embeddings Ali Salehi et.al. 2511.14696 null
2025-11-18 Near-Lossless Model Compression Enables Longer Context Inference in DNA Large Language Models Rui Zhu et.al. 2511.14694 null
2025-11-18 Talk, Snap, Complain: Validation-Aware Multimodal Expert Framework for Fine-Grained Customer Grievances Rishu Kumar Singh et.al. 2511.14693 null
2025-11-18 Attention via Synaptic Plasticity is All You Need: A Biologically Inspired Spiking Neuromorphic Transformer Kallol Mondal et.al. 2511.14691 null
2025-11-18 Ground Truth Generation for Multilingual Historical NLP using LLMs Clovis Gladstone et.al. 2511.14688 null
2025-11-18 Encoding and Understanding Astrophysical Information in Large Language Model-Generated Summaries Kiera McCormick et.al. 2511.14685 null
2025-11-18 SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction Biaojie Zeng et.al. 2511.14684 null
2025-11-18 Quadratic Term Correction on Heaps’ Law Oscar Fontanelli et.al. 2511.14683 null
2025-11-17 Scaling Spatial Intelligence with Multimodal Foundation Models Zhongang Cai et.al. 2511.13719 null
2025-11-17 TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone Xunjie Wang et.al. 2511.13717 null
2025-11-17 TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models Harold Haodong Chen et.al. 2511.13704 null
2025-11-17 Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation Sofia Jamil et.al. 2511.13689 null
2025-11-17 Protein Secondary Structure Prediction Using 3D Graphs and Relation-Aware Message Passing Transformers Disha Varshney et.al. 2511.13685 null
2025-11-17 Training-Free Multi-View Extension of IC-Light for Textual Position-Aware Scene Relighting Jiangnan Ye et.al. 2511.13684 null
2025-11-17 Person-AI Bidirectional Fit - A Proof-Of-Concept Case Study Of Augmented Human-Ai Symbiosis In Management Decision-Making Process Agnieszka Bieńkowska et.al. 2511.13670 null
2025-11-17 Ontology-Driven Model-to-Model Transformation of Workflow Specifications Francisco Abreu et.al. 2511.13661 null
2025-11-17 Why is “Chicago” Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues Jiaming Qu et.al. 2511.13658 null
2025-11-17 Weight-sparse transformers have interpretable circuits Leo Gao et.al. 2511.13653 null
2025-11-17 Part-X-MLLM: Part-aware 3D Multimodal Large Language Model Chunshi Wang et.al. 2511.13647 null
2025-11-17 Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly? Chunqiu Steven Xia et.al. 2511.13646 null
2025-11-17 CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding Shrenik Patel et.al. 2511.13644 null
2025-11-17 Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures Haohui Wang et.al. 2511.13640 null
2025-11-17 Beyond Mimicry: Preference Coherence in LLMs Luhan Mikaelson et.al. 2511.13630 null
2025-11-17 CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product Kaiwen Xue et.al. 2511.13626 null
2025-11-17 Tissue Aware Nuclei Detection and Classification Model for Histopathology Images Kesi Xu et.al. 2511.13615 null
2025-11-17 P1: Mastering Physics Olympiads with Reinforcement Learning Jiacheng Chen et.al. 2511.13612 null
2025-11-17 Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation Hao Wang et.al. 2511.13590 null
2025-11-17 Adaptive Multi-Scale Integration Unlocks Robust Cell Annotation in Histopathology Images Yinuo Xu et.al. 2511.13586 null
2025-11-14 Optimizing Mixture of Block Attention Guangxuan Xiao et.al. 2511.11571 null
2025-11-14 PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning Afra Feyza Akyürek et.al. 2511.11562 null
2025-11-14 Human-AI collaborative autonomous synthesis with pulsed laser deposition for remote epitaxy Asraful Haque et.al. 2511.11558 null
2025-11-14 Multistability of Self-Attention Dynamics in Transformers Claudio Altafini et.al. 2511.11553 null
2025-11-14 DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding Dawei Zhu et.al. 2511.11552 null
2025-11-14 Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping Dena Mujtaba et.al. 2511.11551 null
2025-11-14 Accurate models for recoil velocity distribution in black hole mergers with comparable to extreme mass-ratios and their astrophysical implications Tousif Islam et.al. 2511.11536 null
2025-11-14 Terrain Costmap Generation via Scaled Preference Conditioning Luisa Mao et.al. 2511.11529 null
2025-11-14 Bridging Hidden States in Vision-Language Models Benjamin Fein-Ashley et.al. 2511.11526 null
2025-11-14 CVChess: A Deep Learning Framework for Converting Chessboard Images to Forsyth-Edwards Notation Luthira Abeykoon et.al. 2511.11522 null
2025-11-14 Scalable Policy Evaluation with Video World Models Wei-Cheng Tseng et.al. 2511.11520 null
2025-11-14 Experience-Guided Adaptation of Inference-Time Reasoning Strategies Adam Stein et.al. 2511.11519 null
2025-11-14 W2S-AlignTree: Weak-to-Strong Inference-Time Alignment for Large Language Models via Monte Carlo Tree Search Zhenyu Ding et.al. 2511.11518 null
2025-11-14 Collaborative Representation Learning for Alignment of Tactile, Language, and Vision Modalities Yiyun Zhou et.al. 2511.11512 null
2025-11-14 FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models Yonatan Dukler et.al. 2511.11505 null
2025-11-14 PAS : Prelim Attention Score for Detecting Object Hallucinations in Large Vision–Language Models Nhat Hoang-Xuan et.al. 2511.11502 null
2025-11-14 Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation Mohamad Amin Mohamadi et.al. 2511.11500 null
2025-11-14 ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation Kaishen Wang et.al. 2511.11483 null
2025-11-14 Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective Nhat Chung et.al. 2511.11478 null
2025-11-14 Context-aware Adaptive Visualizations for Critical Decision Making Angela Lopez-Cardona et.al. 2511.11476 null
2025-11-13 Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling Jiahao Wang et.al. 2511.10648 null
2025-11-13 ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference Yesheng Liang et.al. 2511.10645 null
2025-11-13 Black-Box On-Policy Distillation of Large Language Models Tianzhu Ye et.al. 2511.10643 null
2025-11-13 Instella: Fully Open Language Models with Stellar Performance Jiang Liu et.al. 2511.10628 null
2025-11-13 Querying Labeled Time Series Data with Scenario Programs Edward Kim et.al. 2511.10627 null
2025-11-13 SSR: Socratic Self-Refine for Large Language Model Reasoning Haizhou Shi et.al. 2511.10621 null
2025-11-13 Know Your Limits: Entropy Estimation Modeling for Compression and Generalization Benjamin L. Badger et.al. 2511.10618 null
2025-11-13 Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals Shruti Singh Baghel et.al. 2511.10615 null
2025-11-13 Regular Games – an Automata-Based General Game Playing Language Radosław Miernik et.al. 2511.10593 null
2025-11-13 Textual understanding boost in the WikiRace Raman Ebrahimi et.al. 2511.10585 null
2025-11-13 Evaluating Prompting Strategies with MedGemma for Medical Order Extraction Abhinand Balachandran et.al. 2511.10583 null
2025-11-13 DESS: DeBERTa Enhanced Syntactic-Semantic Aspect Sentiment Triplet Extraction Vishal Thenuwara et.al. 2511.10577 null
2025-11-13 Towards Emotionally Intelligent and Responsible Reinforcement Learning Garapati Keerthana et.al. 2511.10573 null
2025-11-13 Belief Net: A Filter-Based Framework for Learning Hidden Markov Models from Observations Reginald Zhiyan Chen et.al. 2511.10571 null
2025-11-13 Impact of Layer Norm on Memorization and Generalization in Transformers Rishi Singhal et.al. 2511.10566 null
2025-11-13 Maximizing Efficiency of Dataset Compression for Machine Learning Potentials With Information Theory Benjamin Yu et.al. 2511.10561 null
2025-11-13 OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Haosong Peng et.al. 2511.10560 null
2025-11-13 URaG: Unified Retrieval and Generation in Multimodal LLMs for Efficient Long Document Understanding Yongxin Shi et.al. 2511.10552 null
2025-11-13 Computing the Formal and Institutional Boundaries of Contemporary Genre and Literary Fiction Natasha Johnson et.al. 2511.10546 null
2025-11-13 Bytes of a Feather: Personality and Opinion Alignment Effects in Human-AI Interaction Maximilian Eder et.al. 2511.10544 null
2025-11-10 Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs Zhongyang Li et.al. 2511.07419 null
2025-11-10 Using Vision Language Models as Closed-Loop Symbolic Planners for Robotic Applications: A Control-Theoretic Perspective Hao Wang et.al. 2511.07410 null
2025-11-10 SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations Manon Berriche et.al. 2511.07405 null
2025-11-10 SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards Hunar Batra et.al. 2511.07403 null
2025-11-10 ConvFill: Model Collaboration for Responsive Conversational Voice Agents Vidya Srinivas et.al. 2511.07397 null
2025-11-10 C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning Antonios Valkanas et.al. 2511.07396 null
2025-11-10 Surgical Agent Orchestration Platform for Voice-directed Patient Data Interaction Hyeryun Park et.al. 2511.07392 null
2025-11-10 Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence Sean McLeish et.al. 2511.07384 null
2025-11-10 Retriv at BLP-2025 Task 2: Test-Driven Feedback-Guided Framework for Bangla-to-Python Code Generation K M Nafi Asib et.al. 2511.07382 null
2025-11-10 Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains Pingjie Wang et.al. 2511.07380 null
2025-11-10 Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization Yu Huang et.al. 2511.07378 null
2025-11-10 Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training Dake Bu et.al. 2511.07372 null
2025-11-10 Consistency Is Not Always Correct: Towards Understanding the Role of Exploration in Post-Training Reasoning Dake Bu et.al. 2511.07368 null
2025-11-10 Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection Vaibhav Mavi et.al. 2511.07364 null
2025-11-10 DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas Zhen Wang et.al. 2511.07338 null
2025-11-10 Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis Yash Mittal et.al. 2511.07329 null
2025-11-10 When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs Shaowen Wang et.al. 2511.07318 null
2025-11-10 RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Zhiyuan Zeng et.al. 2511.07317 null
2025-11-10 ACE-ICD: Acronym Expansion As Data Augmentation For Automated ICD Coding Tuan-Dung Le et.al. 2511.07311 null
2025-11-10 VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models Ying Cheng et.al. 2511.07299 null
2025-11-07 Visual Spatial Tuning Rui Yang et.al. 2511.05491 null
2025-11-07 A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher? Md. Abdul Awal et.al. 2511.05476 null
2025-11-07 Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection Xian-Hong Huang et.al. 2511.05474 null
2025-11-07 SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models Jingxuan Xu et.al. 2511.05459 null
2025-11-07 Steering Language Models with Weight Arithmetic Constanza Fierro et.al. 2511.05408 null
2025-11-07 Large Language Models for Explainable Threat Intelligence Tiago Dinis et.al. 2511.05406 null
2025-11-07 PreResQ-R1: Towards Fine-Grained Rank-and-Score Reinforcement Learning for Visual Quality Assessment via Preference-Response Disentangled Policy Optimization Zehui Feng et.al. 2511.05393 null
2025-11-07 TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework Chao Zhang et.al. 2511.05385 null
2025-11-07 Connectomics Informed by Large Language Models Elinor Thompson et.al. 2511.05383 null
2025-11-07 Dense Motion Captioning Shiyao Xu et.al. 2511.05369 null
2025-11-07 ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations Amr Gomaa et.al. 2511.05359 null
2025-11-07 Turning Adversaries into Allies: Reversing Typographic Attacks for Multimodal E-Commerce Product Retrieval Janet Jenq et.al. 2511.05325 null
2025-11-07 Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE Firoj Ahmmed Patwary et.al. 2511.05324 null
2025-11-07 What Are the Facts? Automated Extraction of Court-Established Facts from Criminal-Court Opinions Klára Bendová et.al. 2511.05320 null
2025-11-07 $\mathbf{S^2LM}$ : Towards Semantic Steganography via Large Language Models Huanqi Wu et.al. 2511.05319 null
2025-11-07 Attention and Compression is all you need for Controllably Efficient Language Models Jatin Prakash et.al. 2511.05313 null
2025-11-07 Cleaning Maintenance Logs with LLM Agents for Improved Predictive Maintenance Valeriu Dimidov et.al. 2511.05311 null
2025-11-07 Listening Between the Lines: Decoding Podcast Narratives with Language Modeling Shreya Gupta et.al. 2511.05310 null
2025-11-07 Code Review Automation using Retrieval Augmented Generation Qianru Meng et.al. 2511.05302 null
2025-11-07 LiveStar: Live Streaming Assistant for Real-World Online Video Understanding Zhenyu Yang et.al. 2511.05299 null
2025-11-06 SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding Ellis Brown et.al. 2511.04668 null
2025-11-06 SAFe-Copilot: Unified Shared Autonomy Framework Phat Nguyen et.al. 2511.04664 null
2025-11-06 VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks Yu Feng et.al. 2511.04662 null
2025-11-06 Benchmark Designers Should “Train on the Test Set” to Expose Exploitable Non-Visual Shortcuts Ellis Brown et.al. 2511.04655 null
2025-11-06 Logit-Entropy Adaptive Stopping Heuristic for Efficient Chain-of-Thought Reasoning Mohammad Atif Quamar et.al. 2511.04654 null
2025-11-06 Optimal Inference Schedules for Masked Diffusion Models Sitan Chen et.al. 2511.04647 null
2025-11-06 When retrieval outperforms generation: Dense evidence retrieval for scalable fake news detection Alamgir Munir Qazi et.al. 2511.04643 null
2025-11-06 PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning Yicheng Xiao et.al. 2511.04601 null
2025-11-06 Neural Computation Without Slots: Steps Towards Biologically Plausible Memory and Attention in Natural and Artificial Intelligence Shaunak Bhandarkar et.al. 2511.04593 null
2025-11-06 Question the Questions: Auditing Representation in Online Deliberative Processes Soham De et.al. 2511.04588 null
2025-11-06 ARETE: an R package for Automated REtrieval from TExt with large language models Vasco V. Branco et.al. 2511.04573 null
2025-11-06 Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Jingqi Tong et.al. 2511.04570 null
2025-11-06 Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment Tao Lin et.al. 2511.04555 null
2025-11-06 LLM-as-a-Judge: Toward World Models for Slate Recommendation Systems Baptiste Bonin et.al. 2511.04541 null
2025-11-06 From Model to Breach: Towards Actionable LLM-Generated Vulnerabilities Reporting Cyril Vallez et.al. 2511.04538 null
2025-11-06 Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics Amir Zur et.al. 2511.04527 null
2025-11-06 Large Language Models for Cyber Security Raunak Somani et.al. 2511.04508 null
2025-11-06 Modeling Clinical Uncertainty in Radiology Reports: from Explicit Uncertainty Markers to Implicit Reasoning Pathways Paloma Rabaey et.al. 2511.04506 null
2025-11-06 RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG Joshua Gao et.al. 2511.04502 null
2025-11-06 Large language models replicate and predict human cooperation across experiments in game theory Andrea Cera Palatsi et.al. 2511.04500 null
2025-11-05 Disentangled Concepts Speak Louder Than Words:Explainable Video Action Recognition Jongseo Lee et.al. 2511.03725 null
2025-11-05 Outbidding and Outbluffing Elite Humans: Mastering Liar’s Poker via Self-Play and Reinforcement Learning Richard Dewey et.al. 2511.03724 null
2025-11-05 LLM-enhanced Air Quality Monitoring Interface via Model Context Protocol Yu-Erh Pan et.al. 2511.03706 null
2025-11-05 Do Androids Dream of Unseen Puppeteers? Probing for a Conspiracy Mindset in Large Language Models Francesco Corso et.al. 2511.03699 null
2025-11-05 AnaFlow: Agentic LLM-based Workflow for Reasoning-Driven Explainable and Sample-Efficient Analog Circuit Sizing Mohsen Ahmadzadeh et.al. 2511.03697 null
2025-11-05 Whisper Leak: a side-channel attack on Large Language Models Geoff McDonald et.al. 2511.03675 null
2025-11-05 Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology Thomas Souverain et.al. 2511.03641 null
2025-11-05 Towards Transparent Stance Detection: A Zero-Shot Approach Using Implicit and Explicit Interpretability Apoorva Upadhyaya et.al. 2511.03635 null
2025-11-05 LiveTradeBench: Seeking Real-World Alpha with Large Language Models Haofei Yu et.al. 2511.03628 null
2025-11-05 PerfDojo: Automated ML Library Generation for Heterogeneous Architectures Andrei Ivanov et.al. 2511.03586 null
2025-11-05 ASVRI-Legal: Fine-Tuning LLMs with Retrieval Augmented Generation for Enhanced Legal Regulation One Octadion et.al. 2511.03563 null
2025-11-05 AILA–First Experiments with Localist Language Models Joachim Diederich et.al. 2511.03559 null
2025-11-05 MultiZebraLogic: A Multilingual Logical Reasoning Benchmark Sofie Helene Bruun et.al. 2511.03553 null
2025-11-05 Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding Ziv Nevo et.al. 2511.03549 null
2025-11-05 SOLVE-Med: Specialized Orchestration for Leading Vertical Experts across Medical Specialties Roberta Di Marino et.al. 2511.03542 null
2025-11-05 U2F: Encouraging SWE-Agent to Seize Novelty without Losing Feasibility Wencheng Ye et.al. 2511.03517 null
2025-11-05 One Battle After Another: Probing LLMs’ Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework Qi Jia et.al. 2511.03508 null
2025-11-05 BanglaSTEM: A Parallel Corpus for Technical Domain Bangla-English Translation Kazi Reyazul Hasan et.al. 2511.03498 null
2025-11-05 RAGBoost: Efficient Retrieval-Augmented Generation with Accuracy-Preserving Context Reuse Yinsicheng Jiang et.al. 2511.03475 null
2025-11-05 Towards Scalable Web Accessibility Audit with MLLMs as Copilots Ming Gu et.al. 2511.03471 null
2025-11-04 Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything Huawei Lin et.al. 2511.02834 null
2025-11-04 In Good GRACEs: Principled Teacher Selection for Knowledge Distillation Abhishek Panigrahi et.al. 2511.02833 null
2025-11-04 TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System Yanjie Ze et.al. 2511.02832 null
2025-11-04 Can LLMs subtract numbers? Mayank Jobanputra et.al. 2511.02795 null
2025-11-04 When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning Chenyu Zhang et.al. 2511.02794 null
2025-11-04 When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought Yiyang Zhou et.al. 2511.02779 null
2025-11-04 XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations Shichao Fan et.al. 2511.02776 null
2025-11-04 Dynamic Reflections: Probing Video Representations with Text Alignment Tyler Zhu et.al. 2511.02767 null
2025-11-04 LLM-Supported Formal Knowledge Representation for Enhancing Control Engineering Content with an Interactive Semantic Layer Julius Fiedler et.al. 2511.02759 null
2025-11-04 ConMeZO: Adaptive Descent-Direction Sampling for Gradient-Free Finetuning of Large Language Models Lejs Deen Behric et.al. 2511.02757 null
2025-11-04 Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning Bowen Jin et.al. 2511.02755 null
2025-11-04 AI Diffusion in Low Resource Language Countries Amit Misra et.al. 2511.02752 null
2025-11-04 Agentic World Modeling for 6G: Near-Real-Time Generative State-Space Reasoning Farhad Rezazadeh et.al. 2511.02748 null
2025-11-04 CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents Jiayu Liu et.al. 2511.02734 null
2025-11-04 LLEXICORP: End-user Explainability of Convolutional Neural Networks Vojtěch Kůr et.al. 2511.02720 null
2025-11-04 ReleaseEval: A Benchmark for Evaluating Language Models in Automated Release Note Generation Qianru Meng et.al. 2511.02713 null
2025-11-04 VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models Zhicheng Zhang et.al. 2511.02712 null
2025-11-04 Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs Georgios Tzannetos et.al. 2511.02690 null
2025-11-04 Optimal Singular Damage: Efficient LLM Inference in Low Storage Regimes Mohammadsajad Alipour et.al. 2511.02681 null
2025-11-04 EasyTUS: A Comprehensive Framework for Fast and Accurate Table Union Search across Data Lakes Tim Otto et.al. 2511.02674 null
2025-11-03 Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training Dayuan Fu et.al. 2510.27630 null
2025-11-03 InnovatorBench: Evaluating Agents’ Ability to Conduct Innovative LLM Research Yunze Wu et.al. 2510.27598 null
2025-10-31 Continuous Autoregressive Language Models Chenze Shao et.al. 2510.27688 null
2025-10-31 Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals Xiangyu Fan et.al. 2510.27684 null
2025-10-31 PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting Danyal Maqbool et.al. 2510.27680 null
2025-10-31 On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection Md Abdul Hannan et.al. 2510.27675 null
2025-10-31 RDMA Point-to-Point Communication for LLM Systems Nandor Licker et.al. 2510.27656 null
2025-10-31 SpecAttn: Speculating Sparse Attention Harsh Shah et.al. 2510.27641 null
2025-10-31 Validity Is What You Need Sebastian Benthall et.al. 2510.27628 null
2025-10-31 Visual Backdoor Attacks on MLLM Embodied Decision Making via Contrastive Trigger Learning Qiusi Zhan et.al. 2510.27623 null
2025-10-31 VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation Heng Ping et.al. 2510.27617 null
2025-10-31 ORGEval: Graph-Theoretic Evaluation of LLMs in Optimization Modeling Zhuohan Wang et.al. 2510.27610 null
2025-10-31 Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning Yuhong Liu et.al. 2510.27606 null
2025-10-31 AMD MI300X GPU Performance Analysis Chandrish Ambati et.al. 2510.27583 null
2025-10-31 MARAG-R1: Beyond Single Retriever via Reinforcement-Learned Multi-Tool Agentic Retrieval Qi Luo et.al. 2510.27569 null
2025-10-31 CodeAlignBench: Assessing Code Generation Models on Developer-Preferred Code Adjustments Forough Mehralian et.al. 2510.27565 null
2025-10-31 Multilingual BERT language model for medical tasks: Evaluation on domain-specific adaptation and cross-linguality Yinghao Luo et.al. 2510.27552 null
2025-10-31 Mechanics of Learned Reasoning 1: TempoBench, A Benchmark for Interpretable Deconstruction of Reasoning System Performance Nikolaus Holzer et.al. 2510.27544 null
2025-10-31 DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models Malik H. Altakrori et.al. 2510.27543 null
2025-10-31 Patient-Centered Summarization Framework for AI Clinical Summarization: A Mixed-Methods Design Maria Lizarazo Jimenez et.al. 2510.27535 null
2025-10-30 Masked Diffusion Captioning for Visual Feature Learning Chao Feng et.al. 2510.26799 null
2025-10-30 Defeating the Training-Inference Mismatch via FP16 Penghui Qi et.al. 2510.26788 null
2025-10-30 ChartAB: A Benchmark for Chart Grounding & Dense Alignment Aniruddh Bansal et.al. 2510.26781 null
2025-10-30 SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models Anushka Sivakumar et.al. 2510.26769 null
2025-10-30 AMO-Bench: Large Language Models Still Struggle in High School Math Competitions Shengnan An et.al. 2510.26768 null
2025-10-30 ProfOlaf: Semi-Automated Tool for Systematic Literature Reviews Martim Afonso et.al. 2510.26750 null
2025-10-30 ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference Zixu Shen et.al. 2510.26730 null
2025-10-30 Unveiling Intrinsic Text Bias in Multimodal Large Language Models through Attention Key-Space Analysis Xinhan Zheng et.al. 2510.26721 null
2025-10-30 Value Drifts: Tracing Value Alignment During LLM Post-Training Mehar Bhatia et.al. 2510.26707 null
2025-10-30 Delegated Authorization for Agents Constrained to Semantic Task-to-Scope Matching Majed El Helou et.al. 2510.26702 null
2025-10-30 Using Copilot Agent Mode to Automate Library Migration: A Quantitative Assessment Aylton Almeida et.al. 2510.26699 null
2025-10-30 The End of Manual Decoding: Towards Truly End-to-End Language Models Zhichao Wang et.al. 2510.26697 null
2025-10-30 Kimi Linear: An Expressive, Efficient Attention Architecture Kimi Team et.al. 2510.26692 null
2025-10-30 LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits Amir Reza Mirzaei et.al. 2510.26690 null
2025-10-30 Evontree: Ontology Rule-Guided Self-Evolution of Large Language Models Mingchen Tu et.al. 2510.26683 null
2025-10-30 The Era of Agentic Organization: Learning to Organize with Language Models Zewen Chi et.al. 2510.26658 null
2025-10-30 Accelerating mathematical research with language models: A case study of an interaction with GPT-5-Pro on a convex analysis problem Adil Salim et.al. 2510.26647 null
2025-10-30 All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles Sayed Pedram Haeri Boroujeni et.al. 2510.26641 null
2025-10-30 Stitch: Step-by-step LLM Guided Tutoring for Scratch Yuan Si et.al. 2510.26634 null
2025-10-30 Low-Altitude UAV-Carried Movable Antenna for Joint Wireless Power Transfer and Covert Communications Chuang Zhang et.al. 2510.26628 null
2025-10-30 PairUni: Pairwise Training for Unified Multimodal Language Models Jiani Zheng et.al. 2510.25682 null
2025-10-30 Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks Davide Romano et.al. 2510.25623 null
2025-10-29 Gaperon: A Peppered English-French Generative Language Model Suite Nathan Godey et.al. 2510.25771 null
2025-10-29 E-Scores for (In)Correctness Assessment of Generative Model Outputs Guneet S. Dhillon et.al. 2510.25770 null
2025-10-29 Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models Sriram Balasubramaniam et.al. 2510.25766 null
2025-10-29 DiagramEval: Evaluating LLM-Generated Diagrams via Graphs Chumeng Liang et.al. 2510.25761 null
2025-10-29 Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks Xu Zheng et.al. 2510.25760 null
2025-10-29 TheraMind: A Strategic and Adaptive Agent for Longitudinal Psychological Counseling He Hu et.al. 2510.25758 null
2025-10-29 Scaling Latent Reasoning via Looped Language Models Rui-Jie Zhu et.al. 2510.25741 null
2025-10-29 The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework Aakriti Shah et.al. 2510.25732 null
2025-10-29 Interpreting LLMs as Credit Risk Classifiers: Do Their Feature Explanations Align with Classical ML? Saeed AlMarri et.al. 2510.25701 null
2025-10-29 Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering Agents Jiayi Kuang et.al. 2510.25694 null
2025-10-29 ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents Tianyu Yang et.al. 2510.25668 null
2025-10-29 User Misconceptions of LLM-Based Conversational Programming Assistants Gabrielle O’Brien et.al. 2510.25662 null
2025-10-29 EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis Yusheng Liao et.al. 2510.25628 null
2025-10-29 Are Language Models Efficient Reasoners? A Perspective from Logic Programming Andreas Opedal et.al. 2510.25626 null
2025-10-29 FARSIQA: Faithful and Advanced RAG System for Islamic Question Answering Mohammad Aghajani Asl et.al. 2510.25621 null
2025-10-29 Don’t Blind Your VLA: Aligning Visual Representations for OOD Generalization Nikita Kachaev et.al. 2510.25616 null
2025-10-29 INT v.s. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats Mengzhao Chen et.al. 2510.25602 null
2025-10-29 Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry Run Peng et.al. 2510.25595 null
2025-10-29 OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Ziyou Hu et.al. 2510.24636 null
2025-10-28 Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance Yujie Wei et.al. 2510.24711 null
2025-10-28 ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games? Shuqing Li et.al. 2510.24706 null
2025-10-28 Tongyi DeepResearch Technical Report Tongyi DeepResearch Team et.al. 2510.24701 null
2025-10-28 Greedy Sampling Is Provably Efficient for RLHF Di Wu et.al. 2510.24700 null
2025-10-28 WebLeaper: Empowering Efficiency and Efficacy in WebAgent via Enabling Info-Rich Seeking Zhengwei Tao et.al. 2510.24697 null
2025-10-28 AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis Xuanzhong Chen et.al. 2510.24695 null
2025-10-28 STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence Zihan Liu et.al. 2510.24693 null
2025-10-28 Dissecting Role Cognition in Medical LLMs via Neuronal Ablation Xun Liang et.al. 2510.24677 null
2025-10-28 Evolving Diagnostic Agents in a Virtual Clinical Environment Pengcheng Qiu et.al. 2510.24654 null
2025-10-28 Optimizing Retrieval for RAG via Reinforced Contrastive Learning Jiawei Zhou et.al. 2510.24652 null
2025-10-28 Advancing site-specific disease and pest management in precision agriculture: From reasoning-driven foundation models to adaptive, feedback-based learning Nitin Rai et.al. 2510.24650 null
2025-10-28 FunReason-MT Technical Report: Overcoming the Complexity Barrier in Multi-Turn Function Calling Zengzhuang Xu et.al. 2510.24645 null
2025-10-28 Relative Scaling Laws for LLMs William Held et.al. 2510.24626 null
2025-10-28 Zero-Shot Cross-Lingual Transfer using Prefix-Based Adaptation Snegha A et.al. 2510.24619 null
2025-10-28 Diffusion LLM with Native Variable Generation Lengths: Let [EOS] Lead the Way Yicun Yang et.al. 2510.24605 null
2025-10-28 ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization Guoxin Chen et.al. 2510.24592 null
2025-10-28 ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? Christine Ye et.al. 2510.24591 null
2025-10-28 Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives Gang Chen et.al. 2510.24551 null
2025-10-28 Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts Seyoung Song et.al. 2510.24541 null
2025-10-28 Multi-Agent Evolve: LLM Self-Improve through Co-evolution Yixing Chen et.al. 2510.23595 null
2025-10-28 PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection Yusu Qian et.al. 2510.23594 null
2025-10-27 PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity Yuqian Yuan et.al. 2510.23603 null
2025-10-27 Alita-G: Self-Evolving Generative Agent for Agent Generation Jiahao Qiu et.al. 2510.23601 null
2025-10-27 Think Twice: Branch-and-Rethink Reasoning Reward Model Yizhu Jiao et.al. 2510.23596 null
2025-10-27 Lightweight Robust Direct Preference Optimization Cheol Woo Kim et.al. 2510.23590 null
2025-10-27 FARMER: Flow AutoRegressive Transformer over Pixels Guangting Zheng et.al. 2510.23588 null
2025-10-27 A Survey of Data Agents: Emerging Paradigm or Overstated Hype? Yizhang Zhu et.al. 2510.23587 null
2025-10-27 RobotArena $\infty$ : Scalable Robot Benchmarking via Real-to-Sim Translation Yash Jangir et.al. 2510.23571 null
2025-10-27 EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT Baoqi Pei et.al. 2510.23569 null
2025-10-27 ReCode: Unify Plan and Action for Universal Granularity Control Zhaoyang Yu et.al. 2510.23564 null
2025-10-27 ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models Bohan Li et.al. 2510.23558 null
2025-10-27 Minimizing Human Intervention in Online Classification William Réveillard et.al. 2510.23557 null
2025-10-27 IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering Jieyong Kim et.al. 2510.23536 null
2025-10-27 Point Convergence of Nesterov’s Accelerated Gradient Method: An AI-Assisted Proof Uijeong Jang et.al. 2510.23513 null
2025-10-27 Deductive Chain-of-Thought Augmented Socially-aware Robot Navigation World Model Weizheng Wang et.al. 2510.23509 null
2025-10-27 Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier Hyeongseop Rha et.al. 2510.23506 null
2025-10-27 VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation Walid Bousselham et.al. 2510.23497 null
2025-10-27 Learning the PTM Code through a Coarse-to-Fine, Mechanism-Aware Framework Jingjie Zhang et.al. 2510.23492 null
2025-10-27 Learning to Reason Efficiently with Discounted Reinforcement Learning Alex Ayoub et.al. 2510.23486 null
2025-10-24 A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection Gaku Morio et.al. 2510.21679 null
2025-10-24 A Data-Centric Approach to Multilingual E-Commerce Product Search: Case Study on Query-Category and Query-Item Relevance Yabo Yin et.al. 2510.21671 null
2025-10-24 The Universal Landscape of Human Reasoning Qiguang Chen et.al. 2510.21623 null
2025-10-24 Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine Wenyi Wang et.al. 2510.21614 null
2025-10-24 Modest-Align: Data-Efficient Alignment for Vision-Language Models Jiaxiang Liu et.al. 2510.21606 null
2025-10-24 RETuning: Upgrading Inference-Time Scaling for Stock Movement Prediction with Large Language Models Xueyuan Lin et.al. 2510.21604 null
2025-10-24 From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene Mojca Brglez et.al. 2510.21575 null
2025-10-24 ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem Fangwen Wu et.al. 2510.21566 null
2025-10-24 Are the LLMs Capable of Maintaining at Least the Language Genus? Sandra Mitrović et.al. 2510.21561 null
2025-10-24 EU-Agent-Bench: Measuring Illegal Behavior of LLM Agents Under EU Law Ilija Lichkovski et.al. 2510.21524 null
2025-10-24 Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech Models Omer Moussa et.al. 2510.21520 null
2025-10-24 Head Pursuit: Probing Attention Specialization in Multimodal Transformers Lorenzo Basile et.al. 2510.21518 null
2025-10-24 Wisdom and Delusion of LLM Ensembles for Code Generation and Repair Fernando Vallecillos Ruiz et.al. 2510.21513 null
2025-10-24 Actionable Cybersecurity Notifications for Smart Homes: A User Study on the Role of Length and Complexity Victor Jüttner et.al. 2510.21508 null
2025-10-24 MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization Chenglong Wang et.al. 2510.21473 null
2025-10-24 Risk Management for Mitigating Benchmark Failure Modes: BenchRisk Sean McGregor et.al. 2510.21460 null
2025-10-24 SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots Adetayo Adebimpe et.al. 2510.21459 null
2025-10-24 ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models Federico Danieli et.al. 2510.21450 null
2025-10-24 MoniTor: Exploiting Large Language Models with Instruction for Online Video Anomaly Detection Shengtian Yang et.al. 2510.21449 null
2025-10-24 REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring Thanh Cong Ho et.al. 2510.21445 null
2025-10-23 KL-Regularized Reinforcement Learning is Designed to Mode Collapse Anthony GX-Chen et.al. 2510.20817 null
2025-10-23 Generative Reasoning Recommendation via LLMs Minjie Hong et.al. 2510.20815 null
2025-10-23 Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation Yuhan Liu et.al. 2510.20812 null
2025-10-23 On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text? Mingmeng Geng et.al. 2510.20810 null
2025-10-23 Video Prediction of Dynamic Physical Simulations With Pixel-Space Spatiotemporal Transformers Dean L Slack et.al. 2510.20807 null
2025-10-23 ARGenSeg: Image Segmentation with Autoregressive Image Generation Model Xiaolong Wang et.al. 2510.20803 null
2025-10-23 Simple Context Compression: Mean-Pooling and Multi-Ratio Training Yair Feldman et.al. 2510.20797 null
2025-10-23 A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text Alicia Sagae et.al. 2510.20782 null
2025-10-23 RAGRank: Using PageRank to Counter Poisoning in CTI LLM Pipelines Austin Jia et.al. 2510.20768 null
2025-10-23 Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations Lorenzo Stacchio et.al. 2510.20743 null
2025-10-23 Learning to Triage Taint Flows Reported by Dynamic Program Analysis in Node.js Packages Ronghao Ni et.al. 2510.20739 null
2025-10-23 Automated Extraction of Fluoropyrimidine Treatment and Treatment-Related Toxicities from Clinical Notes Using Natural Language Processing Xizhi Wu et.al. 2510.20727 null
2025-10-23 User Perceptions of Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios Xiaoyuan Wu et.al. 2510.20721 null
2025-10-23 Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models Xuyang Liu et.al. 2510.20707 null
2025-10-23 Structure-Conditional Minimum Bayes Risk Decoding Bryan Eikema et.al. 2510.20700 null
2025-10-23 Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward Jing Bi et.al. 2510.20696 null
2025-10-23 Exploring Large Language Models for Access Control Policy Synthesis and Summarization Adarsh Vatsa et.al. 2510.20692 null
2025-10-23 Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge Graphs Yanlin Song et.al. 2510.20691 null
2025-10-23 Neural Diversity Regularizes Hallucinations in Small Models Kushal Chakrabarti et.al. 2510.20690 null
2025-10-23 Bayesian Jammer Localization with a Hybrid CNN and Path-Loss Mixture of Experts Mariona Jaramillo-Civill et.al. 2510.20666 null
2025-10-23 Zhyper: Factorized Hypernetworks for Conditioned LLM Fine-Tuning M. H. I. Abdalla et.al. 2510.19733 null
2025-10-23 Fast Inference via Hierarchical Speculative Decoding Clara Mohri et.al. 2510.19705 null
2025-10-22 Semantic World Models Jacob Berg et.al. 2510.19818 null
2025-10-22 olmOCR 2: Unit Test Rewards for Document OCR Jake Poznanski et.al. 2510.19817 null
2025-10-22 Hubble: a Model Suite to Advance the Study of LLM Memorization Johnny Tian-Zheng Wei et.al. 2510.19811 null
2025-10-22 Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning Xichen Zhang et.al. 2510.19807 null
2025-10-22 The Art of Asking: Multilingual Prompt Optimization for Synthetic Data David Mora et.al. 2510.19806 null
2025-10-22 Forbidden Sidon subsets of perfect difference sets, featuring a human-assisted proof Boris Alexeev et.al. 2510.19804 null
2025-10-22 Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models Xiaozhen Qiao et.al. 2510.19802 null
2025-10-22 The Feasibility of Training Sovereign Language Models in the Global South: A Study of Brazil and Mexico Sandra Malagon et.al. 2510.19801 null
2025-10-22 Integrating Transparent Models, LLMs, and Practitioner-in-the-Loop: A Case of Nonprofit Program Evaluation Ji Ma et.al. 2510.19799 null
2025-10-22 Blackbox Model Provenance via Palimpsestic Membership Inference Rohith Kuditipudi et.al. 2510.19796 null
2025-10-22 On Controlled Change: Generative AI’s Impact on Professional Authority in Journalism Tomás Dodds et.al. 2510.19792 null
2025-10-22 ToolDreamer: Instilling LLM Reasoning Into Tool Retrievers Saptarshi Sengupta et.al. 2510.19791 null
2025-10-22 AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders Yuezhou Hu et.al. 2510.19779 null
2025-10-22 The Tail Tells All: Estimating Model-Level Membership Inference Vulnerability Without Reference Models Euodia Dodd et.al. 2510.19773 null
2025-10-22 SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration Xichen Zhang et.al. 2510.19767 null
2025-10-22 Top-P Masking for Cross Language Information Retrieval Joseph Casale et.al. 2510.19758 null
2025-10-22 Review of Tools for Zero-Code LLM Based Application Development Priyaranjan Pattnayak et.al. 2510.19747 null
2025-10-22 RLIE: Rule Generation with Logistic Regression, Iterative Refinement, and Evaluation for Large Language Models Yang Yang et.al. 2510.19698 null
2025-10-22 Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs Haochen Wang et.al. 2510.18876 null
2025-10-21 Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting Howard Chen et.al. 2510.18874 null
2025-10-21 DSI-Bench: A Benchmark for Dynamic Spatial Intelligence Ziang Zhang et.al. 2510.18873 null
2025-10-21 How Do LLMs Use Their Depth? Akshat Gupta et.al. 2510.18871 null
2025-10-21 LightMem: Lightweight and Efficient Memory-Augmented Generation Jizhan Fang et.al. 2510.18866 null
2025-10-21 EffiReasonTrans: RL-Optimized Reasoning for Code Translation Yanlin Wang et.al. 2510.18863 null
2025-10-21 Streamlining Acceptance Test Generation for Mobile Applications Through Large Language Models: An Industrial Case Study Pedro Luís Fonseca et.al. 2510.18861 null
2025-10-21 An Encoder-Decoder Foundation Chemical Language Model for Generative Polymer Design Harikrishna Sahu et.al. 2510.18860 null
2025-10-21 Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning Chenghao Zhu et.al. 2510.18849 null
2025-10-21 See the Text: From Tokenization to Visual Reading Ling Xing et.al. 2510.18840 null
2025-10-21 FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated Learning Yubin Zheng et.al. 2510.18837 null
2025-10-21 MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training Wenxuan Li et.al. 2510.18830 null
2025-10-21 Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework Yujie Xing et.al. 2510.18825 null
2025-10-21 Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring Shuxin Lin et.al. 2510.18817 null
2025-10-21 Integrating Large Language Models and Evaluating Student Outcomes in an Introductory Computer Science Course Annapurna Vadaparty et.al. 2510.18806 null
2025-10-21 FeClustRE: Hierarchical Clustering and Semantic Tagging of App Features from User Reviews Max Tiessler et.al. 2510.18799 null
2025-10-21 ShaRE your Data! Characterizing Datasets for LLM-based Requirements Engineering Quim Motger et.al. 2510.18787 null
2025-10-21 KAT-Coder Technical Report Zizheng Zhan et.al. 2510.18779 null
2025-10-21 Seg the HAB: Language-Guided Geospatial Algae Bloom Reasoning and Segmentation Patterson Hsieh et.al. 2510.18751 null
2025-10-21 Topoformer: brain-like topographic organization in Transformer language models through spatial querying and reweighting Taha Binhuraib et.al. 2510.18745 null
2025-10-21 Verifiable Accuracy and Abstention Rewards in Curriculum RL to Alleviate Lost-in-Conversation Ming Li et.al. 2510.18731 null
2025-10-21 HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models Sidhant Narula et.al. 2510.18728 null
2025-10-21 IF-VidCap: Can Video Caption Models Follow Instructions? Shihao Li et.al. 2510.18726 null
2025-10-21 SemiAdapt and SemiLoRA: Efficient Domain Adaptation for Transformer-based Low-Resource Language Translation with a Case Study on Irish Josh McGiff et.al. 2510.18725 null
2025-10-21 SSD: Spatial-Semantic Head Decoupling for Efficient Autoregressive Image Generation Siyong Jian et.al. 2510.18716 null
2025-10-21 Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options Joongkyu Lee et.al. 2510.18713 null
2025-10-21 Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents Yiqi Lin et.al. 2510.18703 null
2025-10-21 UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation Yibin Wang et.al. 2510.18701 null
2025-10-21 MLMA: Towards Multilingual with Mamba Based Architectures Mohamed Nabih Ali et.al. 2510.18684 null
2025-10-21 Exploring Membership Inference Vulnerabilities in Clinical Large Language Models Alexander Nemecek et.al. 2510.18674 null
2025-10-21 Reasoning Language Model Inference Serving Unveiled: An Empirical Study Qi Li et.al. 2510.18672 null
2025-10-21 Hardness of Learning Regular Languages in the Next Symbol Prediction Setting Satwik Bhattamishra et.al. 2510.18634 null
2025-10-21 Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views Zhangquan Chen et.al. 2510.18632 null
2025-10-21 VAR: Visual Attention Reasoning via Structured Search and Backtracking Wei Cai et.al. 2510.18619 null
2025-10-21 Evaluating Large Language Models in detecting Secrets in Android Apps Marco Alecci et.al. 2510.18601 null
2025-10-21 CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent Haojia Lin et.al. 2510.18596 null
2025-10-21 Tokencake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications Zhuohang Bian et.al. 2510.18586 null
2025-10-21 CLASP: Cost-Optimized LLM-based Agentic System for Phishing Detection Fouad Trad et.al. 2510.18585 null
2025-10-21 CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder Yongmin Lee et.al. 2510.18583 null
2025-10-21 The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability Zijie Xu et.al. 2510.18563 null
2025-10-21 Large language models for folktale type automation based on motifs: Cinderella case study Tjaša Arčon et.al. 2510.18561 null
2025-10-21 Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency Svetlana Maslenkova et.al. 2510.18556 null
2025-10-21 JAUNT: Joint Alignment of User Intent and Network State for QoE-centric LLM Tool Routing Enhan Li et.al. 2510.18550 null
2025-10-21 EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval Zebin Yang et.al. 2510.18546 null
2025-10-21 SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices Pan Zhou et.al. 2510.18544 null
2025-10-21 Noise-Conditioned Mixture-of-Experts Framework for Robust Speaker Verification Bin Gu et.al. 2510.18533 null
2025-10-21 LLMs as Sparse Retrievers:A Framework for First-Stage Product Search Hongru Song et.al. 2510.18527 null
2025-10-21 Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models Hanze Guo et.al. 2510.18526 null
2025-10-21 From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing Yushu Zhao et.al. 2510.18525 null
2025-10-21 Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models Sureyya Akin et.al. 2510.18515 null
2025-10-21 Identity-Aware Large Language Models require Cultural Reasoning Alistair Plum et.al. 2510.18510 null
2025-10-21 Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritization Osama Al Haddad et.al. 2510.18508 null
2025-10-21 Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation Wei-Chia Chang et.al. 2510.18502 null
2025-10-21 One Size Fits All? A Modular Adaptive Sanitization Kit (MASK) for Customizable Privacy-Preserving Phone Scam Detection Kangzhong Wang et.al. 2510.18493 null
2025-10-21 The Attribution Story of WhisperGate: An Academic Perspective Oleksandr Adamov et.al. 2510.18484 null
2025-10-21 StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking Haoran Zhang et.al. 2510.18483 null
2025-10-21 How Efficient Are Diffusion Language Models? A Critical Examination of Efficiency Evaluation Practices Han Peng et.al. 2510.18480 null
2025-10-21 LAFA: Agentic LLM-Driven Federated Analytics over Decentralized Data Sources Haichao Ji et.al. 2510.18477 null
2025-10-21 Probabilistic Modeling of Intentions in Socially Intelligent LLM Agents Feifan Xia et.al. 2510.18476 null
2025-10-21 DART: A Structured Dataset of Regulatory Drug Documents in Italian for Clinical NLP Mariano Barone et.al. 2510.18475 null
2025-10-21 CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Xue Jiang et.al. 2510.18471 null
2025-10-21 CircuitSeer: Mining High-Quality Data by Probing Mathematical Reasoning Circuits in LLMs Shaobo Wang et.al. 2510.18470 null
2025-10-21 IMB: An Italian Medical Benchmark for Question Answering Antonio Romano et.al. 2510.18468 null
2025-10-21 Simple and Efficient Heterogeneous Temporal Graph Neural Network Yili Wang et.al. 2510.18467 null
2025-10-21 CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learning Masato Kikuchi et.al. 2510.18466 null
2025-10-21 Large Language Models in Thematic Analysis: Prompt Engineering, Evaluation, and Guidelines for Qualitative Software Engineering Research Cristina Martinez Montes et.al. 2510.18456 null
2025-10-21 Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models Atharvan Dogra et.al. 2510.18454 null
2025-10-21 PlanU: Large Language Model Decision Making through Planning under Uncertainty Ziwei Deng et.al. 2510.18442 null
2025-10-21 Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation Yasser Hamidullah et.al. 2510.18439 null
2025-10-21 DeepTx: Real-Time Transaction Risk Analysis via Multi-Modal Features and LLM Reasoning Yixuan Liu et.al. 2510.18438 null
2025-10-21 Chain-of-Conceptual-Thought: Eliciting the Agent to Deeply Think within the Response Qingqing Gu et.al. 2510.18434 null
2025-10-21 ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization Yuanhe Guo et.al. 2510.18433 null
2025-10-21 Automated urban waterlogging assessment and early warning through a mixture of foundation models Chenxu Zhang et.al. 2510.18425 null
2025-10-21 Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents Guangfu Guo et.al. 2510.18424 null
2025-10-21 SegTune: Structured and Fine-Grained Control for Song Generation Pengfei Cai et.al. 2510.18416 null
2025-10-21 Adamas: Hadamard Sparse Attention for Efficient Long-Context Inference Siyuan Yan et.al. 2510.18413 null
2025-10-21 MENTOR: A Reinforcement Learning Framework for Model Enhancement via Teacher-Optimized Rewards in Small Models ChangSu Choi et.al. 2510.18383 null
2025-10-21 Training Diverse Graph Experts for Ensembles: A Systematic Empirical Study Gangda Deng et.al. 2510.18370 null
2025-10-21 KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs Donghyeon Ko et.al. 2510.18368 null
2025-10-21 Evaluating LLM-Based Mobile App Recommendations: An Empirical Study Quim Motger et.al. 2510.18364 null
2025-10-21 KrishokBondhu: A Retrieval-Augmented Voice-Based Agricultural Advisory Call Center for Bengali Farmers Mohd Ruhul Ameen et.al. 2510.18355 null
2025-10-21 GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data Yudong Li et.al. 2510.18345 null
2025-10-21 Combining Distantly Supervised Models with In Context Learning for Monolingual and Cross-Lingual Relation Extraction Vipul Rathore et.al. 2510.18344 null
2025-10-21 Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs Jongmin Lee et.al. 2510.18340 null
2025-10-21 ECG-LLM– training and evaluation of domain-specific large language models for electrocardiography Lara Ahrens et.al. 2510.18339 null
2025-10-21 Position: LLM Watermarking Should Align Stakeholders’ Incentives for Practical Adoption Yepeng Liu et.al. 2510.18333 null
2025-10-21 InspectCoder: Dynamic Analysis-Enabled Self Repair through interactive LLM-Debugger Collaboration Yunkun Wang et.al. 2510.18327 null
2025-10-21 Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding Jinlin Li et.al. 2510.18321 null
2025-10-21 Genesis: Evolving Attack Strategies for LLM Web Agent Red-Teaming Zheng Zhang et.al. 2510.18314 null
2025-10-21 ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation Haowei Lou et.al. 2510.18308 null
2025-10-21 The Impact of Image Resolution on Biomedical Multimodal Large Language Models Liangyu Chen et.al. 2510.18304 null
2025-10-21 Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models Lehan Wang et.al. 2510.18303 null
2025-10-21 From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering Lei Li et.al. 2510.18297 null
2025-10-21 BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks Tianyuan Huang et.al. 2510.18288 null
2025-10-21 Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs Yanhong Li et.al. 2510.18279 null
2025-10-21 Enhancing Hotel Recommendations with AI: LLM-Based Review Summarization and Query-Driven Insights Nikolaos Belibasakis et.al. 2510.18277 null
2025-10-21 StreamingTOM: Streaming Token Compression for Efficient Video Understanding Xueyi Chen et.al. 2510.18269 null
2025-10-21 UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding Da Zhang et.al. 2510.18262 null
2025-10-21 DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization Tao Tao et.al. 2510.18257 null
2025-10-21 Illusions of reflection: open-ended task reveals systematic failures in Large Language Models’ reflective reasoning Sion Weatherhead et.al. 2510.18254 null
2025-07-31 Instruction-tuned Large Language Models for Machine Translation in the Medical Domain Miguel Rios et.al. 2408.16440 null
2025-05-28 WizardCoder: Empowering Code Large Language Models with Evol-Instruct Ziyang Luo et.al. 2306.08568 null
2025-05-28 WizardLM: Empowering large pre-trained language models to follow complex instructions Can Xu et.al. 2304.12244 null
2025-03-18 In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models David Ponce et.al. 2503.01611 null
2025-02-19 Efficient Alignment of Large Language Models via Data Sampling Amrit Khera et.al. 2411.10545 null
2025-02-19 Acoustic Prompt Tuning: Empowering Large Language Models with Audition Capabilities Jinhua Liang et.al. 2312.00249 null
2024-11-12 PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications Dingkang Yang et.al. 2405.19266 null
2024-11-06 Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models Shengzhi Li et.al. 2402.10884 null
2024-05-31 X-Instruction: Aligning Language Model in Low-resource Languages with Self-curated Cross-lingual Instructions Chong Li et.al. 2405.19744 null
2024-05-31 Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models Xudong Lu et.al. 2402.14800 null
2024-04-17 Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents Renxi Wang et.al. 2402.11651 null
2024-04-09 A Survey on Transformer Compression Yehui Tang et.al. 2402.05964 null
2024-03-27 Are Compressed Language Models Less Subgroup Robust? Leonidas Gee et.al. 2403.17811 null
2024-02-20 Demystifying Instruction Mixing for Fine-tuning Large Language Models Renxi Wang et.al. 2312.10793 null
2023-11-14 FinGPT: Instruction Tuning Benchmark for Open-Source Large Language Models in Financial Datasets Neng Wang et.al. 2310.04793 null
2023-11-09 PB-LLM: Partially Binarized Large Language Models Yuzhang Shang et.al. 2310.00034 null
2023-11-07 Adapting Language Models to Compress Contexts Alexis Chevalier et.al. 2305.14788 null
2023-10-02 Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models Neha Sengupta et.al. 2308.16149 null
2023-09-06 Making Large Language Models Better Reasoners with Alignment Peiyi Wang et.al. 2309.02144 null
2023-06-27 Low-Rank Prune-And-Factorize for Language Model Compression Siyu Ren et.al. 2306.14152 null

Reinforcement Learning

Publish Date Title Authors PDF Code
2026-09-09 Optimal Intermediate Hamiltonians for Non-Equilibrium Free Energy Calculations: A Numerical Study of Markov Models David Beyer et.al. 2609.10519 null
2026-09-09 Faster Quantum Monte Carlo Simulation by Random Compilation John M. Martyn et.al. 2609.10486 null
2026-09-09 Compact totally separated types Martín Hötzel Escardó et.al. 2609.10447 null
2026-09-09 ConvMem: Convolutional Memory for Long-Context Reasoning Hongming Zhang et.al. 2609.10441 null
2026-09-09 Multi-Agent Reinforcement Learning for Autonomous UAV Exploration in Wildfire Response Caden Chandra et.al. 2609.10433 null
2026-09-09 Dynamic prediction intervals for survival times Lorenzo Carvisiglia et.al. 2609.10409 null
2026-09-09 Searching for New Physics with Reinforcement Learning Jacky Kumar et.al. 2609.10382 null
2026-09-09 Towards new D meson fragmentation functions Manuel Epele et.al. 2609.10327 null
2026-09-09 TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards Rui Sun et.al. 2609.10315 null
2026-09-09 Semiparametric Inference for Conditional Shapley Feature Importance Agostino Gnasso et.al. 2609.10313 null
2026-09-09 On the Limits of Quantum Multiparty Simultaneous Communication Pedro Montealegre et.al. 2609.10289 null
2026-09-09 Learning Terrain-Adaptive Humanoid Locomotion on Granular Terrain Junnosuke Kamohara et.al. 2609.10286 null
2026-09-09 Efficient LOS-Sampled GNSS Direct Position Estimation: An Information-Loss CRB Analysis Wei Gao et.al. 2609.10279 null
2026-09-09 Spatial sparse sampling-based iterative optimization framework for GNSS Direct Position Estimation Wei Gao et.al. 2609.10241 null
2026-09-09 Hierarchical and Permutation-Invariant Feature Transformation Learning via Policy-Guided Embedding Search Rui Liu et.al. 2609.10225 null
2026-09-09 Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection Haoyue Liu et.al. 2609.10221 null
2026-09-09 Adaptive Shared Control with Online Bounded-Rational Human Behavior Estimation Henry Ascencio Trejo et.al. 2609.10215 null
2026-09-09 pyeCE: A Python Implementation of the Embedded Cluster Expansion Yann L. Müller et.al. 2609.10190 null
2026-09-09 Fast, Accurate, and Scalable Fermionic Neural Networks via Translation Equivariance David D. Dai et.al. 2609.10186 null
2026-09-09 A Robust Binary Nonlinear Solver for Multi-stage Decisions Kashif Rashid et.al. 2609.10145 null
2026-09-08 A Data-Driven Framework for Identifying and Prioritizing RPA Opportunities in Healthcare Processes Maria Alejandra Gomez et.al. 2609.09137 null
2026-09-08 Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation Jiacheng Xu et.al. 2609.09135 null
2026-09-08 ExecCritic: Learn to Test, Test to Improve for Coding Agents Leitian Tao et.al. 2609.09133 null
2026-09-08 The Surprising Effectiveness of Approximate Value Iteration in Self-Play Raphael Boige et.al. 2609.09094 null
2026-09-08 ThinkPrior: Zero-Rollout Difficulty Priors for Cold-Start Prompt Selection in RLVR Tommy Sha et.al. 2609.09075 null
2026-09-08 Effects of Interaction Range on Fluid Multicriticality: A Computational Study of an Interconverting Lattice Model Thomas J. Longo et.al. 2609.09074 null
2026-09-08 The Path Integral Monte Carlo Sign Problem Is Not Always NP-Hard: Harmonic Fermions Can Be Solved in Quadratic Time Aarif Chaudhary et.al. 2609.09071 null
2026-09-08 Efficient Quantile-Resolved Hosting Capacity Assessment on Nodal Level for Low-Voltage Grids Maximilian Köhler et.al. 2609.09060 null
2026-09-08 PlayTrain: An Efficient Reinforcement Learning Framework for LLM-Generated Adaptable JavaScript Games Ryan Truong et.al. 2609.09059 null
2026-09-08 Selection Rules for Species Coexistence in a Hierarchical May-Leonard Model Rakesh Samanta et.al. 2609.09027 null
2026-09-08 On the sample complexity of the active subspace method Fabio Nobile et.al. 2609.08940 null
2026-09-08 AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Ziyang Ma et.al. 2609.08936 null
2026-09-08 Probing the Inert Scalar Sector of the Inert Doublet Model via Vector-Boson Fusion at a Muon Collider Abdesslam Arhrib et.al. 2609.08918 null
2026-09-08 Strong Polarization Signatures from Magnetically Stabilized Luminous Thin Accretion Disks P. Chris Fragile et.al. 2609.08895 null
2026-09-08 On Weighted Mathai-Haubold Entropy Measures Oindrali Das et.al. 2609.08889 null
2026-09-08 Equilibria for Time-inconsistent Regular-singular Control Problems Yuting Jia et.al. 2609.08877 null
2026-09-08 CAST: Alternating State-Value Targets and Expanded Policy Gradients for Model-Based Reinforcement Learning Pietro Noah Crestaz et.al. 2609.08853 null
2026-09-08 Asynchronous Model Predictive Control Under Model Mismatch: Stability and Performance Guarantees Changrui Liu et.al. 2609.08836 null
2026-09-08 Graph-Based Safe Reinforcement Learning for Multi-Agent Systems with Time-Varying Topology Xiao Sizhe et.al. 2609.08802 null
2026-09-08 Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Youngrok Park et.al. 2609.08798 null
2026-09-04 Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load Relief Meiyi Li et.al. 2609.05406 null
2026-09-04 Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models Wonje Jeung et.al. 2609.05401 null
2026-09-04 Sharp exponential integrability of conjugate functions David Norrbo et.al. 2609.05348 null
2026-09-04 How dipolar interactions structure molecular droplets Wiiliam Freitas et.al. 2609.05344 null
2026-09-04 Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness Alexander Neubauer et.al. 2609.05314 null
2026-09-04 Human-Human & Human-Robot Interaction Transformer (H2INT) for Robot Navigation in Dense and Uncertain Crowds Ao Shen et.al. 2609.05300 null
2026-09-04 Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning Fatemeh Saberi Khomami et.al. 2609.05298 null
2026-09-04 GUT: Quantifying and Optimizing the Reasoning Uncertainty of LLMs via Graph Complexity Shuang Liang et.al. 2609.05284 null
2026-09-04 A Unified Physics-Aware Quantum Machine Learning Framework across Power GaN HEMTs and Logic Nanowire FETs: Predicting Unseen Process Splits and Held-Out Geometry Combinations with Lower Error and Tighter Split-to-Split Variability Rushat Rai et.al. 2609.05251 null
2026-09-04 Diffusion under competing bulk and surface stopping mechanisms Yilin Ye et.al. 2609.05247 null
2026-09-04 First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves Tianjie Ju et.al. 2609.05224 null
2026-09-04 Stellar Population and Dynamical Modeling of Galaxies with Kinematically Misaligned Components: Age and Metallicity of Structural Components V. S. Goradzhanov et.al. 2609.05216 null
2026-09-04 FluxDisco: Symbolic Regression for Stoichiometric Dynamical Systems via Monte Carlo Graph Search Cassandra Durr et.al. 2609.05207 null
2026-09-04 Morphology and actuation as inductive biases in robotic hand manipulation Zalán Tari et.al. 2609.05206 null
2026-09-04 What Photocurrent Versus Effective Voltage Tells Us About Charge Generation in Organic Solar Cells Ardalan Armin et.al. 2609.05170 null
2026-09-04 Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG Shuyu Guo et.al. 2609.05152 null
2026-09-04 SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding Shenxi Wu et.al. 2609.05141 null
2026-09-04 Modelling Palomar Transients: Constraints from Reflection Geometry and Orbital Altitude Beatriz Villarroel et.al. 2609.05105 null
2026-09-04 Duality for Stochastic Control with non-Markovian Random Coefficients Peter Bank et.al. 2609.05101 null
2026-09-04 Operational Roles of QRNG-Derived Quantum Entropy in Bitcoin Proof-of-Work Architectures Ricardo Fernandes da Silva et.al. 2609.05092 null
2026-09-03 Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning Kevin Du et.al. 2609.04194 null
2026-09-03 A Computationally Feasible Framework for Causal Probabilistic Explanation Rafal Urbaniak et.al. 2609.04177 null
2026-09-03 A Low-Cost, Open Platform for End-to-End Autonomous Driving on a Miniature Ackermann Vehicle Gustavo Claudio Karl Couto et.al. 2609.04147 null
2026-09-03 HyperDet Wavefunction: A Phase-Agnostic Ansatz for Strongly Correlated Systems Xiaodong Hu et.al. 2609.04146 null
2026-09-03 Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR Boyan Li et.al. 2609.04108 null
2026-09-03 Zero sum two-player differential game under three regimes Brahim El Asri et.al. 2609.04100 null
2026-09-03 DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training Shubham Gandhi et.al. 2609.04094 null
2026-09-03 Monitoring antiproton numbers with a CMOS detector in a dense-track environment C. Regenfus et.al. 2609.04078 null
2026-09-03 Subspace Inference Enables Efficient Active Reward Learning from Preferences Yutai Zhou et.al. 2609.04066 null
2026-09-03 Spurious Advantage Hidden in GRPO Jiamian Wang et.al. 2609.04063 null
2026-09-03 When Models Edit Too Much: On the Fidelity of Minimal Code Edits Tongyao Zhu et.al. 2609.04061 null
2026-09-03 Mechanistic Framework for Multicomponent Nanoparticle Assembly: Predicting RNA-lipid and PEI-DNA nanoparticle assembly Turash Haque Pial et.al. 2609.04029 null
2026-09-03 RobustSeiz: An Open-Source Framework for Benchmarking the Robustness of EEG Seizure Detection Models Mohammad Mohammadi et.al. 2609.04007 null
2026-09-03 The Dually Flat Geometry of Planning as Inference Nikola Milosevic et.al. 2609.04005 null
2026-09-03 FiMI Banking: A Sovereign Model for Indian Retail Banking NPCI AI Research Team et.al. 2609.03960 null
2026-09-03 Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs Jiacheng Xu et.al. 2609.03955 null
2026-09-03 WorldReward: Reward Modeling for Camera-Conditioned World Models Yibin Wang et.al. 2609.03952 null
2026-09-03 Revisiting Topological Graphs for Macro Action based Closed-loop Reinforcement Learning of Vision Language Navigation in Continuous Environment Shuhao Ye et.al. 2609.03906 null
2026-09-03 Extremal Families for Matchings in Permutations Mengyu Cao et.al. 2609.03904 null
2026-09-03 Universal Driven Critical Dynamics of Entanglement Entropy Chang-Yu Shen et.al. 2609.03854 null
2026-09-02 Discriminative World Models for Web Agents Kelvin Li et.al. 2609.02885 null
2026-09-02 Post-Training Language Models for Gold-Medal Performance in Coding Competitions Aleksander Ficek et.al. 2609.02849 null
2026-09-02 Model-level synthetic-flux control of hyperchaos order and matched-resource sensing in dissipative optomechanics Stella Rolande Mbokop Tchounda et.al. 2609.02827 null
2026-09-02 Cliff: Learning Process Rewards from the First Mistake Peixuan Han et.al. 2609.02817 null
2026-09-02 GDB-Reward: From Evaluation Metrics to Training Rewards for Graphic Design Adrienne Deganutti et.al. 2609.02813 null
2026-09-02 Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization Giovanni Dispoto et.al. 2609.02677 null
2026-09-02 Minimal Radial Sub-Gamma Envelopes for Infinitely Divisible Random Vectors Yichuan Chen et.al. 2609.02595 null
2026-09-02 Online Reinforcement Learning in the Met Office Unified Model through Distributed Model-Agent Coupling Pritthijit Nath et.al. 2609.02566 null
2026-09-02 Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs Xixiang He et.al. 2609.02548 null
2026-09-02 A Comparative Study of Graph Representations for GNN-Based Power Grid Control in L2RPN Adrian Degenkolb et.al. 2609.02538 null
2026-09-02 Engineering of Non-Hermitian Trajectories and Phase Structure in an Open Bose-Hubbard Model via Rate Operator Transformations Jaakko Luomala et.al. 2609.02490 null
2026-09-02 Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking Siyu Chen et.al. 2609.02414 null
2026-09-02 Leveraging Time-Causal State Variable Aggregation for Real-Time Schedule of Massive Air Conditioners Jingguan Liu et.al. 2609.02410 null
2026-09-02 NE-R1: Enhancing Named Entity Recognition Model via Reinforcement Learning Meixuan Chen et.al. 2609.02366 null
2026-09-02 APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering Jie Ding et.al. 2609.02253 null
2026-09-02 DiffuSearch: How Hybrid Trajectory Planning Benefits from Aligned Objectives in Diffusion and Action Space Steffen Hagedorn et.al. 2609.02252 null
2026-09-02 RideSkill: A Hierarchical Algorithm for Generalized Ride Sharing with LLM-Driven Automatic Evolution Zijian Zhao et.al. 2609.02250 null
2026-09-02 Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RL Hyeonseong Jeon et.al. 2609.02237 null
2026-09-02 PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks Yuyao Zheng et.al. 2609.02236 null
2026-09-02 Scientific performances of the XGIS instrument on-board THESEUS E. Arrigoni et.al. 2609.02235 null
2026-09-01 The Rise of Verbal Reinforcement Learning Kshitij Tayal et.al. 2609.01597 null
2026-09-01 Facet-0: A Robotic Foundation Model for Contact-Rich Precise Manipulation Haoyuan Deng et.al. 2609.01596 null
2026-09-01 Mechanism Design for Alignment and Control Dirk Bergemann et.al. 2609.01595 null
2026-09-01 StudentSim: Training LLM-based Student Simulators Ke Yang et.al. 2609.01591 null
2026-09-01 Concentration of additive functionals of Stratonovich-type Rick Bebon et.al. 2609.01581 null
2026-09-01 Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs Jingtan Wang et.al. 2609.01573 null
2026-09-01 Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers Matteo Merler et.al. 2609.01567 null
2026-09-01 NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games Tomáš Holeček et.al. 2609.01549 null
2026-09-01 Polarized quantum effects in countable signals from intense laser - electron beam interactions Toseo Moritaka et.al. 2609.01494 null
2026-09-01 Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents Xiaofang Yang et.al. 2609.01487 null
2026-09-01 Does Imitation Learning Preserve Temporal Robustness in Dexterous Manipulation? An Expert-Learner Comparison Across Task Execution Speeds Clinton Enwerem et.al. 2609.01453 null
2026-09-01 Provably Safe Sim-to-Real Transfer Tingting Ni et.al. 2609.01418 null
2026-09-01 EdiTikZ: Scientific Figure Editing from Revision Trajectories Christian Greisinger et.al. 2609.01409 null
2026-09-01 Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR Esther Xin et.al. 2609.01354 null
2026-09-01 Scalable Rao-Blackwellized Online Planning for High-Dimensional POMDPs Jiho Lee et.al. 2609.01351 null
2026-09-01 Reconstruction of anomalous air showers with SKA-Low Vital De Henau et.al. 2609.01347 null
2026-09-01 VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models Zhiqi Huang et.al. 2609.01325 null
2026-09-01 Adaptive singular-point method for pricing and hedging surrenderable equity-linked contracts Andrea Molent et.al. 2609.01323 null
2026-09-01 Self-Healing Diffusion Monte Carlo applied to a simple fermionic model: A critical assessment of the method Michel Caffarel et.al. 2609.01301 null
2026-09-01 Sensitivity Oracles for Matroid Packing, Matroid Covering, and Matching Problems with Applications Keerti Choudhary et.al. 2609.01283 null
2026-08-31 PaperGym: Rubric-Centered Evolution for Research-Plan Generation Yuhan Wang et.al. 2608.31119 null
2026-08-31 Static-Field Shielding of Bosonic Molecules: Evaporation to Degeneracy and Self-Bound Droplets Jongheum Jung et.al. 2608.31116 null
2026-08-31 DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Jiashu Zhu et.al. 2608.31106 null
2026-08-31 Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization Jingxiao Yang et.al. 2608.31077 null
2026-08-31 Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Zhiqin Yang et.al. 2608.31075 null
2026-08-31 Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Yi Ding et.al. 2608.31046 null
2026-08-31 Normalized Low-Rank Adaptation Jiale Kang et.al. 2608.31036 null
2026-08-31 When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models Joonyong Park et.al. 2608.31035 null
2026-08-31 Semantic-Aware Sub-Band Allocation for Terahertz Communications Fatima Ismail et.al. 2608.30984 null
2026-08-31 Autonomously Acquiring Robot Manipulation Skills with Language-Driven Quality-Diversity Émiland Garrabé et.al. 2608.30983 null
2026-08-31 Evaluating and Improving LLM Self-Modeling Siqi Zeng et.al. 2608.30980 null
2026-08-31 One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning Armin Dariani et.al. 2608.30952 null
2026-08-31 Compensation in continuous symmetric trilayered planar ferrimagnet Olivia Mallick et.al. 2608.30941 null
2026-08-31 LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation Shaoan Wang et.al. 2608.30935 null
2026-08-31 Beacon: LLM Multi-Agent Driven Hardware Design Space Exploration for Heterogeneous Multi-Chiplet Deep Learning Accelerators Boyu Li et.al. 2608.30932 null
2026-08-31 S3C-LLM: Skill-Code Guided Agentic Language Models for Spectrum-to-Structure Elucidation Xuanle Zhao et.al. 2608.30910 null
2026-08-31 Scalable Statistical Inference in Stochastic Gradient Descent Rahul Singh et.al. 2608.30845 null
2026-08-31 HSRM: Hidden-State Reward Models for Test-Time Verification Xianzhi Li et.al. 2608.30841 null
2026-08-31 Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions Rodrigo Almeida et.al. 2608.30795 null
2026-08-31 Learning to infer and manipulate through distributed whole-arm interaction in a soft robot Chuhan Zhang et.al. 2608.30773 null
2026-08-31 Three Steps at a Time: Learning Representations from Action Sequences in Contrastive RL Michal Korniak et.al. 2608.30640 null
2026-08-31 Test-time Reinforcement Learning in Imperfect Information Games Ondrej Kubicek et.al. 2608.30635 null
2026-08-31 On the Riemann Boundary Value Problem for Poly- and Meta-hyperanalytic Function Spaces over d-summable Curves Yan Dai et.al. 2608.30634 null
2026-08-31 GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning Outongyi Lv et.al. 2608.30632 null
2026-08-31 Efficient measurement schemes for the Monte Carlo projective quantum eigensolver Divye Baid et.al. 2608.30612 null
2026-08-31 Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum Xingjian Wang et.al. 2608.30584 null
2026-08-31 Cascades from ultra-high-energy neutrinos Gaetano Di Marco et.al. 2608.30562 null
2026-08-31 SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards Zixun Guo et.al. 2608.30559 null
2026-08-31 DiffPDE: Masked Diffusion Language Models as PDE Solver Wenxuan Guo et.al. 2608.30532 null
2026-08-31 PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs Yuanqiang Yu et.al. 2608.30528 null
2026-08-31 Bounds on the Posterior-to-Prior Ratios for Inclusion Belief under Bounded Differential Privacy Jan Reiter Sørensen et.al. 2608.30473 null
2026-08-31 ToxLens: A Reproducible Graph-Learning Framework for Leakage-Aware, Uncertainty-Calibrated Molecular Toxicity Prediction Magnus H. Strømme et.al. 2608.30472 null
2026-08-31 Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry Jiani Guo et.al. 2608.30457 null
2026-08-31 Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry Kihun Rhee et.al. 2608.30442 null
2026-08-31 Locally-Guided Actor-Critic: Training a Goal-conditioned Actor with a Subgoal-aware Critic Olivier Serris et.al. 2608.30406 null
2026-08-31 SemPOI-RL: Aligning LLM Semantic Reasoning for Interpretable Out-of-Town POI Sequential Generation Yunqi Liu et.al. 2608.30399 null
2026-08-31 Beyond Polarization: The Generative Constraint of Chain-of-Thought in Pointwise Reranking Xiaoyang Chen et.al. 2608.30398 null
2026-08-31 When LLM Meets Tree Search: A Systematic View of Inference as Search in Large Language Models Jiaqi Wei et.al. 2608.30395 null
2026-08-31 Compact and Infinite-Order Error Analysis for Null-Space SVD Estimation Xin Li et.al. 2608.30374 null
2026-08-31 Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance Donggyu Min et.al. 2608.30317 null
2026-08-28 Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning Nan Wang et.al. 2608.28578 null
2026-08-28 Machine learned designs of functional colloidal foldamers Ryan van Mastrigt et.al. 2608.28554 null
2026-08-28 A reaction volume bias Monte Carlo trial for sampling chemisorption in confinement Samiha Sharlin et.al. 2608.28516 null
2026-08-28 Machine-learning-assisted multiscale topology optimization of functionally graded superimposed lattice structures Prashant Kumar Gupta et.al. 2608.28513 null
2026-08-28 REPLICANT: Learning Policies for Evading and Hardening Malware Detectors Shae McFadden et.al. 2608.28499 null
2026-08-28 Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning Minghui Xu et.al. 2608.28447 null
2026-08-28 Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs Vishvesh Bhat et.al. 2608.28421 null
2026-08-28 Astrophysical Sensitivity Projections for the IceCube Upgrade R. Abbasi et.al. 2608.28395 null
2026-08-28 Real-Time Musculoskeletal Surrogates for Pediatric Cerebral Palsy: a Credibility Pilot Mohammad Arif Ul Alam et.al. 2608.28371 null
2026-08-28 Response Propensity Estimation and Cross-Fitting Alessandro La Rocca et.al. 2608.28324 null
2026-08-28 Scalable Voltage-Stability Dataset Generation Via Boundary-Proximity Indicators Clustering Rock Agon et.al. 2608.28298 null
2026-08-28 A Generalized Model for Disordered Random Sequential Adsorption with Charge-Dependent Deposition G. Palacios et.al. 2608.28289 null
2026-08-28 Searching for vectorlike $T$ quarks in the $T\to tZ$ channel at a future muon-proton collider Yin-Hao Gao et.al. 2608.28285 null
2026-08-28 Quantum many-body effects in the optical response of ideal thin films David Trejo-Garcia et.al. 2608.28261 null
2026-08-28 Fine-Tuning Autobidders with Group Relative Policy Optimization Anton Safin et.al. 2608.28199 null
2026-08-28 HARTS: Efficient Agentic Reinforcement Learning for Hybrid-Attention Models over Arbitrary Rollout Trees Boyuan Meng et.al. 2608.28158 null
2026-08-28 Contact-Guided Exploration for Non-Prehensile Locomanipulation with Multi-Critic RL Simone Tolomei et.al. 2608.28140 null
2026-08-28 VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning Pengcheng Li et.al. 2608.28128 null
2026-08-28 Parameter estimation in Conditional Sequential Monte Carlo algorithms through Particle Learning Alfonso Diz-Lois Palomares et.al. 2608.28079 null
2026-08-28 Extremely Low Mass Ratio Contact Binaries. III. Photometric and Spectroscopic Investigations of Eleven Systems Youmo Lai et.al. 2608.28066 null
2026-08-27 TTPO: Test-Time Policy Optimization Aozhe Wang et.al. 2608.27448 null
2026-08-27 Stochastic Estimation of Transduced Language Models Vésteinn Snæbjarnarson et.al. 2608.27428 null
2026-08-27 Boosting LLM Exploration via Weak-Model Guidance in RLVR Xingyu Shen et.al. 2608.27420 null
2026-08-27 Consolidating RLVR Capabilities Across Domains: A Deep Dive into Fusion Paradigms Siye Wu et.al. 2608.27409 null
2026-08-27 Deep-Control BSDE: Layerwise Brownian-Weighted Regression for High-Dimensional Semilinear PDEs Mingcan Wang et.al. 2608.27369 null
2026-08-27 A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning Zijie Cheng et.al. 2608.27313 null
2026-08-27 Accelerating Optical Photon Simulation in DUNE with Opticks Ilker Parmaksiz et.al. 2608.27306 null
2026-08-27 Astar: Learning to Propose Evolution Directions for Self-Evolving Industrial AI Systems Jinxin Hu et.al. 2608.27287 null
2026-08-27 A Multilevel Interacting Particle System Method for the estimation of Failure Probabilities Rubén Aylwin et.al. 2608.27275 null
2026-08-27 Grain-Boundary Premelting in High-Entropy Transition Metal Carbides Marium M. Mou et.al. 2608.27273 null
2026-08-27 On the approximation of posterior laws in compound loss models by conditional Wasserstein GANs Aleksandar Arandjelovic et.al. 2608.27229 null
2026-08-27 A Trans-Domain Digital Twin for Bio-Aware Control of Climate and Energy in Cattle Fattening Barns Using Single-Episode Optimizer Learning Mansoorali Amiri et.al. 2608.27185 null
2026-08-27 Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation Wendong Li et.al. 2608.27158 null
2026-08-27 GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL Zike Yuan et.al. 2608.27142 null
2026-08-27 Too good to go: Upcycling Phase-Space Points for Multijet Processes Konrad Helms et.al. 2608.27104 null
2026-08-27 Emotional Preferences as Goal-Priority Regulation Shiqi Liu et.al. 2608.27072 null
2026-08-27 Performance Foundations of Parallel & Distributed Reasoning Language Models Maciej Besta et.al. 2608.27046 null
2026-08-27 ProRetrieval: Learning to Orchestrate Hybrid Search via Executable Program Synthesis Chengsong You et.al. 2608.27017 null
2026-08-27 JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols Chen Chen et.al. 2608.26982 null
2026-08-27 RubricRM: Generative Reward Modeling via Dynamic Rubrics for Image Generation and Editing Zijian Kan et.al. 2608.26956 null
2026-08-26 VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Junxiang Xu et.al. 2608.26105 null
2026-08-26 Anatomy of an extensive air shower: building an optimal radio emission calculation Juan Ammerman-Yebra et.al. 2608.26077 null
2026-08-26 Prefix Sliding for efficient test-time scaling Niklas Muennighoff et.al. 2608.26070 null
2026-08-26 $R^3$ : Training Robots to Reason in Natural Language via Reinforcement Learning Lehong Wu et.al. 2608.26053 null
2026-08-26 VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following Min Zeng et.al. 2608.26013 null
2026-08-26 Imitation Learning for Connection-Tableau Construction Fredrik Rømming et.al. 2608.26009 null
2026-08-26 A Weighting Method for Incorporating Mass Resolution Effects in Amplitude Analysis Benhou Xiang et.al. 2608.25969 null
2026-08-26 Optimized Multilevel Sampling Methods under Resource Constraints Niklas Baumgarten et.al. 2608.25958 null
2026-08-26 One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation Justin Robert et.al. 2608.25936 null
2026-08-26 Continually learning neural-operator surrogate for three-dimensional airborne electromagnetic Bayesian inversion Jaehong Chung et.al. 2608.25932 null
2026-08-26 Random Invariance Testing on Quadratic Form Statistics with Application to Autocorrelation Amitakshar Biswas et.al. 2608.25918 null
2026-08-26 Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence Shengyi Pan et.al. 2608.25905 null
2026-08-26 VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation Jiayi Chen et.al. 2608.25872 null
2026-08-26 Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data Rene Glitza et.al. 2608.25794 null
2026-08-26 Sequential Stability of the Value Function and the Solution Mapping in Berge’s Maximum Theorem via Variational Convergence John Cotrina et.al. 2608.25789 null
2026-08-26 TailSFT: Filtered Fine-Tuning Improves Post-Training Performance Sadhika Malladi et.al. 2608.25756 null
2026-08-26 Muonium dynamics as a probe for depth-resolved properties of 4H-SiC Maria Mendes Martins et.al. 2608.25702 null
2026-08-26 AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification Zebei Zhao et.al. 2608.25637 null
2026-08-26 DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search Junzhao Zhang et.al. 2608.25635 null
2026-08-26 Advantage-Driven Explicit Memory for Social Navigation Yeonsoo Park et.al. 2608.25610 null
2026-08-25 SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL Kai Ruan et.al. 2608.24870 null
2026-08-25 Improving Cross-Problem Vehicle Routing with Locally Augmented Preferences and Representation Disentanglement Arthur Corrêa et.al. 2608.24859 null
2026-08-25 Bellman Calibration for Marginalized Importance Weighting in Offline Reinforcement Learning Lars van der Laan et.al. 2608.24858 null
2026-08-25 Learning Whom to Trust : Decision-Generated Credibility in Social Learning Gabriel Bontemps et.al. 2608.24851 null
2026-08-25 Sparse domination implies convex body domination Aapo Laukkarinen et.al. 2608.24802 null
2026-08-25 CAFE: Self-Improving Search Agents Need Co-Evolving Feedback Boyang Liu et.al. 2608.24794 null
2026-08-25 Score-Based Ideal Observer Approximation via Denoising Score Matching for Signal-Known-Exactly Detection Tasks Weimin Zhou et.al. 2608.24768 null
2026-08-25 SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents Shidong Yang et.al. 2608.24747 null
2026-08-25 Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion Zilong Huang et.al. 2608.24730 null
2026-08-25 On-policy Distillation with Verifiable Reward Wenze Lin et.al. 2608.24696 null
2026-08-25 Testing $f(Q)$ Gravity with DESI DR2 and Strong-Lensing Time Delays Darshan Kumar et.al. 2608.24676 null
2026-08-25 EviGraph: Towards Verifiable Evidence Construction for Information-Seeking Agents Jiashun Chen et.al. 2608.24667 null
2026-08-25 On-Policy Self-Distillation in Diffusion Models Wei Zhou et.al. 2608.24646 null
2026-08-25 EVEREST:Endogenous Vision-Language Reinforcement Reasoning Exploration for Urban Socio-Semantic Segmentation Qixiu Li et.al. 2608.24640 null
2026-08-25 IAPO: Influence-Aware Policy Optimization for Credit Assignment in Multi-Turn Service Agents Bo Ren et.al. 2608.24588 null
2026-08-25 Joint Optimization of Tool Creation and Use for Large Language Model Agents Zhi Rui Tam et.al. 2608.24571 null
2026-08-25 RoG-DAgger: Rollout-Guided Post-Training for End-to-End Driving Liangyu Zhong et.al. 2608.24525 null
2026-08-25 Assessing the Credibility of Gamma-Ray QPO Candidates in 41 TeV-Selected Blazars Wen-Xin Yang et.al. 2608.24522 null
2026-08-25 NeuralParker: A Reinforcement Learning Planner for Irregular Parking Environments Zihan Wang et.al. 2608.24485 null
2026-08-25 WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Zihao Wu et.al. 2608.24479 null
2026-08-24 How to Train a Critic Stably and Efficiently Penghui Qi et.al. 2608.23566 null
2026-08-24 Testing selection on observables in parametric models with refreshment samples Grigory Franguridi et.al. 2608.23508 null
2026-08-24 Dynamical Love numbers of analogue rotating black and white holes Satadal Datta et.al. 2608.23496 null
2026-08-24 SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning Jialong Liu et.al. 2608.23493 null
2026-08-24 Reward-Free Continual Adaptation for Resilient Space Robots Andrej Orsula et.al. 2608.23452 null
2026-08-24 Temporal Property-driven Design Space Exploration with Reinforcement Learning for Cyber-Physical Systems Tagir Fabarisov et.al. 2608.23440 null
2026-08-24 Quasiconvexity of the Burkholder function on symmetric matrices André Guerra et.al. 2608.23388 null
2026-08-24 A DQMC study of the spectral and conductive properties of the two-dimensional Holstein model J. Neuhaus et.al. 2608.23379 null
2026-08-24 Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents Wenqi Liu et.al. 2608.23329 null
2026-08-24 Agent-G $^2$ : Gaussian Guidance for Agentic Reinforcement Learning Zixuan Wang et.al. 2608.23318 null
2026-08-24 Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization Xianlei Zhou et.al. 2608.23311 null
2026-08-24 Quantum Monte Carlo in the Age of Many-Body Quantum Information Yi-Ming Ding et.al. 2608.23231 null
2026-08-24 Universality of superdiffusion in simple random graphs Mrinal Sarkar et.al. 2608.23207 null
2026-08-24 Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers Hossein Abdi et.al. 2608.23204 null
2026-08-24 Zeroth-Order Nonsmooth Nonconvex Optimization with Convex Liftings and Its Application to State-Feedback $H_\infty$ Policy Optimization Xuhao Wang et.al. 2608.23178 null
2026-08-24 Distributed Trajectory Planning and Resource Allocation for Dynamic Multi-UAV Collaborative Computing Tiankui Zhang et.al. 2608.23123 null
2026-08-24 Quantum Reservoir Computing with Physics-Informed Correction for Reduced-Order PDE Forecasting Krishna Bhatia et.al. 2608.23119 null
2026-08-24 From Generation to Simulation: How Far Are World Models from Being True Simulators? Tong Wang et.al. 2608.23070 null
2026-08-24 Macro-Action Topological Navigation under Noisy Localization using Reinforcement Learning Simon Hakenes et.al. 2608.23055 null
2026-08-24 MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Yi Zhu et.al. 2608.23035 null
2026-08-21 Efficient Event Generation for High-Multiplicity LHC Processes: An End-to-End GPU Workflow with Normalizing Flows Enrico Bothmann et.al. 2608.21338 null
2026-08-21 Maxwell’s Demon in Markov Chain Monte Carlo: Cooling Information Flow and Entropy Balance Masayuki Ohzeki et.al. 2608.21337 null
2026-08-21 Ultralow-Field Triplon Condensation in a Spin-Ladder Magnet Ankit Labh et.al. 2608.21316 null
2026-08-21 Re $^3$ Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning Haonan Jia et.al. 2608.21305 null
2026-08-21 AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization Huizu Lin et.al. 2608.21292 null
2026-08-21 Neural quantum states in condensed matter: advances, best practices, and prospects Jonas B. Rigo et.al. 2608.21291 null
2026-08-21 Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning Varun Giridhar et.al. 2608.21204 null
2026-08-21 SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control Ruihua Han et.al. 2608.21175 null
2026-08-21 A Fourier Neural Operator for Accelerated Discretization-Invariant Solutions of the Radiative Transfer Equation Daniel Carne et.al. 2608.21173 null
2026-08-21 Distributed synthesis of arbitrary graph states in quantum networks via rank-two GF(2) reduction Xiaoyi Zheng et.al. 2608.21166 null
2026-08-21 Teaching is a Process: The TOSS Framework for Modeling Human Teaching Decisions in Human-Interactive Robot Learning Bernhard Hilpert et.al. 2608.21083 null
2026-08-21 $Z^2$ -ACT: End-to-End Verifiable Agentic Intent Control for Open 6G RAN Sunder Ali Khowaja et.al. 2608.21049 null
2026-08-21 Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight Zhitao Liu et.al. 2608.20948 null
2026-08-21 Artificial spin ice systems on single edge length tilings Ellie Weightman et.al. 2608.20928 null
2026-08-21 Spatial function-on-function quantile regression Eylul Fidan et.al. 2608.20919 null
2026-08-21 Multi-Objective Deep Reinforcement Learning for Secure and Stable Power System Operation Ioannis Papadopoulos et.al. 2608.20914 null
2026-08-21 Decoupling Policy Extraction for Offline Reinforcement Learning Xuyao Lin et.al. 2608.20909 null
2026-08-21 Reinforcement learning for vertical position control on the EXL-50U spherical tokamak Lei Xing et.al. 2608.20901 null
2026-08-21 Sharing the Control Authority Between Deep Reinforcement Learning and Model Predictive Control: Application to Multi-Class Transportation Networks Giray Onur et.al. 2608.20858 null
2026-08-21 Demonstration-Guided Humanoid Stand-Up on an Emulated Deformable Surface Aniruddh Kushwah et.al. 2608.20852 null
2026-08-20 Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models Taihang Hu et.al. 2608.20334 null
2026-08-20 G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation Shiao Xie et.al. 2608.20331 null
2026-08-20 MidTool: Mid-training Data Synthesis for Agentic Tool Use Fengqing Jiang et.al. 2608.20314 null
2026-08-20 Robustness of random-walk Metropolis for steep potentials Sam Power et.al. 2608.20279 null
2026-08-20 Uniform weak type $(1,1)$ bounds for Riesz transforms on stratified Lie groups Sheng-Chen Mao et.al. 2608.20267 null
2026-08-20 Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation Gijs Kassenaar et.al. 2608.20256 null
2026-08-20 Directional Subdifferentials of the Value Function in Asplund Spaces Weihao Mao et.al. 2608.20241 null
2026-08-20 RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation Shaoxuan Wang et.al. 2608.20208 null
2026-08-20 DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing Haoxiang Cao et.al. 2608.20161 null
2026-08-20 Reinforcement LearningtoHarness Approximation Errors for Long-Time QuantumSimulation Yu-Bo Shi et.al. 2608.20139 null
2026-08-20 Multi-Agent Orchestration with the Common-Sense Reasoning Capabilities of LLMs for Autonomous Driving Mehdi Azarafza et.al. 2608.20129 null
2026-08-20 Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo Lohithsai Yadala Chanchu et.al. 2608.20123 null
2026-08-20 Planning-Oriented End-to-End Autonomous Driving: Architectures, Evaluation, and Emerging Paradigms Yanchen Guan et.al. 2608.20111 null
2026-08-20 Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design Poomphob Suwannapichat et.al. 2608.20099 null
2026-08-20 Backstepping-Guided Reinforcement Learning for Wide-Range Saint-Venant Canal Regulation Chenchen Wang et.al. 2608.20089 null
2026-08-20 End-to-end Early Classification of Time Series in Non-Stationary Environments Aurélien Renault et.al. 2608.20044 null
2026-08-20 Emergence of cooperation: A reputation-modulated reinforcement learning Chenyang Zhao et.al. 2608.20016 null
2026-08-20 Learning Highly Dynamic Skills Transition for Quadruped Jumping Through Constrained Space Zeren Luo et.al. 2608.19977 null
2026-08-20 MILD: Tractable Terrain Modeling for Learning Improved Bipedal Locomotion on Deformable Surfaces Zeren Luo et.al. 2608.19955 null
2026-08-20 EXIMO: VLM Guided Exploration of VLA Policies Bhavya Sukhija et.al. 2608.19891 null
2026-08-19 ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning Jayjun Lee et.al. 2608.19182 null
2026-08-19 Continuous-Time Reinforcement Learning for Controlled Hawkes Jump-Diffusions Tomasz R. Bielecki et.al. 2608.19151 null
2026-08-19 Network-Scale Road Disruption from Liquefaction in Cascadia Subduction Zone Earthquakes M. D. Sanger et.al. 2608.19143 null
2026-08-19 PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints Boqiao Zhang et.al. 2608.19121 null
2026-08-19 JANUS: A Multi-modal Foundation Neural Sampler for Disordered Materials Denis Blessing et.al. 2608.19116 null
2026-08-19 Quantum circuit optimization using deep reinforcement learning: Applications across multiple gate sets Khoa Dang Tao et.al. 2608.19103 null
2026-08-19 Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation Huan-ang Gao et.al. 2608.19098 null
2026-08-19 The Radioactive Background of the JUNO Calibration System Rui Li et.al. 2608.19051 null
2026-08-19 Multi-Agent Off-Policy Deep Reinforcement Learning for Smart Campus Coverage Omar Rady et.al. 2608.19049 null
2026-08-19 Sign-problem-resilient singular-value probe in determinant quantum Monte Carlo Wen Chen et.al. 2608.19028 null
2026-08-19 AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL Daniele Lizzio Bosco et.al. 2608.18946 null
2026-08-19 Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck Davide Romano et.al. 2608.18931 null
2026-08-19 Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models Wei Yu et.al. 2608.18884 null
2026-08-19 Falcon Perception-HD: High Density Perception via Reinforcement Learning Sofian Chaybouti et.al. 2608.18881 null
2026-08-19 Think-to-Personalize: Unifying Reasoning and Retrieval for User-Centric Personalized Dense Retrieval Angqing Jiang et.al. 2608.18855 null
2026-08-19 MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models Chenglin Liu et.al. 2608.18827 null
2026-08-19 Dream2Reward: Transition-Alignment Reward Models from Positive Demonstrations for Robotic Manipulation Haoyu Zhang et.al. 2608.18787 null
2026-08-19 Size-Mass Relation Shows Its Colours: Contrasting Physical Imprints of Galaxy Evolution in Rest-Frame UV and Optical Angelo George et.al. 2608.18776 null
2026-08-19 To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization Taehyung Kim et.al. 2608.18770 null
2026-08-19 Robust Modeling of Extremes in the Presence of Inliers with Enhanced Tail Estimation Shivshankar Nila et.al. 2608.18735 null
2026-08-18 The concentration game: Bayesian updating, regret, and information Akshay Balsubramani et.al. 2608.18061 null
2026-08-18 Extending and Unifying the Fundamental Tasks of Hamilton-Jacobi Reachability Analysis Dylan Hirsch et.al. 2608.18060 null
2026-08-18 Runs Above Expected (RAE) and Wicket Effect (WE): A Context-adjusted and Unified Impact Metric for Twenty20 Cricket Rhitankar Bandyopadhyay et.al. 2608.18020 null
2026-08-18 Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents Christophe D. Hounwanou et.al. 2608.18008 null
2026-08-18 Hamiltonian dynamics for sampling on discrete spaces Raphaël Barboni et.al. 2608.17961 null
2026-08-18 Towards Zero-Shot Task Transfer with Neurosymbolic World Models Isidoro Tamassia et.al. 2608.17959 null
2026-08-18 Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation Zhizhao Liu et.al. 2608.17941 null
2026-08-18 Steady-State Equivalent Circuit Model for Data Center Loads Muhammad Hamza Ali et.al. 2608.17925 null
2026-08-18 TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification Neelesh Kumar Shukla et.al. 2608.17795 null
2026-08-18 Interference Engineering for Quantum Imaginary-Time Evolution through Multiple Energy Shifts Hong-Jian Tang et.al. 2608.17792 null
2026-08-18 Debate Training Reduces Reward Hacking in RLAIF Zachary Kenton et.al. 2608.17776 null
2026-08-18 Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See Ayoub Kirouane et.al. 2608.17744 null
2026-08-18 Offline Multi-Agent Reinforcement Learning with a Physics-Informed World Model for Cooperative Mixed Traffic Control Lu Liu et.al. 2608.17739 null
2026-08-18 Fault detection on manifolds of nonlinear dynamical systems with dual autoencoders Bulut Kuşkonmaz et.al. 2608.17698 null
2026-08-18 Picard Proximal Monte Carlo for Parallel Bayesian Imaging with Score-Based Generative Priors Deliang Wei et.al. 2608.17666 null
2026-08-18 rl-triton: High-Performance Triton GPU Kernels for Reinforcement Learning Credit Assignment Lars Simon Zehnder et.al. 2608.17641 null
2026-08-18 COS-TT-CHF: A Tensor-Train Characteristic-Function COS Method for Multi-Asset Option Pricing Lucas Arenstein et.al. 2608.17636 null
2026-08-18 Gauge-constrained Spinon Complexes Near Deconfined Quantum Criticality Zhi-Yao Ning et.al. 2608.17631 null
2026-08-18 Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision Amir Arsalan Nematollahi et.al. 2608.17628 null
2026-08-18 tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots Markus D. Kobelrausch et.al. 2608.17596 null
2026-08-17 Q-based Variational Inverse Reinforcement Learning Ondrej Bajgar et.al. 2608.16888 null
2026-08-17 AutoSR: Automatic Symbolic Regression by Searching Research States Kejia Zhang et.al. 2608.16876 null
2026-08-17 HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL Langzhe Gu et.al. 2608.16837 null
2026-08-17 ClawGym II: Exploring Black-Box RL on Agent Harness Huatong Song et.al. 2608.16798 null
2026-08-17 Neurosymbolic Embodied Agents Mohammad Albinhassan et.al. 2608.16794 null
2026-08-17 A Cross-Band (X-ray $\times$ Optical) Periodicity Search for Supermassive Black Hole Binaries: A Null Result and the First Completeness-Corrected Constraint Karan Akbari et.al. 2608.16787 null
2026-08-17 Le Critique: Privileged Value Functions for LLM Reinforcement Learning Siddarth Venkatraman et.al. 2608.16739 null
2026-08-17 A Stable Transport-Mechanism Descriptor for Per-Pixel Rendering Difficulty Po-Ting Lin et.al. 2608.16730 null
2026-08-17 MatchingPolicy: Correspondence-Aware Policy Enables Cross-Object In-Context Learning Qijin She et.al. 2608.16715 null
2026-08-17 The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback Thomas Mbrice et.al. 2608.16710 null
2026-08-17 Configurational Temperature in the 3D XY Model Kutloano Nkojoana et.al. 2608.16677 null
2026-08-17 Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents Amrapali Pednekar et.al. 2608.16666 null
2026-08-17 A nuclear-quantum-corrected machine-learning potential reveals quantum-enhanced hydrogen segregation at general grain boundaries in alpha-iron Kazuma Ito et.al. 2608.16652 null
2026-08-17 A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory Zhihui Chen et.al. 2608.16626 null
2026-08-17 Interactive Whole Slide Images for RL-based Tumour Segmentation Mohamad Mohamad et.al. 2608.16607 null
2026-08-17 CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction Tianqi Xiang et.al. 2608.16594 null
2026-08-17 Bessel-Debiased Pseudo-Marginal MCMC for Generalised Bayesian Inference Yingkai Lu et.al. 2608.16573 null
2026-08-17 Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning Yongqi Tong et.al. 2608.16554 null
2026-08-17 FLEET: Token-Based Feature Extraction for Event Camera-based Reinforcement Learning Tristan Gottwald et.al. 2608.16523 null
2026-08-17 Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation Marc Pérez-Roig et.al. 2608.16482 null
2026-08-14 Cosmic Ray Diffusion and the Origin of Very High Energy Gamma-Ray Emission in Young Massive Stellar Clusters Lucas Barreto-Mota et.al. 2608.14547 null
2026-08-14 Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training Hanfeng Lu et.al. 2608.14498 null
2026-08-14 A Fixed Universal Determinant is Variationally Complete for Continuum Fermions Giuseppe Carleo et.al. 2608.14476 null
2026-08-14 Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View Yixian Xu et.al. 2608.14430 null
2026-08-14 Uncertainty-Aware Jacobi Set Computation Daniel Klötzl et.al. 2608.14409 null
2026-08-14 Offline Deep Q* Estimation with Diffusion Models Xiaohong Chen et.al. 2608.14401 null
2026-08-14 Submodular Policy Learning for Distributed Task Allocation in Open Multi-Agent Systems Jing Liu et.al. 2608.14390 null
2026-08-14 Linearised quantum signal processing Marek Arsenault et.al. 2608.14387 null
2026-08-14 CoRun: Padding is Simple and Efficient for Deterministic LLM Inference Shiju Zhao et.al. 2608.14376 null
2026-08-14 Multi-Agent Reinforcement Learning for Joint Handover Management and Power Allocation in Multi-Orbit Satellite Networks Yassine Afif et.al. 2608.14335 null
2026-08-14 CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving Anisa Saleem et.al. 2608.14332 null
2026-08-14 Quantum Multi-Armed Bandits and Linear Bandits: Lower Bounds and Algorithms Maoli Liu et.al. 2608.14319 null
2026-08-14 Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL Xiaojun Wu et.al. 2608.14312 null
2026-08-14 PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment Yuyang Liu et.al. 2608.14284 null
2026-08-14 MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement Lushi Pu et.al. 2608.14221 null
2026-08-14 APTER: Adaptive Post-Training with Expert-Grounded Rubrics Xukai Wang et.al. 2608.14212 null
2026-08-14 Probing the Single Production of First-Generation Singlet Vector-like Leptons at Future $e^+e^-$ Colliders Yao-Bei Liu et.al. 2608.14195 null
2026-08-14 Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine Chenran Weng et.al. 2608.14157 null
2026-08-14 Deep Reinforcement Learning solution for pickup and delivery routing problems with time window and capacity constraints Andrew Soroka et.al. 2608.14156 null
2026-08-14 AgilePE: Autonomous UAV Pursuit-Evasion via Self-Play Reinforcement Learning Wenhao Tang et.al. 2608.14135 null
2026-08-13 Intern-S2-Preview: Scientific Agentic Foundation Model Lei Bai et.al. 2608.13505 null
2026-08-13 Landau theory and exchange instabilities in Mn $_5$Si$_3$ : A case against altermagnetism K. D. Belashchenko et.al. 2608.13483 null
2026-08-13 Imaginary-time correlations in time-sliced stochastic series expansion Ryan Flynn et.al. 2608.13477 null
2026-08-13 On-Off Digital Noise Modulation with Fluid Antenna Systems over $κ$-$μ$ Fading Channels Daniel C. Araújo et.al. 2608.13471 null
2026-08-13 FIRE-VLA: Failure-Informed Self-Evolution for Vision-Language-Action Models in Autonomous Driving Hao Dou et.al. 2608.13395 null
2026-08-13 Weighted cumulative past inaccuracy and Kullback-Leibler divergence based on extropy: properties, estimation, and applications Bighneswar Sahoo et.al. 2608.13363 null
2026-08-13 Rules or Character? Scaling Laws for AI Safety Design Satoshi Takahashi et.al. 2608.13345 null
2026-08-13 Radio Monitoring of Classical Novae using the ASKAP Variable and Slow Transients Survey Aishani Majumder et.al. 2608.13330 null
2026-08-13 The Time Value of Evolution Matthew Siper et.al. 2608.13297 null
2026-08-13 Emergent Symmetry-Protected Topological Phases via Polyakov Confinement in Quantum Spin Systems Li-Wei He et.al. 2608.13238 null
2026-08-13 TANGCO: Learning Topology-Aware Capacity Allocation for Overload-driven Cascading Failures Orkun Irsoy et.al. 2608.13212 null
2026-08-13 Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents Zechuan Wang et.al. 2608.13179 null
2026-08-13 Splat-based Metal Artifact Reduction in Cone-Beam CT via Polychromatic Modeling Kiseok Choi et.al. 2608.13159 null
2026-08-13 Pego theorem for Hilbert space-valued functions on compact groups Anaté Kodjovi Lakmon et.al. 2608.13142 null
2026-08-13 S2-HWM: Sparse Event-Structured Hierarchical World Model for Long-Horizon Surgical Robot Manipulation Shuzhe Zhang et.al. 2608.13103 null
2026-08-13 AoI-Guaranteed Dynamic Route Planning for Connected Vehicles Sajedeh Norouzi et.al. 2608.13083 null
2026-08-13 Pareto-Aware Hierarchical Reinforcement Learning for Online Resource Allocation in RIS-assisted Large-Scale IoT Systems Wenhan Xu et.al. 2608.13032 null
2026-08-13 Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning Yao Zhou et.al. 2608.13026 null
2026-08-13 OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways Mao Jiayang et.al. 2608.12995 null
2026-08-13 Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning Zijie Cheng et.al. 2608.12973 null
2026-08-12 Completeness properties of the space of quasicontinuous functions Ľubica Holá et.al. 2608.12318 null
2026-08-12 Redistribution-based Cost Inference Improves Sparse Safe Offline RL Ebenezer Gelo et.al. 2608.12306 null
2026-08-12 Finite-depth scaling and an exact Bernoulli-leaf identity for the min-plus process on the binary tree José Ricardo G. Mendonça et.al. 2608.12295 null
2026-08-12 When should one stop the most exciting game? Sequential Inference for win-martingales Steven Campbell et.al. 2608.12291 null
2026-08-12 Hemispheric Asymmetry of Solar Active Regions Arises from a Nested Population Aimee A. Norton et.al. 2608.12263 null
2026-08-12 SelectLight: Learning to Select Signal Plans Generated by Distributed Model Predictive Control for Urban Traffic Networks Lyuzhou Luo et.al. 2608.12256 null
2026-08-12 One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL Simon Yu et.al. 2608.12253 null
2026-08-12 SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward Zile Zhou et.al. 2608.12220 null
2026-08-12 Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation Md Yassir Mottalib et.al. 2608.12190 null
2026-08-12 Empirical likelihood confidence regions for ordered bivariate means Naresh Garg et.al. 2608.12174 null
2026-08-12 RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning Yibo Shen et.al. 2608.12146 null
2026-08-12 Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models Shukrullo Nazirjonov et.al. 2608.12078 null
2026-08-12 Background decomposition of the CONUS+ run 1 data N. Ackermann et.al. 2608.12065 null
2026-08-12 Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL Martin Schuck et.al. 2608.12063 null
2026-08-12 Token-Level Credit Assignment Optimization for Generative Document Retrieval Xinpeng Zhao et.al. 2608.12049 null
2026-08-12 Energy-Dependent Dechanneling in Cu: Insights from Monte Carlo Channeling Simulations Przemyslaw Jozwik et.al. 2608.12017 null
2026-08-12 Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection Chaoran Chen et.al. 2608.11977 null
2026-08-12 LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation Zhixin Zhang et.al. 2608.11967 null
2026-08-12 LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured – Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence Po-Jen Ko et.al. 2608.11922 null
2026-08-12 HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Wenshuo Peng et.al. 2608.11913 null
2026-08-11 VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics Bowei Liu et.al. 2608.11201 null
2026-08-11 Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation Shiyu Xuan et.al. 2608.11191 null
2026-08-11 Scheduling Mixed RL Rollouts Beyond Prefix Locality Zetao Hong et.al. 2608.11152 null
2026-08-11 Mastering Stochastic OLG Models in Continuous Time Yves Achdou et.al. 2608.11134 null
2026-08-11 Sum rules and density-wave modes in spin-singlet fractional quantum Hall fluids Ritajit Kundu et.al. 2608.11133 null
2026-08-11 Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting Kiran Madhusudhanan et.al. 2608.11114 null
2026-08-11 Diffusion Quasi-Monte Carlo Jianlong Chen et.al. 2608.11055 null
2026-08-11 Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study Sepideh Saran et.al. 2608.11054 null
2026-08-11 Efficient Hypergradient Descent for Inverse Reinforcement Learning Nikita Sevriukov et.al. 2608.11052 null
2026-08-11 Jamming transition in an active exclusion process Kavita Jain et.al. 2608.11041 null
2026-08-11 A parametric framework for assessing and estimating sufficient follow-up time in cure models Luiz Silva-Resende et.al. 2608.11029 null
2026-08-11 Influence of interactions on the chiral effect in $1D$ Dirac semimetal Maksim Ulybyshev et.al. 2608.11004 null
2026-08-11 ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering Taojie Zhu et.al. 2608.10996 null
2026-08-11 Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes Zhaoyang Wei et.al. 2608.10954 null
2026-08-11 Laser-Diode LiFi With Diffused-Beam Optics: System-Level Modeling and a Cross-Validated ns-3 Simulation Framework Hussain Ahmad et.al. 2608.10950 null
2026-08-11 Threshold Structure of Optimal Policies in Restart POMDPs Konstantin Avrachenkov et.al. 2608.10936 null
2026-08-11 Endogeneity-Aware Cognitive Diagnostic Model for Multidomain Ordinal Assessments Zhiyu Huang et.al. 2608.10913 null
2026-08-11 Partially Observable Learning for Multi-Platform Dispatch Optimization Fengming Yao et.al. 2608.10897 null
2026-08-11 Enabling Scalable Kinesthetic Teaching via Observer-based Hand-guiding with Active Support Anna Tuma et.al. 2608.10847 null
2026-08-11 Deep reinforcement learning for separation control in turbulent wind-tunnel flow Sofia Avdiiv et.al. 2608.10829 null
2026-08-10 Realization Variance of Gravitational Wave Background Anisotropies from Shot Noise for Pulsar Timing Arrays Meng-Xiang Lin et.al. 2608.09929 null
2026-08-10 Consilience for Verifier-Free Test-Time Scaling Lecheng Kong et.al. 2608.09898 null
2026-08-10 The Earth Moves, But So Does the Bias: Systematic Upward Bias of the Wasserstein (Earth Mover’s) Distance and Permutation-Based Null Calibration Ho Ting et.al. 2608.09863 null
2026-08-10 RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance Dongchi Huang et.al. 2608.09853 null
2026-08-10 Mismatch Matters: On-Policy Distillation Beyond Token Agreement Zichao Yu et.al. 2608.09836 null
2026-08-10 Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation Yubo Jiang et.al. 2608.09826 null
2026-08-10 Parameter Exploration for RLVR via Variational Learning Vatsal Venkatkrishna et.al. 2608.09805 null
2026-08-10 A Singular Control Problem for Data Center Electricity Cost Minimization Rene Carmona et.al. 2608.09794 null
2026-08-10 Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Changhao Li et.al. 2608.09762 null
2026-08-10 SR-OPSD: Self-Referenced On-Policy Self-Distillation Zhuo Sun et.al. 2608.09745 null
2026-08-10 Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models Kevin Murphy et.al. 2608.09696 null
2026-08-10 Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance Logan Luna et.al. 2608.09628 null
2026-08-10 Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search Youssef A. Elhagrasy et.al. 2608.09622 null
2026-08-10 Bayesian Symbolic Regression with Entropic Reinforcement Learning Oussama Boussif et.al. 2608.09617 null
2026-08-10 FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving Guolei Huang et.al. 2608.09591 null
2026-08-10 AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning Yan Rong et.al. 2608.09559 null
2026-08-10 Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents Tianjun Pan et.al. 2608.09555 null
2026-08-10 Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning Yuting Liu et.al. 2608.09507 null
2026-08-10 Autoregressive Projective Quantum Monte Carlo: From a Hermitian to a Non-Hermitian Perspective Lavoisier Wah et.al. 2608.09496 null
2026-08-10 Walk-on-Spheres Monte Carlo and deep neural network approximations of elliptic PDEs with drift and killing Konrad Kleinberg et.al. 2608.09494 null
2026-08-07 SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Zongchuang Zhao et.al. 2608.07468 null
2026-08-07 CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity Ananya Sahu et.al. 2608.07460 null
2026-08-07 Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing Jiacheng Miao et.al. 2608.07437 null
2026-08-07 Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control Zhaoyu Zhu et.al. 2608.07433 null
2026-08-07 ResidencyRL: Reinforcement Learning in Simulated Clinical Environments Valentin Liévin et.al. 2608.07418 null
2026-08-07 LYRA: Label-Free Structural Synchronization and Resource Allocation for UAV Edge Networks Feng He et.al. 2608.07392 null
2026-08-07 A Formalization of the Laplace Transform and Its Inversion in Lean 4 Daniel Goldberg et.al. 2608.07384 null
2026-08-07 Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Haoyu Zheng et.al. 2608.07371 null
2026-08-07 Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks Taha Shieenavaz et.al. 2608.07335 null
2026-08-07 Learning Fault-Tolerant Locomotion with Adaptive Gait Timing Giovanbattista Gravina et.al. 2608.07328 null
2026-08-07 TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Ziheng Liu et.al. 2608.07314 null
2026-08-07 Learning Long-Term Educational Investment Policies under Residential Sorting Honglei Guo et.al. 2608.07295 null
2026-08-07 Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction Assaf Caftory et.al. 2608.07280 null
2026-08-07 Homojunction-induced thermopower enhancement in polymer films Zhen Xu et.al. 2608.07266 null
2026-08-07 Inverse reinforcement learning for indefinite mean-field social optimization with multiplicative noise Ying Cao et.al. 2608.07252 null
2026-08-07 Multiscale probing of a Hernquist-type environmental black hole spacetime with the Sgr A* shadow and S2 orbital dynamics Lai Zhao et.al. 2608.07229 null
2026-08-07 Learning Suffers More Than the Policy Class Under Partial Observability: A Closed-Form Analysis Idil Gözel et.al. 2608.07228 null
2026-08-07 Redshift Dependence of $H_0$ Dipole in Pantheon+ Supernovae M. H. Jalali-Kanafi et.al. 2608.07209 null
2026-08-07 Distribution of the Radius of Gyration for an ISAW Antony Lesage et.al. 2608.07195 null
2026-08-07 Momba: Network Modernization Improves Multi-Objective Reinforcement Learning Adam Štafa et.al. 2608.07180 null
2026-08-06 $ω$ -0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation Zhe Li et.al. 2608.06375 null
2026-08-06 RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Chenglong Wang et.al. 2608.06310 null
2026-08-06 Stochastic Dynamics on Persistence Diagram Space via Reinforcement Learning Farzana Nasrin et.al. 2608.06276 null
2026-08-06 DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models ZhiYan Hou et.al. 2608.06243 null
2026-08-06 Optimal Designs in Multicomponent Stress Strength Reliability for the Unit Generalized Rayleigh Distribution Rajat Das et.al. 2608.06214 null
2026-08-06 VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations Hisham Khalil et.al. 2608.06210 null
2026-08-06 EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Zishan Xu et.al. 2608.06197 null
2026-08-06 MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration Jia Xiong et.al. 2608.06183 null
2026-08-06 iARCS: Iterative Agentic RL for Controllable 3D Scene Generation Saugat Adhikari et.al. 2608.06161 null
2026-08-06 Quantum Amplitude Estimation for Travel Time Estimation in Stochastic Vehicle Routing Problems Xingyue Wang et.al. 2608.06145 null
2026-08-06 Large-Market Discipline in Combinatorial Double Auctions: No Assembly, Bundle Selection, and Complementarities Konstantinos E. Zachariadis et.al. 2608.06134 null
2026-08-06 Contextual Information Policy Optimization for Search Agents Xingyu Guo et.al. 2608.06128 null
2026-08-06 Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training Rui Li et.al. 2608.06125 null
2026-08-06 Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping Vaishnav Vaidheeswaran et.al. 2608.06105 null
2026-08-06 Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval Zelong Sun et.al. 2608.06060 null
2026-08-06 Criteria for Feasible Monte Carlo Stochastic Simulations of Bosonic Markovian Open Quantum Dynamics Toma Yoneya et.al. 2608.06056 null
2026-08-06 Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference Jiming Su et.al. 2608.06025 null
2026-08-06 ProDVI: Programmatic Dynamics Priors for Value Network Initialization Xinwei Liu et.al. 2608.06015 null
2026-08-06 OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction Jiahao Huang et.al. 2608.06013 null
2026-08-06 Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation He Kong et.al. 2608.05999 null
2026-08-05 Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning Boxiu Li et.al. 2608.05144 null
2026-08-05 Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning Jai Malegaonkar et.al. 2608.05111 null
2026-08-05 Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming Yanting Wang et.al. 2608.05108 null
2026-08-05 ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment Yijun Lu et.al. 2608.05102 null
2026-08-05 Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control Rohit Kumar Salla et.al. 2608.05084 null
2026-08-05 Generalized Glauber theorem for dark-matter axion and graviton detection Jakub Bręczewski et.al. 2608.05082 null
2026-08-05 Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Zheyuan Zhang et.al. 2608.05080 null
2026-08-05 Exact Model-Free Policy Iteration for Co-safe LTL Planning Zetong Xuan et.al. 2608.05047 null
2026-08-05 Exact simulation of diffusions and improved algorithms for log-concave sampling Fan Chen et.al. 2608.05022 null
2026-08-05 ORACLE: A Multi-Objective Reinforcement Learning-Based Analog Circuit Design Optimizer with Large Language Models-Guided Exploration Osei Brempong et.al. 2608.04999 null
2026-08-05 A Pairwise Differencing Distribution Regression Approach for Network Models Gabriela Miyazato Szini et.al. 2608.04983 null
2026-08-05 Delocalized Coupled-Cluster Theory for Polaron Structure and Dynamics Hamlin Wu et.al. 2608.04979 null
2026-08-05 WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models Bohai Gu et.al. 2608.04964 null
2026-08-05 SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts Nhat Minh Pham et.al. 2608.04962 null
2026-08-05 State2State: Environment-Derived Mid-Training for LLM Agents Xuanyu Lei et.al. 2608.04934 null
2026-08-05 PRIMAL3: Pathfinding via Reinforcement and Imitation Multi-Agent Learning - Leveraging LaCAM3 Chengyang He et.al. 2608.04905 null
2026-08-05 Structural Chirality from Short-Range Order in Heteroanionic Materials Benjamin J. Morgan et.al. 2608.04841 null
2026-08-05 Circular polarization as a probe of cloud properties and asymmetries in giant exoplanet atmospheres M. B. Michaelis et.al. 2608.04837 null
2026-08-05 DEGR: Dual Exploration-Driven Generative Re-Ranking for Adaptive Cross-Request Context Bridging Binglei Zhao et.al. 2608.04809 null
2026-08-05 Spatial distribution of water ice in the protoplanetary silhouette disk d216-0939 J. S. Martin et.al. 2608.04803 null
2026-08-04 TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning Changle Qu et.al. 2608.04007 null
2026-08-04 Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Zhen Fang et.al. 2608.03979 null
2026-08-04 Stochastic Multiple Shooting Trajectory Optimization via Sequential Local Policy Evaluation Ashwin Gupta et.al. 2608.03978 null
2026-08-04 Information-Geometric Forward Policy Training in GFlowNets Yordan Raykov et.al. 2608.03967 null
2026-08-04 Kappa distributions as asymptotic marginals of exponential family ensembles Sergio Davis et.al. 2608.03960 null
2026-08-04 Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level Jetson Ekansh Singh et.al. 2608.03938 null
2026-08-04 Simulation-Based Neural Policies for Portfolio Choice: Architecture, Training, and Interpretability Jules Viard et.al. 2608.03933 null
2026-08-04 Latent Reward Registers for Diffusion Preference Alignment Yuanshen Guan et.al. 2608.03929 null
2026-08-04 Accelerated quantum Monte Carlo simulations of the attractive Hubbard model on the kagome lattice Jie Zhang et.al. 2608.03894 null
2026-08-04 Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning Pyrros Koussios et.al. 2608.03875 null
2026-08-04 EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning Shuoqin Zhang et.al. 2608.03872 null
2026-08-04 FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs Amin Farajzadeh et.al. 2608.03852 null
2026-08-04 Local magnetic order in vacancy-disrupted spin ice Ho2TiO5 Raju Baral et.al. 2608.03850 null
2026-08-04 History Matters: Meta-policy Delegation with Heterogeneous Multi-agent Reinforcement Learning Ziqing Lu et.al. 2608.03833 null
2026-08-04 Explicit formulas of strict BV relaxed energies for polyconvex functionals with linear growth Domanico Mucci et.al. 2608.03821 null
2026-08-04 AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding Yuxiang Duan et.al. 2608.03779 null
2026-08-04 GORDON: Graph-based Object-centric Rewards for Decomposition of Long-Horizon Manipulation Andrea Protopapa et.al. 2608.03753 null
2026-08-04 Inverse Design of Quantum Control Sequences with Fourier Neural Operators Anastasia Pipi et.al. 2608.03702 null
2026-08-04 PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Chenghua Wang et.al. 2608.03682 null
2026-08-04 DiagLoop: A Counterfactual Data Flywheel with Stage-Localized Reinforcement for Diagnostic LLMs Jian Zhang et.al. 2608.03674 null
2026-08-04 CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning Jian Zhang et.al. 2608.03673 null
2026-08-04 Quantum Impurities as Probes of Finite-Temperature Fluctuations in Two-Dimensional Bose Gases Victor Velasco et.al. 2608.03665 null
2026-08-04 Group Perspective Matters: Regulating Debate Relationships Can Mitigate Blind Conformity in Multi-Agent Debate Hao Wu et.al. 2608.03648 null
2026-08-04 Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details Maksymilian Wolski et.al. 2608.03644 null
2026-08-04 Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR Yuan Xie et.al. 2608.03610 null
2026-08-04 Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents William Bolton et.al. 2608.03606 null
2026-08-04 SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Kejian Zhu et.al. 2608.03573 null
2026-08-04 Robust General Utility for Reinforcement Learning Zixuan Liu et.al. 2608.03562 null
2026-08-04 Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Kunbin Xu et.al. 2608.03545 null
2026-08-04 Training Documents Reranker with Search Rubrics for Deep Research Agent Wenhan Liu et.al. 2608.03527 null
2026-08-04 When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs Omatharv Bharat Vaidya et.al. 2608.03506 null
2026-08-04 Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks Christophe D. Hounwanou et.al. 2608.03502 null
2026-08-04 Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution Weichen Xu et.al. 2608.03483 null
2026-08-04 ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning Xiuhui You et.al. 2608.03468 null
2026-08-04 When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO Zhe Cao et.al. 2608.03467 null
2026-08-04 Analysis of inverse stochastic resonance: Effects of neural excitability and timescale separation Marius E. Yamakou et.al. 2608.03454 null
2026-08-03 Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework Junjie Yin et.al. 2608.02599 null
2026-08-03 CCAT: Optical Design of the 410 GHz Prime-Cam Module Tilak M. Patel et.al. 2608.02579 null
2026-08-03 Analytic Planning under Uncertainty with Moment Closure Shishir Sharma et.al. 2608.02519 null
2026-08-03 RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States Yi Yang et.al. 2608.02508 null
2026-08-03 FCC precision requests: challenges for Monte Carlos and phenomenology tools Z. Was et.al. 2608.02476 null
2026-08-03 Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+ Christoph Weinhuber et.al. 2608.02454 null
2026-08-03 Foundations of Reinforcement Learning and Control:Connections and New Perspectives Claire Vernade et.al. 2608.02433 null
2026-08-03 Agentic Incident Response through Digital Twin-Enhanced Multiscale Planning Yiran Gao et.al. 2608.02422 null
2026-08-03 Antares: Foundation Models for Agentic Vulnerability Localization Supriti Vijay et.al. 2608.02407 null
2026-08-03 Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training Zhiyuan Wang et.al. 2608.02391 null
2026-08-03 Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcement Learning Patrick Oberlin et.al. 2608.02379 null
2026-08-03 Qwen-CUA: Native Computer Use for (almost) Everything Dunjie Lu et.al. 2608.02352 null
2026-08-03 Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection Patrick Helm et.al. 2608.02343 null
2026-08-03 Diffusion Policy with Behavioral Advantage Correction for Offline Reinforcement Learning Botao Dong et.al. 2608.02332 null
2026-08-03 Learnable yet not simulable: a quantum resource theory of learning models Xinbiao Wang et.al. 2608.02325 null
2026-08-03 BRiG-AFA: Bellman Risk-to-Go Learning for Non-Myopic Active Feature Acquisition Jiaorong Feng et.al. 2608.02305 null
2026-08-03 Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories Shuai Shao et.al. 2608.02276 null
2026-08-03 Valley-controlled chiral magnetism in transition metal dichalcogenide monolayers Igor S. Krivenko et.al. 2608.02260 null
2026-08-03 Phase-Drift Limits and Adaptive Quadrature Readout in Programmable Photonic Processors Gökhan Elmas et.al. 2608.02249 null
2026-08-03 VC-Tooler: Learning Compositional and Adaptive Visual Tool Use Yizheng Wu et.al. 2608.02217 null
2026-08-02 adabay: an R package for rapid evaluation and calibration of Bayesian group sequential designs across common endpoint types Zhangyi He et.al. 2608.01068 null
2026-08-02 Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception Xinheng Han et.al. 2608.01055 null
2026-08-02 Cloud-ScPO: Hidden-State Geometry for Semi-Supervised Preference Optimization in LLM Reasoning Yuzhou Liu et.al. 2608.01014 null
2026-08-02 RL Bootstrapping of OpenVLA-OFT for a Novel Robot Embodiment Damir Nurtdinov et.al. 2608.01013 null
2026-08-02 MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models Ofir Ben Shoham et.al. 2608.01012 null
2026-08-02 Search-GRT: Guided Retrieval Training of Search Agents to Optimize for Complex Question Answering Aounon Kumar et.al. 2608.00974 null
2026-08-02 PROGRESS: Coverage-guided RL to Train Search-augmented LLM Agent Sudipta Paul et.al. 2608.00969 null
2026-08-02 Gaokerena: A Small Persian Medical Language Model Family Mehrdad Ghassabi et.al. 2608.00932 null
2026-08-02 Battery Storage Co-Optimization in Day-Ahead and Real-Time Markets with Bayesian Optimization Thiha Aung et.al. 2608.00911 null
2026-08-02 Learning Not to Optimize: Physics-Informed Action-Space Reshaping for Intent-Based Network Control Zuyuan Zhang et.al. 2608.00908 null
2026-08-01 Responsible AI and Algorithmic Adoption in Methodology Development for National Statistical Offices Siu-Ming Tam et.al. 2608.00896 null
2026-08-01 Bicycle Acrobatics with Reinforcement Learning Shamel Fahmi et.al. 2608.00880 null
2026-08-01 Minute-Scale Training for Microrobot Navigation Yinghan Sun et.al. 2608.00854 null
2026-08-01 Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback Keertana Chidambaram et.al. 2608.00816 null
2026-08-01 webSME: An online tool to infer stellar parameters and abundances Johannes Puschnig et.al. 2608.00787 null
2026-08-01 Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Zhuowen Han et.al. 2608.00782 null
2026-08-01 Variational Inference Using a Differentiable Multigrid Linear Solver Andrés Ramírez et.al. 2608.00760 null
2026-08-01 LUT: Latent Utility Training for Visual Reasoning Jiaxuan Kang et.al. 2608.00743 null
2026-08-01 Quantitative Particle Approximation for Controlled Nonlinear Filtering Erhan Bayraktar et.al. 2608.00686 null
2026-08-01 HetGPS: Scalable Graph Multi-Agent Reinforcement Learning with Physics-Anchored Adaptive Safety for EV Charging Xiangwei Wang et.al. 2608.00679 null
2026-07-31 An optimal quadratic estimator for window-free cosmic shear power spectra Taisei Terawaki et.al. 2607.29652 null
2026-07-31 CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding Wenxin Tang et.al. 2607.29637 null
2026-07-31 RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning Qian Wang et.al. 2607.29622 null
2026-07-31 When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Luca Viano et.al. 2607.29617 null
2026-07-31 WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Senyu Fei et.al. 2607.29613 null
2026-07-31 Recursive rounding of sample size estimation for multi-fidelity Monte Carlo Jiaxing Liang et.al. 2607.29607 null
2026-07-31 Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment Yanwei Jia et.al. 2607.29593 null
2026-07-31 Spindrift: Learning quantum degeneracy from thermal purity in restricted path integral Monte Carlo Jarvist Moore Frost et.al. 2607.29590 null
2026-07-31 DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat Ismayil Ismayilov et.al. 2607.29577 null
2026-07-31 LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Manith Adikari et.al. 2607.29559 null
2026-07-31 STAGE: STyle-controllable Action GEneration for personalized autonomous driving Zihao Liu et.al. 2607.29517 null
2026-07-31 DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search Jiayang Niu et.al. 2607.29491 null
2026-07-31 Discovery Sensitivity for a Counting Experiment with Background Uncertainty Enzo Canonero et.al. 2607.29436 null
2026-07-31 Explore Beyond the Boundary Using Entropic Information Bumgeun Park et.al. 2607.29419 null
2026-07-31 OnlineCache: Learning Dynamic Caching Policies with Error Correction for Efficient Diffusion Inference Zhikang Xie et.al. 2607.29398 null
2026-07-31 pylhe: A Lightweight Python interface to Les Houches Event files Alexander Puck Neuwirth et.al. 2607.29352 null
2026-07-31 BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning BWM Team et.al. 2607.29302 null
2026-07-31 Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification Anders Jonsson et.al. 2607.29294 null
2026-07-31 Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation Yongshi Ye et.al. 2607.29287 null
2026-07-31 TRACT: Temporally Routed Action Chunks with Chronological Phase Authority for Contact-Rich Manipulation Jiahao Liu et.al. 2607.29285 null
2026-07-30 ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine Yukang Cao et.al. 2607.28625 null
2026-07-30 OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Qiushi Sun et.al. 2607.28609 null
2026-07-30 Beacon: Knowing When and How to Perform Agentic Visual Reasoning Qixun Wang et.al. 2607.28595 null
2026-07-30 ABC methods for IoT Emitter Geolocalisation using LEO Satellite Doppler Measurements B. Ristic et.al. 2607.28585 null
2026-07-30 $β$ -OPSD: Deriving with Policy Optimization, Training with Self-Distillation Jiawei Xu et.al. 2607.28582 null
2026-07-30 X-NavDP: Generalizing Navigation Diffusion Policy to Novel Behavior and Embodiments with Group Q-score Reweighted Matching Tianyu Yang et.al. 2607.28560 null
2026-07-30 Can Vision-Language Models Reason about AI Edits in Images? Darsha Udayanga et.al. 2607.28464 null
2026-07-30 Cybersecurity Detection Classification with Reasoning-enabled Language Models Amol Khanna et.al. 2607.28460 null
2026-07-30 SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute Hongyu Chen et.al. 2607.28457 null
2026-07-30 Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs Ankur Naskar et.al. 2607.28390 null
2026-07-30 Fast optics-based modeling enabling large-scale optimization of the H4 and M2 beamlines in the CERN SPS North Area Giovanni Dal Maso et.al. 2607.28370 null
2026-07-30 HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks Tiangang Li et.al. 2607.28301 null
2026-07-30 Bootstrap inference in autoregressive duration models Giuseppe Cavaliere et.al. 2607.28294 null
2026-07-30 Synchronization, Kinematic Waves and Spike-Phase-Separation in Feedback Ising Neural Networks on Heterogeneous Graphs Anna Poggialini et.al. 2607.28275 null
2026-07-30 Uncertainty quantification for trustworthy deep learning: Methods and measures H. Martin Gillis et.al. 2607.28248 null
2026-07-30 Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Hanzhang Zhou et.al. 2607.28227 null
2026-07-30 Semi-supervised Hopfield model: Theoretical and Numerical results Linda Albanese et.al. 2607.28173 null
2026-07-30 LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning Mohand Mezmaz et.al. 2607.28135 null
2026-07-30 FinSMART: Financial Sentiment Analysis for Algorithmic Trading through Market-Aligned Reinforcement Learning Giorgos Iacovides et.al. 2607.28127 null
2026-07-30 Approximate sampling from decoded quantum interferometry via Markov chain Monte Carlo methods Elies Gil-Fuster et.al. 2607.28120 null
2026-07-29 Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning? Perry Dong et.al. 2607.27203 null
2026-07-29 Improved Methods for Determining Quantum Error Correcting Code Performance and Fault Tolerance Michael Mullan et.al. 2607.27153 null
2026-07-29 Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes Zuyuan Zhang et.al. 2607.27132 null
2026-07-29 Formation of $\mathrm{L}1_2$-ordered $γ’$-$\mathrm{Ni}_3\mathrm{Al}$ precipitates in ternary Cu-Ni-Al alloys modelled using an ab initio concentration wave theory and atomistic simulations Christopher D. Woodgate et.al. 2607.27108 null
2026-07-29 Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering Nicolas Béreux et.al. 2607.27077 null
2026-07-29 Correlated Chance Sampling for Monte Carlo Counterfactual Regret Minimization Boning Li et.al. 2607.27035 null
2026-07-29 RL $^2$ -VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models Derek Ming Siang Tan et.al. 2607.26991 null
2026-07-29 Belief-Guided Decision Making with Uncertainty Gating in the Game of Go Mehrad Yaghoubi et.al. 2607.26946 null
2026-07-29 BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories Zhe Liu et.al. 2607.26914 null
2026-07-29 Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising Zipeng Chen et.al. 2607.26893 null
2026-07-29 The HRT Conjecture for Symmetric Configurations and Real-Valued Functions Shuang Guan et.al. 2607.26878 null
2026-07-29 Revising Indirect Dark Matter Constraints with Updated Astrophysical $J$ -Factor Priors Giacomo D’Amico et.al. 2607.26876 null
2026-07-29 SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Jianze Wang et.al. 2607.26873 null
2026-07-29 ReCo: Reweighting GRPO Against Distributional Concentration Junoh Park et.al. 2607.26862 null
2026-07-29 Uniform Convergence of Generalized Conditional Fréchet Means with Applications to Weighted Fréchet Aggregation and Exceedance Set Estimation Houren Hong et.al. 2607.26837 null
2026-07-29 SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution Zhiyuan Yao et.al. 2607.26784 null
2026-07-29 Automated NRQCD and NRQED simulations of quarkonium and leptonium production with P-wave states and physical-mass effects Luca Maxia et.al. 2607.26739 null
2026-07-29 The CROSS experiment: detector construction, background projection, and sensitivity to $^{100}$Mo $0\nu2β$ decay D. Auguste et.al. 2607.26732 null
2026-07-29 Robust Interpolated Quantile Estimators: Asymptotic Theory and Efficiency Saïd Maanan et.al. 2607.26714 null
2026-07-29 Reinforcement Learning applied to Optimization of LHC beams in the CERN Proton Synchrotron Joel Axel Wulff et.al. 2607.26697 null
2026-07-29 Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL Mingxuan Che et.al. 2607.26680 null
2026-07-29 Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection Yikun Li et.al. 2607.26656 null
2026-07-29 Bayesian nonparametric estimation of correlated gravitational wave detector network noise using matrix-gamma process priors Yixuan Liu et.al. 2607.26619 null
2026-07-28 S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information Kaneyoshi Hiratsuka et.al. 2607.26047 null
2026-07-28 Photonuclear Neutron Production in OpenMC: Verification Against MCNPX, FLUKA, and a First-Collision Analytical Solution Lorenzo Loi et.al. 2607.26045 null
2026-07-28 Reinformed Dreamer: An Asymmetric World Model Efficiently Trained through Latent Guidance Gaspard Lambrechts et.al. 2607.26040 null
2026-07-28 Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing Fengxiang Wang et.al. 2607.25993 null
2026-07-28 Physics-Aware End-to-End Deep Reinforcement Learning for Quadcopter Control with Actuator Dynamics Ya-Chia Shen et.al. 2607.25985 null
2026-07-28 Schrödinger’s Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics Timy Phan et.al. 2607.25984 null
2026-07-28 Superfluidity without charge order in the attractive Hubbard model on the kagome lattice Xiaodong Jin et.al. 2607.25983 null
2026-07-28 Reinforcement Learning for Code Optimization Pierre Chambon et.al. 2607.25970 null
2026-07-28 Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification Chenrui Shi et.al. 2607.25904 null
2026-07-28 RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation Guohong Mu et.al. 2607.25901 null
2026-07-28 The calibration of large-radius jets using the Run 2 dataset with the ATLAS detector ATLAS Collaboration et.al. 2607.25893 null
2026-07-28 Scaling universal Fermi network toward ground states: A diffusion-Monte-Carlo assessment Yu-Sheng Li et.al. 2607.25872 null
2026-07-28 Laser power transmission in space: Plasma-based power cell Li Lin et.al. 2607.25843 null
2026-07-28 General Relativistic Entropic Acceleration at the perturbation level: a CLASS implementation and first Boltzmann-code constraints Simone D’Onofrio et.al. 2607.25841 null
2026-07-28 WarmTuner: Program-Specific Warm Starts for Compiler Autotuning via Offline-to-Online Reinforcement Learning Tianlu Qiao et.al. 2607.25831 null
2026-07-28 Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL Jiabao Ji et.al. 2607.25816 null
2026-07-28 Variance-Reduced Conditional Gradient Methods under Markovian Sampling for Nonconvex Composite Optimization Zhaojun Peng et.al. 2607.25785 null
2026-07-28 A Hierarchical Optimisation Framework for Integrated Electric-Hydrogen-Transport Systems Fulong Yao et.al. 2607.25776 null
2026-07-28 Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning Yu Su et.al. 2607.25754 null
2026-07-28 A systematic evaluation of machine learning classifiers for event-by-event background rejection in LAFOV PET scanners Konrad Klimaszewski et.al. 2607.25732 null
2026-07-27 Sign-optimized Quantum Monte Carlo Julius S. Herz et.al. 2607.24679 null
2026-07-27 Explainable Reinforcement Learning via Physics-Aware Policy Distillation Shaker Al-Tamari et.al. 2607.24672 null
2026-07-27 Kimi K3: Open Frontier Intelligence Kimi Team et.al. 2607.24653 null
2026-07-27 Next-to-leading order FsQED corrections to radiative pion pair production Carlo M. Carloni Calame et.al. 2607.24642 null
2026-07-27 Emergence of the halo in $^{11}$ Li from full nuclear many-body dynamics Yilong Yang et.al. 2607.24636 null
2026-07-27 Recommended Second Virial Coefficients for Nitrogen and Oxygen Robert Hellmann et.al. 2607.24634 null
2026-07-27 PYPM-GGD: Pitman-Yor Process Mixture with Generalized Gaussian Density using ADAM Kart-Leong Lim et.al. 2607.24583 null
2026-07-27 Constraining young massive cluster properties with radio-continuum observations: The Arches cluster M. Cano-González et.al. 2607.24580 null
2026-07-27 Evaluating Fuzz Testing for Reinforcement Learning Agents Zhibin Kang et.al. 2607.24577 null
2026-07-27 The balance between compactness and forecast accuracy of data-driven latent-space reduced-order models in controlled wake flows Alberto Solera-Rico et.al. 2607.24569 null
2026-07-27 Entanglement Distillation and Swapping Scheduling in Quantum Repeaters with Noisy Memories Siddharth Chander et.al. 2607.24557 null
2026-07-27 VulnGym: Evaluating Vulnerability Management Strategies against Advanced Persistent Threats Sofia Della Penna et.al. 2607.24552 null
2026-07-27 Amortized Posteriors for Estimation of Material Constitutive Parameters from Multimodal Measurements on Small Punch Tests Mohammad Ali Seyed Mahmoud et.al. 2607.24534 null
2026-07-27 What do Reward Models Memorize? Ivo Verhoeven et.al. 2607.24484 null
2026-07-27 Locally Robust Kernel Specification Tests for Conditional Moment Restrictions Juan Carlos Escanciano et.al. 2607.24382 null
2026-07-27 Bayesian Feature Extraction using Gaussian and Diffused-gamma Priors for High Dimensional Spatio-Temporal Data Garrett Frady et.al. 2607.24378 null
2026-07-27 Möbius-Invariant Goodness-of-Fit Tests for the Spherical Cauchy Model Diego Bolón et.al. 2607.24376 null
2026-07-27 Continual-RL for Generalization in Autonomous Racing on the RoboRacer Platform Joel Siegert et.al. 2607.24320 null
2026-07-27 CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models Mingxuan Sun et.al. 2607.24312 null
2026-07-27 Learning Adaptive Multi-Task Guidance, Navigation, and Control via Hypernetworks Ricard Marsal I Castan et.al. 2607.24292 null
2026-07-26 Robust estimation of the autocorrelation function via forward ratios A. Montañés et.al. 2607.23744 null
2026-07-26 Zing: Social Mind for LLMs Zing Team et.al. 2607.23740 null
2026-07-26 TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs Muhammad Umar Farooq Qaisar et.al. 2607.23734 null
2026-07-26 Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning Zahra Abdalla Elashaal et.al. 2607.23726 null
2026-07-26 Offline-Online Curriculum RL for Multimodal Reasoning Wendi Deng et.al. 2607.23700 null
2026-07-26 Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set Heyang Zhao et.al. 2607.23679 null
2026-07-26 Optimal Reward Shaping: Autonomous Car Parking Case Study Emre Özkaya et.al. 2607.23617 null
2026-07-26 Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning Wenxuan Zhang et.al. 2607.23605 null
2026-07-26 Hierarchical Gaussian-process test of DESI’s dynamical dark-energy preference Yu-Hao Mu et.al. 2607.23593 null
2026-07-26 Anticipatory Risk-Guided Reinforcement Learning for Safe Flight Through Dynamic Clutter Yuchao Mei et.al. 2607.23565 null
2026-07-26 ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness Qiao Yan et.al. 2607.23537 null
2026-07-26 LEACL: LLM-Enhanced Automatic Curriculum Learning for Reinforcement Learning in Long-Horizon Manipulation Tasks Faraz Heravi et.al. 2607.23515 null
2026-07-26 Learning Sampling Parameters for Diffusion Models Arisrei Lim et.al. 2607.23488 null
2026-07-26 Simultaneous Color Glass Condensate fit to deep inelastic scattering and forward hadron production at HERA, RHIC, and the LHC Piotr Korcyl et.al. 2607.23485 null
2026-07-26 Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning Minh Vu et.al. 2607.23474 null
2026-07-26 PRISM: Polynomial Representations for Interaction-Structured Motor Control Seung Hyun Lee et.al. 2607.23473 null
2026-07-26 When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design Longying Wen et.al. 2607.23469 null
2026-07-26 Learning to Optimize: Joint Routing and Flow Allocation on Sparse Non-Euclidean Networks Haomiao Sun et.al. 2607.23467 null
2026-07-26 Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations Young Hyun Cho et.al. 2607.23434 null
2026-07-26 LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction Xiao You et.al. 2607.23420 null
2026-07-24 Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Siyuan Huang et.al. 2607.22529 null
2026-07-24 Explainable Reinforcement Learning for assisting Air Traffic Controllers Anduel Mehmeti et.al. 2607.22525 null
2026-07-24 Learning to Prepare Molecular Ground States with Transformer Models Alex Koziell-Pipe et.al. 2607.22468 null
2026-07-24 Scaling Results for Piecewise Deterministic Monte Carlo : A Survey Joris Bierkens et.al. 2607.22449 null
2026-07-24 Nonlinear Boosting with Multiple Testing in High-Dimensional Generalised Linear Models with Binary Responses Charisios Grivas et.al. 2607.22440 null
2026-07-24 Highly indistinguishable photons from a tin-vacancy spin qubit in diamond Dennis Herrmann et.al. 2607.22439 null
2026-07-24 Conformal Constraint Tightening for Chance-Constrained Motion Planning with Unknown Dynamics Shubham Natraj et.al. 2607.22409 null
2026-07-24 Active few-shot segmentation by reinforcing data selection Chenlan Zhao et.al. 2607.22371 null
2026-07-24 A Small-Noise Analysis of Controlled Functional Differential Equations with Gaussian Noise David Criens et.al. 2607.22362 null
2026-07-24 Integrated Order Dispatching and Routing for Last-Mile Pickup via Deep Reinforcement Learning Yida Xu et.al. 2607.22356 null
2026-07-24 A Hierarchical Likelihood Model for Non-linear Inverse Problems under Additive and Multiplicative Noise Nicolas Goeman et.al. 2607.22330 null
2026-07-24 The macroscopic precession model of quasi-periodic oscillations for rotating compact objects Orlando Luongo et.al. 2607.22322 null
2026-07-24 Learning Bidirectional Causal Interactions with Heteroscedastic Neural Networks Masahiro Tanaka et.al. 2607.22313 null
2026-07-24 Hidden Truchet Architecture in Zinc $p$ -Hydroxybenzoate Hunter J. Windsor et.al. 2607.22307 null
2026-07-24 smartcor: Intelligent Correlation Method Selection for Mixed Variable Types M Harshvardhan et.al. 2607.22285 null
2026-07-24 When Can a Cavity Move a Mott Transition? A Spectral-Density Criterion within Gutzwiller Theory Nikhil Vamsodharakan Seshadri et.al. 2607.22283 null
2026-07-24 Unbiased Diffusion Monte Carlo for non local operators Carlos Rodriguez Perez et.al. 2607.22273 null
2026-07-24 General Value Functions for Remaining Useful Life and Failure-Mode Prediction Hao Yan et.al. 2607.22268 null
2026-07-24 Safe Learning Predictive Control for Ego-World Robotic Systems Davide Valenti et.al. 2607.22225 null
2026-07-24 LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR Xudong Liu et.al. 2607.22200 null
2026-07-23 Parallel Tempered Metadynamics for full QCD Timo Eichhorn et.al. 2607.21575 null
2026-07-23 MIRROR: Learning from the Other View for Multi-Modal Reasoning Wen Ye et.al. 2607.21552 null
2026-07-23 Bayesian evidence adaptive pursuit to identify neutron sources with scatter-based spectrometers David Breitenmoser et.al. 2607.21543 null
2026-07-23 Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control Xin Chen et.al. 2607.21520 null
2026-07-23 Group boarding for airplanes: benchmarking static policies and optimizing dynamic assignment with deep reinforcement learning Minyu Shen et.al. 2607.21512 null
2026-07-23 Strong correlations and local self-energies from on-site ensembles Alberto Carta et.al. 2607.21490 null
2026-07-23 Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections Gil Lifshits et.al. 2607.21488 null
2026-07-23 A story about a tipsy kangaroo: Reversible jump MCMC for model selection in the analysis of gravitational-wave signals from the coalescence of compact objects Anna Puecher et.al. 2607.21484 null
2026-07-23 AREX: Towards a Recursively Self-Improving Agent for Deep Research Shuqi Lu et.al. 2607.21461 null
2026-07-23 Stochastic Quantization as Optimal Control Lingxiao Wang et.al. 2607.21436 null
2026-07-23 PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning Yipeng Shi et.al. 2607.21419 null
2026-07-23 DISCO: Distributed Spectrum Compliance and Orchestration for Scalable IoT Coexistence Lyes Saad Saoud et.al. 2607.21387 null
2026-07-23 3D Uncertainty Quantification for the Photo-Acoustic Tomography Babak Maboudi Afkham et.al. 2607.21373 null
2026-07-23 How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-Tuning Kaizhen Tan et.al. 2607.21351 null
2026-07-23 Uniformly Consistent Semi-nonparametric Demand Estimation with Micro-Data Richard Grigorian et.al. 2607.21323 null
2026-07-23 Expert Behavior Prior Reinforcement Learning Gong Gao et.al. 2607.21302 null
2026-07-23 Determination of fundamental properties of nitrogen from first principles. III. Temperature and frequency dependence of the molecular polarizability and magnetic susceptibility Jakub Lang et.al. 2607.21261 null
2026-07-23 FORGE-plus: Force-Budgeted Recovery for Contact-Rich Assembly with a Frozen LLM Supervisor Kyupaeck Jeff Rah et.al. 2607.21227 null
2026-07-23 A Kuramoto phase model to explore the synchronisation of a network of circadian clocks Franck Delaunay et.al. 2607.21214 null
2026-07-23 Intrinsic coupling between transverse spherocity and elliptic flow in heavy-ion collisions Subikash Choudhury et.al. 2607.21161 null
2026-07-21 Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning Lizhe Fang et.al. 2607.19345 null
2026-07-21 OmniReasoner: Thinking with Long Audio-Video via Native Tool Use Yu Chen et.al. 2607.19339 null
2026-07-21 ISO: An RLVR-Native Optimization Stack Hanqing Zhu et.al. 2607.19331 null
2026-07-21 Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information Priyank Agrawal et.al. 2607.19313 null
2026-07-21 GARTFIMA Models: A Class of Observation-Driven Models with Tempered Fractional Dynamics Guilherme Pumi et.al. 2607.19311 null
2026-07-21 Stochastic Multi-Objective Kinodynamic Planning Against Adversaries Thomas Marshall Vielmetti et.al. 2607.19284 null
2026-07-21 A Reinforcement-Learning-Augmented Liquid-Fueled Reactor Network Model for Predicting Lean Blowout in Gas Turbine Combustors Philip John et.al. 2607.19281 null
2026-07-21 S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning Kshitij Kumar Srivastava et.al. 2607.19232 null
2026-07-21 The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation Michael Jungo et.al. 2607.19226 null
2026-07-21 Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards Xuefeng Jin et.al. 2607.19219 null
2026-07-21 Conservative Query and Adaptive Regularization for Offline RL Under Uncertainty Estimation Li-Rong Zhou et.al. 2607.19199 null
2026-07-21 ATLAS: A Foundation Neural Sampler for Amorphous Materials Mouyang Cheng et.al. 2607.19198 null
2026-07-21 Numerical methods for Langevin-type SPDE: an implicit Milstein approach and multilevel Monte Carlo techniques Sascha Portaro et.al. 2607.19188 null
2026-07-21 Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning Aixiu An et.al. 2607.19181 null
2026-07-21 Coherence in Control: Bridging Many-Core Mapping and Routing through Cost Unification Guochu Xiong et.al. 2607.19158 null
2026-07-21 Parallel Noising in Neural Markov Logic Networks Peter Jung et.al. 2607.19126 null
2026-07-21 Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning Ubayd Ali Bapoo et.al. 2607.19117 null
2026-07-21 Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation Mingxuan Ouyang et.al. 2607.19044 null
2026-07-21 Pricing options on illiquid assets using liquid market benchmarks: an application to energy markets Federico Aluigi et.al. 2607.19030 null
2026-07-21 Bayesian Sequential Quantum Amplitude Estimation for Rare-Event Structural Failure Probability Alireza Tabarraei et.al. 2607.18996 null
2026-07-20 On off-diagonal operators in matrix-weighted spaces David Cruz-Uribe et.al. 2607.18175 null
2026-07-20 OR Else: A Differentiable Trust Region for Policy Optimization Chinmay Rane et.al. 2607.18163 null
2026-07-20 Isaac Sim-to-Real: Reinforcement Learning based Locomotion for Quadrupeds Jordan Dowdy et.al. 2607.18135 null
2026-07-20 LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Tianzhu Ye et.al. 2607.18110 null
2026-07-20 Importance Sampling and PCA for Finding Failures in Commercial Autonomous Vehicles Hailey Warner et.al. 2607.18106 null
2026-07-20 Study of ordering in (MoCrTi) $_{100-x}$Al$_x$ refractory high-entropy alloys using machine learning interatomic potential Jiyao Zhang et.al. 2607.18099 null
2026-07-20 Real-Time Flight Test Maneuver Selection with Monte Carlo Tree Search Nicholas E. Bostock et.al. 2607.18089 null
2026-07-20 Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection Haochen Zhao et.al. 2607.18080 null
2026-07-20 Generalised Bellman recurrence and three dualities in sequential decision-making Fernando E. Rosas et.al. 2607.18077 null
2026-07-20 Predicting subjective rage and facial expressions in human driving: A Bayesian network approach with beta-distributed nodes Zaïra Méndez-Porcar et.al. 2607.18030 null
2026-07-20 MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models Martino M. L. Pulici et.al. 2607.18006 null
2026-07-20 PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning Daegyeong Roh et.al. 2607.18004 null
2026-07-20 AlphaZeroBeta: Deep Reinforcement Learning for Market-Neutral Portfolios Boris Belyakov et.al. 2607.18001 null
2026-07-20 Long-time behavior and turnpike properties of linear-quadratic graphon mean field control problems Erhan Bayraktar et.al. 2607.18000 null
2026-07-20 Gate-tunable giant anomalous Hall effect in magnetic topological insulator bilayer Basavaraja G et.al. 2607.17988 null
2026-07-20 Information-Based Exploration via Random Features for Reinforcement Learning Waris Radji et.al. 2607.17981 null
2026-07-20 MEVION: Low-Cost Open-Source Data Collection System for Powerful and High-Speed Dual-Arm Manipulation Kento Kawaharazuka et.al. 2607.17970 null
2026-07-20 A Geometric Perspective on Stabilizing Value Conflict Resolution Saket Reddy et.al. 2607.17946 null
2026-07-20 Probabilistic Stellar Age Estimation for Gaia XP Stars with NGBoost Xiaokun Hou et.al. 2607.17932 null
2026-07-20 Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization Zijian Zhao et.al. 2607.17924 null
2026-07-17 Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Ce Zhang et.al. 2607.16189 null
2026-07-17 Handroid: Bridging Dexterous Hand and Humanoid Ruogu Li et.al. 2607.16187 null
2026-07-17 Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems Matteo Tomasetto et.al. 2607.16177 null
2026-07-17 When Does Muon Help Agentic Reinforcement Learning? Kai Ruan et.al. 2607.16169 null
2026-07-17 Radiopurity material assays and radiation exposure projections for superconducting qubit measurements at SNOLAB Y. Ahmed et.al. 2607.16151 null
2026-07-17 Learning Standard Model structure from LHC data with Riemannian flow matching Midori Kato et.al. 2607.16144 null
2026-07-17 ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning Binglin Zhou et.al. 2607.16131 null
2026-07-17 Nonadiabatic excited-state dynamics with quantum Monte Carlo-trained machine learning: azomethane as a stringent test Alfonso Annarelli et.al. 2607.16129 null
2026-07-17 Quantum-classical crossover in fault-tolerant quantum dynamics simulation Jinzhao Sun et.al. 2607.16116 null
2026-07-17 Understanding Reasoning from Pretraining to Post-Training Jingyan Shen et.al. 2607.16097 null
2026-07-17 DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning Hanyang Chen et.al. 2607.16090 null
2026-07-17 JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models Haoran Sun et.al. 2607.16074 null
2026-07-17 When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis S. Aaron McClendon et.al. 2607.16062 null
2026-07-17 Deconfined quantum critical point in a dissipative spin-1/2 chain Longye Lu et.al. 2607.16039 null
2026-07-17 CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading – An Alpha-Reward Approach Andrei Neagu et.al. 2607.16028 null
2026-07-17 Infrared spectroscopy of gas-phase hydrogenated and methylated pyrenes: from laboratory spectra to the simulated 3.4 $μ$ m emission band Karine Demyk et.al. 2607.16018 null
2026-07-17 Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids Josef Hoppe et.al. 2607.16004 null
2026-07-17 BayesPO: Bayesian Prompt Optimization via Parallel-Tempered Gradient-Guided Discrete MCMC Junjie Zhou et.al. 2607.16001 null
2026-07-17 Data and Learning Where it Matters for Contact-Rich Manipulation Oliver Hausdörfer et.al. 2607.15982 null
2026-07-17 Non-thermal emission from the vicinity of the magnetar CXOU J171405.7-381031 Manoel F. Sousa et.al. 2607.15955 null
2026-07-16 MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators Yushi Huang et.al. 2607.15273 null
2026-07-16 Delocalization of bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin Yifan Chen et.al. 2607.15208 null
2026-07-16 Mask-Aware Policy Gradients for Diffusion Language Models Haran Raajesh et.al. 2607.15200 null
2026-07-16 SciPhy Reinforcement Learning for Portfolio Optimization Igor Halperin et.al. 2607.15195 null
2026-07-16 Can We Trust Item Response Theory for AI Evaluation? Han Jiang et.al. 2607.15190 null
2026-07-16 On-Policy Delta Distillation Byeongho Heo et.al. 2607.15161 null
2026-07-16 Concept-Guided Spatial Regularization for World Models in Atari Pong Yukuan Lu et.al. 2607.15142 null
2026-07-16 Perfectly equidistributed Quasi-Monte Carlo sequences from Artin-Schreier polynomials Nicolas Bonneel et.al. 2607.15141 null
2026-07-16 Learning in Infinitesimal Non-Compositional Sketches Sridhar Mahadevan et.al. 2607.15107 null
2026-07-16 Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents Dylan Van Mulders et.al. 2607.15095 null
2026-07-16 Evaluating covariate balance for long time horizon Markov decision processes Joshua Spear et.al. 2607.15080 null
2026-07-16 Learning the Fermion sign structure in path-integral Monte Carlo Jarvist Moore Frost et.al. 2607.15060 null
2026-07-16 Robust Optimal Control of Arbitrarily Switched Systems: A Path-Complete Framework Léa Ninite et.al. 2607.15055 null
2026-07-16 High resolution Lyman-α forest constraints on dark matter-neutrino scattering Markus R. Mosbech et.al. 2607.15020 null
2026-07-16 Risk-Aware Belief Control Barrier Functions over Random Finite Sets Shaohang Han et.al. 2607.15016 null
2026-07-16 CosFly-VLA: A Spatially Aware Vision-Language-Action Model for UAV Tracking Ruilong Ren et.al. 2607.15004 null
2026-07-16 SMC-ES: Automated synthesis of formally verified control policies Riccardo Curcio et.al. 2607.15003 null
2026-07-16 Achievable-Rate Analysis of MISO Systems with Transmit-Side Multiport Matching Networks Wendong Cheng et.al. 2607.14992 null
2026-07-16 Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation Ku Onoda et.al. 2607.14962 null
2026-07-16 Statistical Modelling of Planetary Boundary Layer Height and Its Measurement Uncertainty Using GRUAN Profiles Tommaso Locatelli et.al. 2607.14960 null
2026-07-15 Exact Decomposition of Adversarial Dual-Objective Value Functions, with Applications to Optimal Drug Dosing Dylan Hirsch et.al. 2607.14023 null
2026-07-15 Industrial Dexterity Benchmark: A Hardware-Software Benchmarking Platform for Industrial Dexterous Manipulation Honglu He et.al. 2607.14021 null
2026-07-15 Lighthouse RL: Sample-Efficient Circuit Optimization via Strategic Reset Points Mustafa Emre Gürsoy et.al. 2607.14008 null
2026-07-15 Lyapunov Exponent as Physics-Informed Dense Reward: RL Discovery of Stabilization Beyond the Kapitza Pendulum Slava Andrejev et.al. 2607.14001 null
2026-07-15 TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents Leitian Tao et.al. 2607.13988 null
2026-07-15 Kaleidoscopic-ray-tracing-based model of the scintillation flash energy deposition in the photomultipliers attached to a strip scintillator I. V. Vovchenko et.al. 2607.13963 null
2026-07-15 Discriminative Barrier Functions for Safe Adversarial Imitation Learning from Observation Anubhav Vishwakarma et.al. 2607.13938 null
2026-07-15 SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning Cheng Tang et.al. 2607.13931 null
2026-07-15 Hybrid Time-Frequency Domain Frequency Offset Compensation Under GHz Doppler Shift for LEO Satellite-to-Ground Coherent Free-Space Optical Communication Tiankuo Jiao et.al. 2607.13904 null
2026-07-15 Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation Yizhou Zhang et.al. 2607.13903 null
2026-07-15 Task-Oriented Sensing and Covert Transmissions for Collaborative Multi-AUV Systems Xueyao Zhang et.al. 2607.13880 null
2026-07-15 SPyCE: Skill-Policy Co-evolution for Multimodal Agents Ru Zhang et.al. 2607.13854 null
2026-07-15 Learning Robust Execution in Robotic Manipulation with Agentic Reinforcement Learning Xiaopeng Zhang et.al. 2607.13818 null
2026-07-15 Vision-Based Obstacle Separation for Strawberry Harvesting in Clusters Using Hierarchical Reinforcement Learning Teng Li et.al. 2607.13799 null
2026-07-15 Bridging Frustration and Non-Hermiticity via COMPASS: An Adaptive Biorthogonal Neural Quantum State Framework Lavoisier Wah et.al. 2607.13790 null
2026-07-15 jQMC: A JAX-based ab initio quantum Monte Carlo package designed for GPU-accelerated computing Kousuke Nakano et.al. 2607.13781 null
2026-07-15 Mono-Z Dark Matter Search with Neural Spline Flows Using CMS Run 2015D Open Data Hitesh Rasineni et.al. 2607.13771 null
2026-07-15 Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibration Shuhao Li et.al. 2607.13753 null
2026-07-15 DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention Xing Lei et.al. 2607.13731 null
2026-07-15 Stoner transitions beyond mean-field in two-dimensional electronic systems: a diagrammatic Monte Carlo study Yueh-Chen Lee et.al. 2607.13675 null
2026-07-14 DenseReward: Dense Reward Learning via Failure Synthesis for Robotic Manipulation Yu Fang et.al. 2607.13033 null
2026-07-14 TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale Zhouchonghao Wu et.al. 2607.13028 null
2026-07-14 Dynamic Resource Allocation for Ensemble Determinization MCTS Jakub Kowalski et.al. 2607.13007 null
2026-07-14 One Shot, Twenty-One Balls: Existence and Rarity of a Total Clearance in a Single Stroke of Snooker Avner Kantor et.al. 2607.12995 null
2026-07-14 A Noise-Aware Quantum Algorithm for Credit Valuation Adjustments on Real Quantum Hardware Guillem Borràs Espert et.al. 2607.12990 null
2026-07-14 RecRec: Latent Interests Recursive Reasoning for Sequential Recommendation Wenhao Deng et.al. 2607.12945 null
2026-07-14 ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning Yilun Kong et.al. 2607.12931 null
2026-07-14 Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes Jonas Ehrhardt et.al. 2607.12924 null
2026-07-14 LatentFlow: A General Framework for Conditioning Stochastic Processes Louis Sharrock et.al. 2607.12922 null
2026-07-14 Accelerated Mixing Time of Randomized Hamiltonian Monte Carlo Siddharth Mitra et.al. 2607.12902 null
2026-07-14 Unveiling Complex Collective Behaviors from Simple Rewards Yize Mi et.al. 2607.12861 null
2026-07-14 Verifier-Based Reinforcement Fine-Tuning of Reasoning Models for Thermal Energy Storage Control Takumi Shioda et.al. 2607.12856 null
2026-07-14 Hash-augmented adaptive multilevel splitting Monte Carlo algorithm for accurate estimation of two-sample permutation test p-values Nikita Golikov et.al. 2607.12853 null
2026-07-14 AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning Yanghai Wang et.al. 2607.12820 null
2026-07-14 UniVR: Thinking in Visual Space for Unified Visual Reasoning Zhongwei Ren et.al. 2607.12800 null
2026-07-14 Directional Constraints for Efficient Exploration in Safe Reinforcement Learning Paolo Magliano et.al. 2607.12784 null
2026-07-14 Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination Usman Haider et.al. 2607.12763 null
2026-07-14 Standard basis operator method for ground-state and temperature properties of single and two-component Bose-Hubbard model Oliwier Urbański et.al. 2607.12718 null
2026-07-14 GRAFT: Graph-Matched Retrieval and Fusion of Tables in Data Lakes Daomin Ji et.al. 2607.12717 null
2026-07-14 Vision-Based Dribbling for Humanoid Soccer via Privileged Representation Learning Flavio Maiorana et.al. 2607.12702 null
2026-07-14 From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation Mehak Dhaliwal et.al. 2607.12687 null
2026-07-14 A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism Chengguang Gan et.al. 2607.12640 null
2026-07-14 Gradient-free learning of a closed-loop wall controller for turbulent drag reduction Giorgio Maria Cavallazzi et.al. 2607.12626 null
2026-07-14 Direct Retrieval of Protoplanetary Disk Dust Properties using Auto-differentiable Gaussian Processes and Its Application to the HD 169142 Disk Tomohiro C. Yoshida et.al. 2607.12618 null
2026-07-14 Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning Amber Srivastava et.al. 2607.12590 null
2026-07-14 Lattice Configuration Generation with a Self-Learning Diffusion Model Akio Tomiya et.al. 2607.12587 null
2026-07-14 OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning Emil Mittag et.al. 2607.12523 null
2026-07-14 TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments Edward Y. Chang et.al. 2607.12480 null
2026-07-14 Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences Taehyung Kim et.al. 2607.12466 null
2026-07-13 Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Runhui Huang et.al. 2607.11886 null
2026-07-13 A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation Yunhai Feng et.al. 2607.11874 null
2026-07-13 Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search Romain Amigon et.al. 2607.11826 null
2026-07-13 Active Noise Floor Estimation for Reliability-Optimal POMDPs: A Value-of-Noise-Information Approach Hyung-Jin Yoon et.al. 2607.11822 null
2026-07-13 Analytical Markov Chain for Spatiotemporal Flux Evolution of the Inner Filter Effect in Fluorescent Media Xuhui Yang et.al. 2607.11815 null
2026-07-13 Synchronized Three-Dimensional Vocal-Tract Motion for Speech Synchronization via Joint-Embedding Predictive Architecture Alignment Sheng Li et.al. 2607.11772 null
2026-07-13 Multi-Agent Reinforcement Learning for C-V2X RAT Selection Moritz Schaffenroth et.al. 2607.11744 null
2026-07-13 Exchange topology and criticality in ferrite and chromium spinels: a unified Monte Carlo analysis Keltoum Khallouq et.al. 2607.11729 null
2026-07-13 Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories Ziheng Zhang et.al. 2607.11725 null
2026-07-13 Active Offline-to-Online Reinforcement Learning Alper Kamil Bozkurt et.al. 2607.11720 null
2026-07-13 One Vote, Several Parliaments: An Empirical Analysis of the Algorithmic Ambiguity of the Italian Electoral Law on the 2022 General Election Data Paolo Coppola et.al. 2607.11676 null
2026-07-13 Thermal phase transitions in a mixed-spin Ising model on the Lieb lattice: Exact results beyond zero magnetic field Jozef Strecka et.al. 2607.11661 null
2026-07-13 Markov Chain Monte Carlo with Diffusion Paths Han Chen et.al. 2607.11631 null
2026-07-13 On the Policy Convergence of Policy Mirror Descent Methods Wenye Li et.al. 2607.11626 null
2026-07-13 SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning Evelyn D’Elia et.al. 2607.11624 null
2026-07-13 Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns Yong Yang et.al. 2607.11621 null
2026-07-13 Auditing the Risk Claims of Distributional Reinforcement Learning Hari Prasad et.al. 2607.11607 null
2026-07-13 Copositive Characterizations of Convex Hull Pricing Madhusudan Ghosh et.al. 2607.11590 null
2026-07-13 Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Xin Zhang et.al. 2607.11581 null
2026-07-13 DiffEEG: A Self-Supervised Denoising Diffusion Model for Learning EEG Generic Representations Abdulkader Helwan et.al. 2607.11578 null
2026-07-10 Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection Cláudio Lúcio do Val Lopes et.al. 2607.09641 null
2026-07-10 Beyond the Cube: Overlapping Grid Methods for Debris Collision Risk Assessment Yacob Medhin et.al. 2607.09634 null
2026-07-10 PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers Yujie Pang et.al. 2607.09590 null
2026-07-10 CORAL-AUV: CFD Oriented Reinforcement Learning for Autonomous Underwater Vehicles Steven Roche et.al. 2607.09557 null
2026-07-10 A Boosted Energy Extraction from the CapMix Process by Grafting with Titratable Polymers Mamta Yadav et.al. 2607.09554 null
2026-07-10 Fused Constrained Policy Reuse Optimization for Wireless Resource Allocation An Liu et.al. 2607.09498 null
2026-07-10 Multimodal Reward Hacking in Reinforcement Learning Jiayu Yao et.al. 2607.09492 null
2026-07-10 Symmetry-Protected Pinch Curves in Classical Spin Liquids Takumi Fukushima et.al. 2607.09470 null
2026-07-10 Deep Learning for Dynamic Programming with Recursive Utility Using First-order Conditions Xianhua Peng et.al. 2607.09461 null
2026-07-10 Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning Edwin De Nicolo et.al. 2607.09422 null
2026-07-10 Tests for Increasing Convex Ordering Based on Generalized Tsallis Entropy Measures Aritra Saha et.al. 2607.09418 null
2026-07-10 Two observables of one wall: how surface relaxivity can bias the diffusion intra-axonal fraction and the myelin water fraction Rutger H. J. Fick et.al. 2607.09401 null
2026-07-10 Mach-Mind-4-Flash Technical Report Foundation Model Team et.al. 2607.09375 null
2026-07-10 Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning Guanquan Wang et.al. 2607.09336 null
2026-07-10 Clock-noise subtraction in geometric time-delay interferometry for space-based gravitational-wave parameter estimation Rui Luo et.al. 2607.09335 null
2026-07-10 Risk-Aware General-Utility Markov Decision Processes Pedro P. Santos et.al. 2607.09298 null
2026-07-10 Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC Mohammad Farhoudi et.al. 2607.09295 null
2026-07-10 Diffusion Monte Carlo study of deuteron-like fully light hexaquarks M. C. Gordillo et.al. 2607.09288 null
2026-07-10 An Improved Deep Reinforcement Learning Control Strategy for Traction Dual Rectifiers in EMUs Zhigang Liu et.al. 2607.09276 null
2026-07-10 When Does Order Flow Matter? State-Dependent L2 Liquidity-State Transitions in Crypto Futures Joohyoung Jeon et.al. 2607.09230 null
2026-07-09 Force convergence in Monte Carlo Lyman-alpha radiative transfer Joshua Kasiri et.al. 2607.08726 null
2026-07-09 Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference Chuning Zhu et.al. 2607.08724 null
2026-07-09 Optimal-Transport-Based Cell Resampling for Negative and Pathological Event Weights Regan Doherty et.al. 2607.08723 null
2026-07-09 On improving the estimates of the sampling variances via Global-Local priors in Small Area Estimation Sirapat Watakajaturaphon et.al. 2607.08720 null
2026-07-09 MPFlow: Learning Budgeted Max-Flow Optimization on the Lightning Network with Deep Graph Reinforcement Learning Harrison Rush et.al. 2607.08703 null
2026-07-09 Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution Ethan Leung et.al. 2607.08700 null
2026-07-09 Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning Ali Larian et.al. 2607.08647 null
2026-07-09 A Novel Hadronic Calorimeter With A Direct Neutron Readout I. Giomataris et.al. 2607.08587 null
2026-07-09 Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Yiyang Fang et.al. 2607.08572 null
2026-07-09 PhononScore: a phonon-aware scoring function for dynamical stability Xiao-Qi Han et.al. 2607.08518 null
2026-07-09 Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Feng Wang et.al. 2607.08497 null
2026-07-09 Learning LDPC codes with quantized density evolution over relaxed protographs Gennady Shutkov et.al. 2607.08484 null
2026-07-09 Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning Zijie Cheng et.al. 2607.08444 null
2026-07-09 ADORN: Adaptive Drift handling for Open RAN using Reinforcement Learning Ashit Kumar Subudhi et.al. 2607.08443 null
2026-07-09 A note on the convergence of the eigenvalues in a subdomain to the continuous spectrum Miroslav Bulíček et.al. 2607.08433 null
2026-07-09 When Synthetic Speech Is All You Have: Better Call GRPO Shashi Kumar et.al. 2607.08409 null
2026-07-09 Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks Hong Zhao et.al. 2607.08406 null
2026-07-09 DrugGen 2: A disease-aware language model for enhancing drug discovery Ali Motahharynia et.al. 2607.08404 null
2026-07-09 Testing Covariance Separability in High Dimensions Tomas Masak et.al. 2607.08388 null
2026-07-09 Self-Adaptive Anomaly Detection with Reinforcement Learning and Human Feedback in Connected Vehicles Matthias Weiß et.al. 2607.08373 null
2026-07-08 Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF Eric Zhu et.al. 2607.07693 null
2026-07-08 Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning Vladislav Beliaev et.al. 2607.07690 null
2026-07-08 Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops Mingguang Chen et.al. 2607.07663 null
2026-07-08 Unlearning to Protect: A Distilled Reinforcement Learning Framework with Privacy-Preserving Feature Unlearning and XAI for IoT Security Md. Nahid Hasan et.al. 2607.07635 null
2026-07-08 PHaul: A PPO-based forwarding agent for Sub6 enhanced Integrated Access and Backhaul networks Jorge Pueyo et.al. 2607.07584 null
2026-07-08 On the Robustness in Data-Driven Nonlinear Optimal Control: From Stability to Optimality Yicheng Lin et.al. 2607.07570 null
2026-07-08 RubriQ: Rubric-Guided Group Relative Policy Optimization for Constraint-Aware Quantum Circuit Synthesis Ziqing Guo et.al. 2607.07554 null
2026-07-08 Gradient-free Riemannian Langevin Sampler Ricardo Baptista et.al. 2607.07519 null
2026-07-08 Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning Zhenyu Hou et.al. 2607.07508 null
2026-07-08 Reward-Adaptive Iterative Discovery: A Case Study on Automated Game Testing for NHL26 Florian Fuchs et.al. 2607.07498 null
2026-07-08 5G Positioning Reference Signal impact assessment in Non-Terrestrial Networks communication service Alejandro Gonzalez-Garrido et.al. 2607.07466 null
2026-07-08 EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI Xinjie Wang et.al. 2607.07459 null
2026-07-08 A Multi-Scale Machine Learning Framework for Coupled Chemical, Spin, and Structural Disorder in Alloys Zhenyao Fang et.al. 2607.07456 null
2026-07-08 Collaborate to decorrelate in path space: Hamiltonian replica exchange transition interface sampling (HRETIS) Sina Safaei et.al. 2607.07453 null
2026-07-08 RLVP: Penalize the Path, Reward the Outcome Bojie Li et.al. 2607.07435 null
2026-07-08 Immersive Social Interaction with VR and LLM-Assisted Humanoids Niraj Pudasaini et.al. 2607.07430 null
2026-07-08 Improving greenhouse fruit-production control by integrating reinforcement learning into short-horizon model predictive control Bart van Laatum et.al. 2607.07365 null
2026-07-08 BUS: Brain-Inspired Unsupervised Self-Reflection for Advanced Multimodal Reasoning Jiacheng Yang et.al. 2607.07361 null
2026-07-08 The Joneses Visit an Economics Lab Mikhail Freer et.al. 2607.07353 null
2026-07-08 Taming nonlinear energy diffusion: The case of time-crystal energy condensates P. I. Hurtado et.al. 2607.07325 null
2026-07-07 Embodied Human-Robot Interaction via Acoustics: A MARL Approach with AcoustoBots for Spatial Data Physicalization Shiqi Liu et.al. 2607.06563 null
2026-07-07 RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation Haoyu Zhao et.al. 2607.06558 null
2026-07-07 Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment Han-Jun Ko et.al. 2607.06522 null
2026-07-07 FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games Chase McDonald et.al. 2607.06514 null
2026-07-07 Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine Juan S. Santillana et.al. 2607.06495 null
2026-07-07 Multi-Agent Deep Reinforcement Learning for Multi Objective Battery Management in Dairy Farms Marcos Eduardo Cruz Victorio et.al. 2607.06489 null
2026-07-07 SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models Changti Wu et.al. 2607.06442 null
2026-07-07 A Machine-Learning-Compatible Omnibus Test for Treatment Effect Heterogeneity Elia Lapenta et.al. 2607.06412 null
2026-07-07 A Definition and Roadmap for World Models Xinyuan Chen et.al. 2607.06401 null
2026-07-07 Learning to Throw Objects Safely in Multi-Obstacle Environments Mohammadreza Kasaei et.al. 2607.06388 null
2026-07-07 LAMP: Latent Motion Prior-Guided Real-World Learning for Dexterous Hand Manipulation Xinye Yang et.al. 2607.06323 null
2026-07-07 Joint Probabilistic and Geometric Constellation Shaping for Complexity-Constrained Direct Detection Optical Systems Rodrigo Fischer et.al. 2607.06300 null
2026-07-07 A Convex Approximation Framework for Neural Likelihood-Based Bayesian Inverse Problems Fabian Schneider et.al. 2607.06252 null
2026-07-07 Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale Ziting Wang et.al. 2607.06233 null
2026-07-07 Entanglement as a Structural Complexity Axis: A PAC-Bayesian View of Generalization in Quantum Policies and Value Functions Jian Xu et.al. 2607.06230 null
2026-07-07 Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents Yijun Zhang et.al. 2607.06223 null
2026-07-07 Arbitrage-Free Multi-Maturity Risk-Neutral Marginals Hao Qin et.al. 2607.06204 null
2026-07-07 Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design Alexander Rombach et.al. 2607.06175 null
2026-07-07 CurateEvo: Data-Curation Evolving for Agentic Post-Training Dingzirui Wang et.al. 2607.06140 null
2026-07-07 Self-Bound Droplets of Ultracold Dipolar Molecules under Tunable Double Microwave Shielding Roger Melero et.al. 2607.06130 null
2026-07-06 Weak-to-Strong Generalization via Direct On-Policy Distillation Shiyuan Feng et.al. 2607.05394 null
2026-07-06 CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents Yujiang Li et.al. 2607.05378 null
2026-07-06 Fitted Occupancy-Ratio Evaluation without Bellman Completeness Lars van der Laan et.al. 2607.05375 null
2026-07-06 Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP Planning Idan Lev-Yehudi et.al. 2607.05359 null
2026-07-06 Partial Gateaux and Frechet Derivatives and Applications to Variational Analysis Jinlu Li et.al. 2607.05313 null
2026-07-06 Adaptive Inference Batching using Policy Gradients Ruslan Sharifullin et.al. 2607.05272 null
2026-07-06 Optimal Base Station Placement for Beyond 5G Networks with Non-Convex Topology Mohamed Shalma et.al. 2607.05210 null
2026-07-06 Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models Raj Jaiswal et.al. 2607.05199 null
2026-07-06 When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents Yechao Zhang et.al. 2607.05189 null
2026-07-06 SMART: A Machine Learning and Monte Carlo Framework for Rapid Analysis of Stochastic Transistor Aging and Process Variation in Digital Circuits Arash Esshaghi et.al. 2607.05187 null
2026-07-06 Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets Enrique Adrian Villarrubia-Martin et.al. 2607.05179 null
2026-07-06 Efficient classical simulation of two-dimensional long-range systems: Rydberg arrays and beyond Jia-Lin Chan et.al. 2607.05178 null
2026-07-06 Variance reduction with probing and Multilevel Monte Carlo in Lattice QCD Andreas Frommer et.al. 2607.05157 null
2026-07-06 Claim-Level Rubric Rewards for Video Caption Reinforcement Learning Mingqi Gao et.al. 2607.05150 null
2026-07-06 TimeThink: Reasoning with Time for Video LLMs Handong Li et.al. 2607.05089 null
2026-07-06 Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization Junqi Tu et.al. 2607.05064 null
2026-07-06 Watts per event: evaluating Sustainability of HEP Event Generators beyond the LHC era Szabolcs Molnár et.al. 2607.05018 null
2026-07-06 Non-Convex Sparse Reinforcement Learning via Non-Monotone Inclusions Kyohei Suzuki et.al. 2607.04990 null
2026-07-06 STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training Qiuyi Qi et.al. 2607.04963 null
2026-07-06 Closing the Reality Gap: Zero-Shot Sim-to-Real Deployment for Dexterous Force-Based Grasping and Manipulation Zhe Zhao et.al. 2607.04940 null
2026-07-02 Seek to Segment: Active Perception for Panoramic Referring Segmentation Song Tang et.al. 2607.02497 null
2026-07-02 Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning Liyan Tang et.al. 2607.02490 null
2026-07-02 Learning Agile Intruder Interception using Differentiable Quadrotor Dynamics Michael Anoruo et.al. 2607.02472 null
2026-07-02 Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation Zhuowei Chen et.al. 2607.02460 null
2026-07-02 WorldSample: Closed-loop Real-robot RL with World Modelling Yuquan Xue et.al. 2607.02431 null
2026-07-02 DecompRL: Solving Harder Problems by Learning Modular Code Generation Juliette Decugis et.al. 2607.02390 null
2026-07-02 Mesoscopic Linear Statistics for Two Ensembles of Quantum Graphs Anna Maltsev et.al. 2607.02356 null
2026-07-02 SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces Jinwei Hu et.al. 2607.02345 null
2026-07-02 Sensitivity Analysis and Robust Optimal Control for Coupled Evolution Inclusions with State-Dependent Maximal Monotone Operators Jinsheng Du et.al. 2607.02339 null
2026-07-02 Hybridizing a Grouping Metaheuristic with Reinforcement Learning for the One-Dimensional Bin Packing Problem Zitouni Rania et.al. 2607.02315 null
2026-07-02 One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective Juan Agustín Duque et.al. 2607.02292 null
2026-07-02 Optimizing Visual Generative Models via Distribution-wise Rewards Ruihang Li et.al. 2607.02291 null
2026-07-02 Generalization in offline RL: The structure is more important than the amount of pessimism Max Weltevrede et.al. 2607.02288 null
2026-07-02 DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation Zijun Li et.al. 2607.02220 null
2026-07-02 Differentiable inverse design of short-range order in high-entropy alloys: from target sro to target property Tiancheng Ding et.al. 2607.02219 null
2026-07-02 The Binary Crisis Clock: Controlled by Sparse Ternary Interventions Małgorzata Nowak-Kȩpczyk et.al. 2607.02207 null
2026-07-02 Actuator Reality Shaping for Zero-Shot Sim-to-Real Robot Learning Satoshi Yamamori et.al. 2607.02205 null
2026-07-02 ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning Yilie Huang et.al. 2607.02137 null
2026-07-02 Reachability-Based Safe-Start Regions for Approach to a Tumbling Target with Rotating LOS Constraints Omer Burak Iskender et.al. 2607.02128 null
2026-07-02 Coverage Analysis in Terahertz Clustered HetNets Hadeel Obaid et.al. 2607.02125 null
2026-07-01 Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training Zijian Zhang et.al. 2607.01232 null
2026-07-01 Language-Critique Imitation Learning from Suboptimal Demonstrations Chih-Han Yang et.al. 2607.01225 null
2026-07-01 Computationally Efficient Near-Optimal Control for Current Ripple Reduction and Optimization of Three-Phase Motors via LMIs Huu-Thinh Do et.al. 2607.01215 null
2026-07-01 Quantum vs. Classical Machine Learning: A Unified Empirical Comparison Chuanming Yu et.al. 2607.01197 null
2026-07-01 Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Hongxing Li et.al. 2607.01191 null
2026-07-01 QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling Michael Y. Li et.al. 2607.01179 null
2026-07-01 Diffusion-GR2: Diffusion Generative Reasoning Re-ranker Zhuoxuan Zhang et.al. 2607.01170 null
2026-07-01 Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents Ran Yan et.al. 2607.01120 null
2026-07-01 Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use Song-Lin Lv et.al. 2607.01084 null
2026-07-01 AutoRestTest at the SBFT 2026 Tool Competition Tyler Stennett et.al. 2607.01063 null
2026-07-01 AMBUSH: Collaborative Capture in Complex Environments with Neural Acceleration Junfeng Chen et.al. 2607.01029 null
2026-07-01 Effect of radially heterogeneous band gap collapse on formation of swift heavy ion tracks in Al2O3 Roman Voronkov et.al. 2607.01016 null
2026-07-01 The Milky Way Atlas for Linear Filaments III: Giant filaments and magnetic fields as evidence of a bubbly Galactic disk Naval K. Bhadari et.al. 2607.00976 null
2026-07-01 DRL-Based Joint Beamforming and Surface Shape Optimization for Flexible Intelligent Metasurface-Aided ISAC Systems Maoyuan Wang et.al. 2607.00951 null
2026-07-01 Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm Midhun Parakkal Unni et.al. 2607.00926 null
2026-07-01 Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination Subhadeep Pal et.al. 2607.00924 null
2026-07-01 Tail Risk Management with Puts and Trend Following: A CVaR Framework for Crashes and Drawdowns Miquel Noguer I Alonso et.al. 2607.00883 null
2026-07-01 EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection Wenhao Zhang et.al. 2607.00867 null
2026-07-01 From Pixels to Temporal Correlations: Learning Informative Representations for Reinforcement Learning Pre-training Jinwen Wang et.al. 2607.00811 null
2026-07-01 Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos Jinwen Wang et.al. 2607.00808 null
2026-06-30 Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs Gabrielle Kaili-May Liu et.al. 2606.32032 null
2026-06-30 Freeform Preference Learning for Robotic Manipulation Marcel Torne et.al. 2606.32027 null
2026-06-30 TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning Yuanda Xu et.al. 2606.32017 null
2026-06-30 On the Comparison of Reinforcement Learning and Adaptive Control for Linear Systems under Packet Loss and Uncertainty Moh Kamalul Wafi et.al. 2606.32003 null
2026-06-30 OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation Arnav Balaji et.al. 2606.31993 null
2026-06-30 GR2 Technical Report Yufei Li et.al. 2606.31984 null
2026-06-30 Adapting Generalist Robot Policies with Semantic Reinforcement Learning Jagdeep Singh Bhatia et.al. 2606.31958 null
2026-06-30 LeCropFollow: Latent Space Planning for Navigation in Unstructured Crop Fields Felipe Tommaselli et.al. 2606.31941 null
2026-06-30 Learning Locomotion on Discrete Terrain via Minimal Proximity Sensing Jiale Fan et.al. 2606.31912 null
2026-06-30 CoDex: Learning Compositional Dexterous Functional Manipulation without Demonstrations Bowen Jiang et.al. 2606.31909 null
2026-06-30 Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models Lang Cao et.al. 2606.31846 null
2026-06-30 RoboTacDex: A Dexterous Visual-Tactile-Action Dataset for Humanoid Manipulation Xinyi Wang et.al. 2606.31836 null
2026-06-30 Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning Junha Jung et.al. 2606.31825 null
2026-06-30 Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR Ruijia Zhang et.al. 2606.31813 null
2026-06-30 Reinforcement Learning-Based Control for an Inline Skating Humanoid Robot Ethan Marot et.al. 2606.31807 null
2026-06-30 Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision Xianda Zheng et.al. 2606.31800 null
2026-06-30 Information-Epidemic Dynamics in Cyber-Physical Systems: A Hypergraph Framework with Interpersonal Relationships Shanchao Peng et.al. 2606.31782 null
2026-06-30 Addressing Over-Refusal in LLMs with Competing Rewards Taeyoun Kim et.al. 2606.31748 null
2026-06-30 Dynamic Scheduling for Flexible Manufacturing Systems Based on Multi-Agent Deep Reinforcement Learning and Petri Nets Zhou He et.al. 2606.31737 null
2026-06-30 UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization Yaozhi Zheng et.al. 2606.31732 null
2026-06-29 Pessimism’s Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models Subramanyam Sahoo et.al. 2606.30627 null
2026-06-29 Prescriptions for the stochasticity effect on the integrated X-ray luminosity of star-forming galaxies:Implications for selecting star-forming galaxies and AGN in X-ray surveys Elias Kyritsis et.al. 2606.30624 null
2026-06-29 When and Which Sensor to Observe? Timely Tracking of a Joint Markov Source Ismail Cosandal et.al. 2606.30623 null
2026-06-29 Equilibrium and non-equilibrium phases of microwave-dressed polar molecules beyond rotational symmetries Matteo Ciardi et.al. 2606.30589 null
2026-06-29 Staged Hybridisation for Visual Quantum Reinforcement Learning via Knowledge Distillation Javier Lazaro et.al. 2606.30520 null
2026-06-29 Bayesian Analysis with Markov Chain Monte Carlo for Global Optimization and Degeneracy Diagnosis in Nuclear Mass Models Xiangnan Lee et.al. 2606.30519 null
2026-06-29 Discovering the Kalman-Bucy-Koopman Filter Umesh Vaidya et.al. 2606.30487 null
2026-06-29 Grasp-Oriented Non-Prehensile Manipulation via Learning a Graspability Field Licheng Zhong et.al. 2606.30474 null
2026-06-29 When Does Online Imitation Learning Help in LLM Post-Training? The Role of (Non-)Realizability Beyond Horizon Huaqing Zhang et.al. 2606.30445 null
2026-06-29 Experience Augmented Policy Optimization for LLM Reasoning Jinda Lu et.al. 2606.30420 null
2026-06-29 Quantum-enhanced Monte Carlo Tree Search framework for combinatorial optimization problems Yohan Finet et.al. 2606.30415 null
2026-06-29 Diffusion Fine-tuning with Rewarded Moment Matching Distillation Alexis Jacq et.al. 2606.30414 null
2026-06-29 MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training Wenhan Ma et.al. 2606.30406 null
2026-06-29 FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification Zheming Fu et.al. 2606.30376 null
2026-06-29 DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training Haisen Luo et.al. 2606.30345 null
2026-06-29 Value Functions of Separable Convex Integer Programs are Periodically Convex Koen Ligthart et.al. 2606.30330 null
2026-06-29 Toward an Energy-Optimized Operation of Data Centers Located in Wind Farms Using Reinforcement Learning Jan Stenner et.al. 2606.30316 null
2026-06-29 Highly Data Parallelizable Estimation of the Sliced-Wasserstein Distance Using Cumulative Distribution Functions Christophe Vauthier et.al. 2606.30310 null
2026-06-29 Pathway variability, coat stiffening and mechanical adaptation during clathrin-mediated endocytosis Johannes H. H. Dreckhoff et.al. 2606.30267 null
2026-06-29 Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation Shihao Zhang et.al. 2606.30248 null
2026-06-29 KYON: Semi-Modular Wheel-Legged Quadruped With Agile Bimanual Capability Luca Rossini et.al. 2606.30243 null
2026-06-29 Sparse Sensor Placement in Multi-Agent Reinforcement Learning Control of Rayleigh-Bénard Convection Jan Stenner et.al. 2606.30238 null
2026-06-29 EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures Buğra Alperen Uluırmak et.al. 2606.30219 null
2026-06-29 Precision measurement of radiative neutron \b{eta}-decay: methodology and systematic effects J. S. Nico et.al. 2606.30205 null
2026-06-29 Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data Hyunwoo Park et.al. 2606.30192 null
2026-06-29 Reactive Graphs for Efficient Markov Chain Monte Carlo Inference in Probabilistic Programming Languages Viktor Palmkvist et.al. 2606.30137 null
2026-06-29 Kinetic energy from the cubic sum rule of the dynamic structure factor Fotios Kalkavouras et.al. 2606.30123 null
2026-06-29 Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts Chunhui Bai et.al. 2606.30092 null
2026-06-29 ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning Daiki E. Matsunaga et.al. 2606.30072 null
2026-06-29 Phase Boundary of a Stochastic Watts-Threshold SIS Model on Random Networks Yasmine Beji et.al. 2606.30069 null
2026-06-29 Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management Byeong Hoon Yoon et.al. 2606.30067 null
2026-06-29 Joint Outage Detection and Compensation for Self-Healing 5G RAN via Deep Reinforcement Learning Sajjad Hussain et.al. 2606.30031 null
2026-06-29 Error bounds for simultaneous Wasserstein contractive adaptive increasingly rare MCMC Julian Hofstadler et.al. 2606.30018 null
2026-06-29 Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting Xiaobiao Du et.al. 2606.30017 null
2026-06-29 AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills Xinyuan Song et.al. 2606.29999 null
2026-06-29 Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning Peng et.al. 2606.29984 null
2026-06-26 WARP-RM: A Warp-Augmented Relative Progress Reward Model for Data Curation Justin Yu et.al. 2606.28320 null
2026-06-26 Constraining primordial oscillations and inflationary particle production with Planck, ACT DR6, and DESI DR2 Simran K. Nerval et.al. 2606.28310 null
2026-06-26 Accretion-Driven Evolution of Compact-Object Populations in Gas-Rich Environments and the Origin of Massive Gravitational-Wave Sources Mor Rozner et.al. 2606.28293 null
2026-06-26 Composing Quantum Instruments Robert I. Booth et.al. 2606.28291 null
2026-06-26 Numerical model of fast electron energy deposition in interstellar molecular gas Aleksandr Nesterenok et.al. 2606.28259 null
2026-06-26 HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech Sihang Nie et.al. 2606.28249 null
2026-06-26 Learning Stable In-Grasp Manipulation in a Non-Dropping Action Space Ha Thang Long Doan et.al. 2606.28196 null
2026-06-26 Tandem Reinforcement Learning with Verifiable Rewards Difan Jiao et.al. 2606.28166 null
2026-06-26 EchoSonar-R: A Multi-View Reasoning-Enabled Model for Disease Classification and Report Generation in Echocardiography Darya Taratynova et.al. 2606.28164 null
2026-06-26 Regularized Reward-Punishment Reinforcement Learning Jiexin Wang et.al. 2606.28152 null
2026-06-26 Configurational Temperature in Matrix Models and Random Matrix Ensembles Anosh Joseph et.al. 2606.28148 null
2026-06-26 A statistically robust framework for detecting and classifying hysteresis patterns in astrophysical spectral evolution Tomislav Terzić et.al. 2606.28146 null
2026-06-26 The QCD energy-momentum tensor on the lattice: non-perturbative renormalization with $N_f=3$ Matteo Bresciani et.al. 2606.28035 null
2026-06-26 TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL Jing Wang et.al. 2606.28016 null
2026-06-26 Latent Visual Diffusion Reasoning with Monte Carlo Tree Search Xirui Teng et.al. 2606.27988 null
2026-06-26 Local Fokker–Planck Geometry for Score Estimation: Heat-Ball Mean-Value Representations and Exact High-Dimensional Sampling Jiayao Bai et.al. 2606.27954 null
2026-06-26 Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition Violeta Basten-Romero et.al. 2606.27939 null
2026-06-26 (In)Efficient Market States and Rough Volatility Detected via Grunwald-Letnikov Fractional Derivative Daniele Angelini et.al. 2606.27932 null
2026-06-26 Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing Can Li et.al. 2606.27926 null
2026-06-26 Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding Shuimu Chen et.al. 2606.27922 null
2026-06-25 Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards Ritesh Thawkar et.al. 2606.27376 null
2026-06-25 World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays Manish Kumar Govind et.al. 2606.27374 null
2026-06-25 Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models Shravan Venkatraman et.al. 2606.27373 null
2026-06-25 Don’t Settle at the Mode! Mitigating Diversity Collapse in Pretrained Flow Models via Feature Self-Guidance Pradhaan S Bhat et.al. 2606.27371 null
2026-06-25 Reinforcement Learning without Ground-Truth Solutions can Improve LLMs Yingyu Lin et.al. 2606.27369 null
2026-06-25 Bridging Performance and Generalization in Reinforcement Learning for Agile Flight Jonathan Green et.al. 2606.27348 null
2026-06-25 VibeAct: Vibration to Actions for Contact-Rich Reactive Robot Dexterity Yuemin Mao et.al. 2606.27344 null
2026-06-25 Sculpting NeRF Geometry: Human-Preference Fine-Tuning of a 3D-Aware Face GAN Archer Moore et.al. 2606.27305 null
2026-06-25 Designing Reward Signals for Portable Query Generation: A Case Study in Industrial Semantic Job Search Ping Liu et.al. 2606.27291 null
2026-06-25 Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC Alina Bazarova et.al. 2606.27286 null
2026-06-25 Quasi-Feynman formulas that provide fast converging Chernoff approximations to solution of parabolic differential equation on the real line Ivan D. Remizov et.al. 2606.27232 null
2026-06-25 Paved with True Intents: Intent-Aware Training Improves LLM Safety Classification Across Training Regimes Jeremias Ferrao et.al. 2606.27210 null
2026-06-25 Finite temperature precursors of Mottness in the Fermi Hubbard model Sayantan Roy et.al. 2606.27204 null
2026-06-25 Towards a Theory of Dobrakov-Sobolev Spaces Artem Yurievich Dudko et.al. 2606.27194 null
2026-06-25 Numerical Approximation for Path-Dependent McKean-Vlasov Control with Non-Asymptotic Error Estimates Olivier Bokanowski et.al. 2606.27181 null
2026-06-25 Automating Potential-based Reward Shaping with Vision Language Model Guidance Henrik Müller et.al. 2606.27180 null
2026-06-25 On Fourier Phase Retrieval from Differential Intensity Measurements with Applications to Wavefront Sensing Simon Hubmer et.al. 2606.27176 null
2026-06-25 Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline) Ilia Larchenko et.al. 2606.27163 null
2026-06-25 fTNN: a tensor neural network for fractional PDEs Qingkui Ma et.al. 2606.27140 null
2026-06-25 Heavy-Ball Q-Learning with Residual Weighting Correction Donghwan Lee et.al. 2606.27112 null
2026-06-24 On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Andrei Liviu Nicolicioiu et.al. 2606.26091 null
2026-06-24 On the entropic convergence for piecewise deterministic samplers: speedup and obstruction Pierre Monmarché et.al. 2606.26086 null
2026-06-24 Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents Changdae Oh et.al. 2606.26080 null
2026-06-24 Deep Reinforcement Learning-Enhanced Event-Triggered Data-Driven Predictive Control for a 3D Cable-Driven Soft Robotic Arm Cheng Ouyang et.al. 2606.26048 null
2026-06-24 Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations Han Bao et.al. 2606.26047 null
2026-06-24 Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It Yupu Hao et.al. 2606.26027 null
2026-06-24 G2DP: Diffusion Planning with Spatio-Temporal Grid Guidance Hang Yu et.al. 2606.26017 null
2026-06-24 Is Variational Monte Carlo Robust? Sharp Moment Thresholds and Heavy-tailed Stochastic Optimization Philipp Grohs et.al. 2606.26009 null
2026-06-24 FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation Shuyi Zhang et.al. 2606.26006 null
2026-06-24 Hierarchical Reinforcement Learning for Neural Network Compression (HiReLC): Pruning and Quantization Kamar Hibatallah Baghdadi et.al. 2606.26002 null
2026-06-24 Exploring Pareto smoothing in sequential Monte Carlo Jia Le Tan et.al. 2606.25983 null
2026-06-24 Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and Factorized Branch-and-Bound Thiago Thomas et.al. 2606.25978 null
2026-06-24 Studentized Cheap Bootstrap: Achieving Higher-Order Coverage Accuracy with Low Computation Shengyi He et.al. 2606.25968 null
2026-06-24 Slice Monte Carlo Integration Johannes K. Krondorfer et.al. 2606.25967 null
2026-06-24 Mixture-of-Experts RL for Fault-Tolerant Legged Locomotion Giulio Turrisi et.al. 2606.25965 null
2026-06-24 WinDOM: Self-Family Distillation for Small-Model GUI Grounding Chengheng Li-Chen et.al. 2606.25964 null
2026-06-24 Complementary probes of Bilinear RPV SUSY models with a wino-like LSP via Neutrino Oscillation and LHC Arghya Choudhury et.al. 2606.25936 null
2026-06-24 Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data Shangkun Li et.al. 2606.25894 null
2026-06-24 A Physics-Informed Statistical Learning Model for Long-Term Fragmentation Cloud Propagation Yema Paul et.al. 2606.25893 null
2026-06-24 Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents Peng Xu et.al. 2606.25852 null
2026-06-23 Sequential Probability Ratio Test using Z-Statistics (SPRT-z): A Practical Approach for Online Experimentation Derek L. Ho et.al. 2606.24871 null
2026-06-23 Exact log-odds representation and mean-field criticality of a growing social group model Xingfu Ke et.al. 2606.24818 null
2026-06-23 Finite Spectral-Band Optimal Control of Acoustic Waves via Subwavelength Point-Like Resonant Actuators Arpan Mukherjee et.al. 2606.24788 null
2026-06-23 World Value Models for Robotic Manipulation Zhihao Wang et.al. 2606.24742 null
2026-06-23 Reweighting Underlying Event and Colour Reconnection parameter variations in Sherpa Moritz Pabst et.al. 2606.24702 null
2026-06-23 LaGO: Latent Action Guidance for Online Reinforcement Learning Kuan-Yen Liu et.al. 2606.24669 null
2026-06-23 CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning Xinyu Mao et.al. 2606.24636 null
2026-06-23 Beyond Monotonic Progress: Retry-Supervised Value Learning for Robot Imitation Xinyao Qin et.al. 2606.24633 null
2026-06-23 Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback Andreas Chouliaras et.al. 2606.24622 null
2026-06-23 Quantum-enabled active matter at the atomic scale Sabrina Burgardt et.al. 2606.24615 null
2026-06-23 Degeneracy-Aware Resilient Resource Allocation in Cell-Free Cache-Aided MU-MIMO Networks Sayanti Ghosh et.al. 2606.24611 null
2026-06-23 ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering Zhentao Guo et.al. 2606.24602 null
2026-06-23 ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning Anurag Akula et.al. 2606.24601 null
2026-06-23 PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought Ling Li et.al. 2606.24539 null
2026-06-23 VisCritic: Visual State Comparison as Process Reward for GUI Agents Jiachen Qian et.al. 2606.24525 null
2026-06-23 What Do Flow-Based Inverse Solvers Approximate? A Posterior-Transport View Jian Xu et.al. 2606.24516 null
2026-06-23 Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation Marta Sumyk et.al. 2606.24515 null
2026-06-23 Uncovering Latent Structures in Robust Pulse Sequences: A Model-Based Reinforcement Learning Approach for Adaptable Quantum Control Tobias Kiermeyer et.al. 2606.24507 null
2026-06-23 video-SALMONN-R $^3$ : Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Yixuan Li et.al. 2606.24477 null
2026-06-23 Entanglement and non-separability of momenta and coordinates at colliders Marco Fabbrichesi et.al. 2606.24468 null
2026-06-22 CoorDex: Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation Sikai Li et.al. 2606.23680 null
2026-06-22 AIR: Adaptive Interleaved Reasoning with Code in MLLMs Cong Han et.al. 2606.23678 null
2026-06-22 Resolving support-mismatch by local basis rotation in variational Monte Carlo Jia-Lin Chen et.al. 2606.23657 null
2026-06-22 Dynamic estimation of slowly varying sequences Prashant Gokhale et.al. 2606.23655 null
2026-06-22 Optimal Stopping for a Diffusion with Unobserved Bernoulli Drift Georgy Gaitsgori et.al. 2606.23648 null
2026-06-22 Learning Process Rewards via Success Visitation Matching for Efficient RL Raymond Tsao et.al. 2606.23640 null
2026-06-22 DiT-Reward: Generative Representations for Text-to-Image Reward Modeling Yuanming Yang et.al. 2606.23626 null
2026-06-22 Learning to See While Learning to Act: Diffusion Models for Active Perception in Robot Imitation Kuancheng Wang et.al. 2606.23625 null
2026-06-22 dVLA-RL: Reinforcement Learning over Denoising Trajectories for Discrete Diffusion Vision-Language-Action Models Yuhao Wu et.al. 2606.23623 null
2026-06-22 RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models Ulas Berk Karli et.al. 2606.23617 null
2026-06-22 MORL-A2C: Multi-Objective Reinforcement Learning Reranker for Optimizing Healthiness in MOPI-HFRS Aarya Vasantlal et.al. 2606.23603 null
2026-06-22 SPIRAL: Learning to Search and Aggregate Jubayer Ibn Hamid et.al. 2606.23595 null
2026-06-22 Two-Stage Optimization for Dynamic Line Rating and Energy Storage Deployment Abanish Tiwari et.al. 2606.23586 null
2026-06-22 Decentralized Autonomous Traffic Management through Corridor Networks Jasmine Jerry Aloor et.al. 2606.23585 null
2026-06-22 Protection Switching in Hybrid Hollow-Core and Single-Mode Fiber Networks: Challenges, Analysis, and Mitigation Strategies Md Ghulam Saber et.al. 2606.23554 null
2026-06-22 Structure-Aware Variance Reduction for Unbiased Randomized Hamiltonian Simulation Joshua W. Dai et.al. 2606.23544 null
2026-06-22 VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct Haoling Li et.al. 2606.23543 null
2026-06-22 SQLConductor: Search-to-Policy Learning for Step-wise Text-to-SQL Orchestration Yizhang Zhu et.al. 2606.23537 null
2026-06-22 BiliVLA: Scene-Aware Vision-Language-Action Model with Reinforcement Learning for Autonomous Biliary Endoscopic Navigation Jinsong Lin et.al. 2606.23531 null
2026-06-22 Towards an Automated Reasoning Tool for Complexity Analysis of Automated Reasoners Louis Rustenholz et.al. 2606.23516 null
2026-06-21 On the Position Bias of On-Policy Distillation Yan Xie et.al. 2606.22600 null
2026-06-21 Stationary Robust Mean-Field Games under Model Mismatches Yue Wang et.al. 2606.22579 null
2026-06-21 OASIS: Observation-Aware Simulation-Based Inference via Distributional Matching Arya Farahi et.al. 2606.22572 null
2026-06-21 What are Key Factors for Updates in RL for LLM Reasoning? Peidong Wang et.al. 2606.22570 null
2026-06-21 PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models Xianghui Wang et.al. 2606.22540 null
2026-06-21 Imagine to Ensure Safety in Hierarchical Reinforcement Learning Gregory Gorbov et.al. 2606.22509 null
2026-06-21 WebCQ: Cooperative Multi-Agent Deep Reinforcement Learning for Scalable Web GUI Testing Yujia Fan et.al. 2606.22502 null
2026-06-21 ARP: Enhancing Quantized Skill Abstractions via Visual Alignment and Iterative Refinement for Robotic Manipulation Yuntian Wang et.al. 2606.22480 null
2026-06-21 Scalable Multi-Task Data Generation via Reinforcement Learning for Language-Conditioned Bimanual Dexterous Manipulation Zechu Li et.al. 2606.22471 null
2026-06-21 Exact Nonnegative Matrix Factorization via Cone-Ray Witnesses: Obtuseness Ranking, Saturation Curves, and an Augmented Alt-LP Breakthrough Mithil Ramteke et.al. 2606.22451 null
2026-06-21 A Differentiable Atari VCS:A Complex, Fully Known Ground Truth for Explainable AI Andreas Maier et.al. 2606.22447 null
2026-06-21 Distribution-Aware Robust Bilevel Optimization: Quantile-Guided Huber Updates in Two-Timescale Stochastic Approximation Zhiyu Li et.al. 2606.22436 null
2026-06-21 Escaping the Variance Trap: Jacobian-Free Dynamics for Root-Finding Bilevel Optimization Zhiyu Li et.al. 2606.22433 null
2026-06-21 SVGym (SciVerseGym): An Environment for Reinforcement Learning and Bayesian Optimization in Crystal Discovery Bin Cao et.al. 2606.22425 null
2026-06-21 Reinforcement learning to improve large language model-based automated code compliance systems Jack Wei Lun Shi et.al. 2606.22402 null
2026-06-21 Do Rigid-Body Simulators Dream of Soft Robots? Learning Contact-Rich Manipulation for Tendon-Driven Continuum Robots Chengnan Shentu et.al. 2606.22397 null
2026-06-21 Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning Songtao Tian et.al. 2606.22394 null
2026-06-21 Select-to-Act: Hierarchical Reinforcement Learning via Adaptive Language Guidance Hanping Zhang et.al. 2606.22350 null
2026-06-21 Full Configuration Interaction Quantum Monte Carlo for Accurate $\textit{Ab Initio}$ Nuclear Structure Calculations Rongzhe Hu et.al. 2606.22341 null
2026-06-21 Lattice-quantile estimation of π and convex-region integrals from coined two-dimensional quantum walks Jen-Yu Chang et.al. 2606.22334 null
2026-06-18 Generating Robot Hands from Human Demonstrations Sha Yi et.al. 2606.20549 null
2026-06-18 Fast Human Attention Prediction for Fixation-guided Active Perception in Autonomous Navigation Fatma Youssef Mohammed et.al. 2606.20491 null
2026-06-18 Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users Haw-Shiuan Chang et.al. 2606.20482 null
2026-06-18 Correlated Mott semi-metal in the topological heavy fermion model Emile Pangburn et.al. 2606.20466 null
2026-06-18 ARC: Adaptive Robust Joint State and Covariance Estimation Alexandre Hadji-Thomas et.al. 2606.20428 null
2026-06-18 TaCauchy: An Extensible FEM Framework for Vision-Based Tactile Simulation Hengfei Zhao et.al. 2606.20426 null
2026-06-18 Neural network surrogates with uncertainty quantification for inverse problems in partial differential equations Christian Jimenez-Beltran et.al. 2606.20417 null
2026-06-18 Direct Advantage Estimation for Scalable and Sample-efficient Deep Reinforcement Learning Hsiao-Ru Pan et.al. 2606.20411 null
2026-06-18 CoLI: A Reproducible Platform for Continuum Robot Learning via Monolithic 3D Printing and Isomorphic Teleoperation Ziyuan Tang et.al. 2606.20389 null
2026-06-18 CRAX: Fast Safe Reinforcement Learning Benchmarking Tristan Tomilin et.al. 2606.20376 null
2026-06-18 Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining Yuexing Hao et.al. 2606.20363 null
2026-06-18 On the Variance of Temporal Difference Learning and its Reduction Using Control Variates Hsiao-Ru Pan et.al. 2606.20357 null
2026-06-18 A Model-Driven Approach for Developing Families of Reinforcement Learning Environments Xiaoran Liu et.al. 2606.20324 null
2026-06-18 Through the PRISM: Preference Representation in Intermediate States of Video Diffusion Models Haoxuan Wu et.al. 2606.20310 null
2026-06-18 ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval Yuhan Liu et.al. 2606.20280 null
2026-06-18 Effects of nonlocal interactions on s- and d-wave superconducting correlations in the extended Hubbard model Pavol Farkasovsky et.al. 2606.20260 null
2026-06-18 A Multi-Agent system for Multi-Objective constrained optimization Federica Filippini et.al. 2606.20236 null
2026-06-18 Reliable ORIS-assisted FSO Communications via HARQ Georgios D. Chondrogiannis et.al. 2606.20222 null
2026-06-18 Addressing uncertainties of model predictions for extensive air showers initiated by high energy cosmic rays Sergey Ostapchenko et.al. 2606.20221 null
2026-06-18 Probing Strange-Quark Hadronization via (Multi-)Strange Hadron Multiplicity Distributions in Small Collision Systems with ALICE Sara Pucillo et.al. 2606.20213 null
2026-06-17 Native Active Perception as Reasoning for Omni-Modal Understanding Zhenghao Xing et.al. 2606.19341 null
2026-06-17 Learning User Simulators with Turing Rewards Yingshan Susan Wang et.al. 2606.19336 null
2026-06-17 UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning Mohamed Nabail et.al. 2606.19328 null
2026-06-17 Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation Siyi Gu et.al. 2606.19327 null
2026-06-17 Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation Xin Ci Wong et.al. 2606.19300 null
2026-06-17 Accelerating Network-Agent Dispersion: Territorial Behavior and Directionally Biased Lazy Random Walks Li Zeng et.al. 2606.19294 null
2026-06-17 Quantum-Classical Auxiliary-Field Quantum Monte Carlo at the Edge of Practicability Francesco Nappi et.al. 2606.19239 null
2026-06-17 STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability Haipeng Luo et.al. 2606.19236 null
2026-06-17 A Human-in-the-Loop Bayesian Optimization Framework for Constraint-Aware Bioprocess Development Samuel Stricker et.al. 2606.19230 null
2026-06-17 Discovering a well-conditioned analytic continuation problem via dictionary learning Thomas Chuna et.al. 2606.19205 null
2026-06-17 Forecasting what Matters: Decision-Focused RL for Controlled EV Charging with Unknown Departure Times Giuseppe Gabriele et.al. 2606.19199 null
2026-06-17 Direct large-area observation of subsurface plastic activity in conditioned copper electrodes Yinon Ashkenazy et.al. 2606.19192 null
2026-06-17 The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL Nicolas Beltran-Velez et.al. 2606.19162 null
2026-06-17 Pareto Q-Learning with Reward Machines Arnaud Lequen et.al. 2606.19134 null
2026-06-17 Quantifying Compromise Risk in Exceptional Access Architectures Under Sparse and Indirect Evidence Alan Woodward et.al. 2606.19106 null
2026-06-17 ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL Mukund Khanna et.al. 2606.19103 null
2026-06-17 Byzantine-Resilient Federated Multi-Agent Optimization Framework for Cyber-Secure Interconnected Microgrids Ali Peivand et.al. 2606.19080 null
2026-06-17 A Measure-Valued Obstacle Problem for an Obliquely Reflected Diffusion with a Max-Type Payoff Louis Shuo Wang et.al. 2606.19070 null
2026-06-17 Model-Free Reinforcement Learning Control for Resilient Cyber-Physical Systems Hugo O. Garcés et.al. 2606.19069 null
2026-06-17 An extendable, integrated, and dynamic approach to forecasting and stress-testing credit risk Marcel Muller et.al. 2606.19052 null
2026-06-16 Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification Wujian Peng et.al. 2606.18249 null
2026-06-16 Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents Ankita Samaddar et.al. 2606.18223 null
2026-06-16 Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Byung-Kwan Lee et.al. 2606.18216 null
2026-06-16 Beyond Plane Waves: Coherent Network Response to Collimated Gravitational-Wave Wavepackets S. D. Campos et.al. 2606.18184 null
2026-06-16 High-temperature ferromagnetism and antiferromagnetism in monolayer \ce{CrTe2}: Roles of strong spin-lattice coupling and charge doping Anupama S et.al. 2606.18148 null
2026-06-16 Spatial Disease Mapping and Disparity Detection Using Generative AI: An Amortized Bayesian Learning Framework Luca Aiello et.al. 2606.18146 null
2026-06-16 Knowledge Reutilization in Meta-Reinforcement Learning Yuan Meng et.al. 2606.18132 null
2026-06-16 Undocumented Behavior in the gsynth R package and its Consequences for Three Published Studies Beniamino Green et.al. 2606.18113 null
2026-06-16 Learning Fair Pareto-Optimal Policies in Multi-Objective Reinforcement Learning Umer Siddique et.al. 2606.18111 null
2026-06-16 Deep Reinforcement Learning for Minimum Zero-Forcing Sets Steve Halley et.al. 2606.18106 null
2026-06-16 OmniPlan: An Adaptive Framework for Timely and Near-Optimal Network Planning Optimization Longlong Zhu et.al. 2606.18105 null
2026-06-16 WireCraft: A Simulation Benchmark for Industrial DLO Manipulation Chongyu Zhu et.al. 2606.18097 null
2026-06-16 From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Lingjing Kong et.al. 2606.18089 null
2026-06-16 Beyond Parallel Trends in Staggered Difference-in-Differences: Identification under Higher-Order Parallelism Zecharias Anteneh et.al. 2606.17977 null
2026-06-16 Endogenous business cycles via state-dependent saving and noise-induced metastability Shenglan Yuan et.al. 2606.17946 null
2026-06-16 WAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFT Zezhong Qian et.al. 2606.17906 null
2026-06-16 Time-Slotted Multi-Cluster UAV AirComp with Energy-Awareness: A Pointer Network-Assisted Soft Actor-Critic Learning Framework Xunqiang Lan et.al. 2606.17900 null
2026-06-16 Detectability of deuterium in spectra of early-type stars Veronika Mitrokhina et.al. 2606.17896 null
2026-06-16 Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models Zihao Wei et.al. 2606.17890 null
2026-06-16 StepGuard: Guarding Web Navigation via Single-Step Calibration Zhihao Cui et.al. 2606.17871 null
2026-06-15 The Value Axis: Language Models Encode Whether They’re on the Right Track Nick Jiang et.al. 2606.17056 null
2026-06-15 Context-Aware RL for Agentic and Multimodal LLMs Peiyang Xu et.al. 2606.17053 null
2026-06-15 Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes Tongyan Fang et.al. 2606.17043 null
2026-06-15 R2RDreamer: 3D-aware Data Augmentation for Spatially-generalized 2D Manipulation Policies Xiuwei Xu et.al. 2606.17040 null
2026-06-15 DEEPRUBRIC: Evidence-Tree Rubric Supervision for Efficient Reinforcement Learning of Deep Research Agents Minghang Zhu et.al. 2606.17029 null
2026-06-15 ExpRL: Exploratory RL for LLM Mid-Training Violet Xiang et.al. 2606.17024 null
2026-06-15 ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning Wei Xiao et.al. 2606.17011 null
2026-06-15 TuneJury: An Open Metric for Improving Music Generation Preference Alignment Yonghyun Kim et.al. 2606.17006 null
2026-06-15 When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning Nathan Gavenski et.al. 2606.16995 null
2026-06-15 DreamX-World 1.0: A General-Purpose Interactive World Model DreamX Team et.al. 2606.16993 null
2026-06-15 Grassmannian quantum cohomology in the infinite limit and total positivity Ines Chung-Halpern et.al. 2606.16983 null
2026-06-15 Task-Error Residual Learning for Real-Robot Five-Ball Juggling Kai Ploeger et.al. 2606.16978 null
2026-06-15 Exploring Extrinsic and Intrinsic Properties for Effective Reasoning with Code Interpreter Patomporn Payoungkhamdee et.al. 2606.16934 null
2026-06-15 A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning Ardianto Wibowo et.al. 2606.16933 null
2026-06-15 Greed Is Learned: Visible Incentives as Reward-Hacking Triggers Tong Che et.al. 2606.16914 null
2026-06-15 Latent Space Reinforcement Learning for Inverse Material Estimation in Food Fracture Simulation Adrian Ramlal et.al. 2606.16870 null
2026-06-15 Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement Learning Tung M. Luu et.al. 2606.16856 null
2026-06-15 Deep Q-Learning on Hölder Spaces Qian Qi et.al. 2606.16846 null
2026-06-15 Understanding the Behaviors of Environment-aware Information Retrieval Ruifeng Yuan et.al. 2606.16817 null
2026-06-15 OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models Tianyi Lin et.al. 2606.16774 null
2026-06-12 Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning Pengxin Wang et.al. 2606.14693 null
2026-06-12 CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment Jiayue Cao et.al. 2606.14691 null
2026-06-12 AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition Jixuan Chen et.al. 2606.14674 null
2026-06-12 HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities Yijun Liu et.al. 2606.14657 null
2026-06-12 Graph Structured Combinatorial Semi-Bandit with Nonlinear Reward Associations through Separable Signals Christoph Bauschmann et.al. 2606.14650 null
2026-06-12 Open Wilson chain numerical renormalization group approach to steady-state non-equilibrium quantum transport Anand Manaparambil et.al. 2606.14635 null
2026-06-12 Safe Reinforcement Learning of Autonomous Highway Driving: A Unified Framework for Safety and Efficiency Chufei Yan et.al. 2606.14609 null
2026-06-12 A Statistical and Machine Learning Framework for Operational Threshold Detection and Deployable Dispatch Controller Development in Hydrogen Multi-Energy Systems Shadi Heenatigala et.al. 2606.14601 null
2026-06-12 VISTA: View-Consistent Self-Verified Training for GUI Grounding Xinyu Qiu et.al. 2606.14579 null
2026-06-12 Tomography of Atomic Nuclei Noemi Rocco et.al. 2606.14558 null
2026-06-12 On the design distribution for predictive Bayesian regression Wanyue Sun et.al. 2606.14544 null
2026-06-12 Provably Safe, Yet Scalable Reinforcement Learning Kai S. Yun et.al. 2606.14536 null
2026-06-12 Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera Seoyoon Kim et.al. 2606.14535 null
2026-06-12 From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Yongheng Zhang et.al. 2606.14502 null
2026-06-12 Quantum Horizon: An evaluation of quantum computing as a threat to Bitcoin and Ethereum Iosif M. Gershteyn et.al. 2606.14484 null
2026-06-12 Kine2Go: Kinematic dataset for the Unitree Go2 robot with diverse gaits and motions Władysław Pałucki et.al. 2606.14433 null
2026-06-12 SmoQyElPhQMC.jl: An open-source Julia package for efficient and scalable quantum Monte Carlo simulations of electron-phonon coupled models Benjamin Cohen-Stead et.al. 2606.14425 null
2026-06-12 Causal Object-Centric Models for Planning with Monte Carlo Tree Search Rodion Vakhitov et.al. 2606.14418 null
2026-06-12 CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning Ayoub Belouadah et.al. 2606.14415 null
2026-06-12 Predictive Concordance for Parameter Optimisation and Mixture Synthesis Tobias Adrian et.al. 2606.14382 null
2026-06-11 Mana: Dexterous Manipulation of Articulated Tools Zhao-Heng Yin et.al. 2606.13677 null
2026-06-11 Improving Robotic Generalist Policies via Flow Reversal Steering Andy Tang et.al. 2606.13675 null
2026-06-11 Aerial Wildfire Suppression Planning with a Hybrid CNN-Cellular Automata Fire Model Ion Matei et.al. 2606.13633 null
2026-06-11 Beyond Runtime Enforcement: Shield Synthesis as Defensibility Analysis for Adversarial Networks Achraf Hsain et.al. 2606.13621 null
2026-06-11 Distribution-Agnostic Robust Trajectory Optimization via Chance-Constrained Reinforcement Learning Yashdeep Chaudhary et.al. 2606.13605 null
2026-06-11 Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch Haochen Wu et.al. 2606.13604 null
2026-06-11 Feasibility of up-the-ramp sampling under variable sky for ground-based spectrographs Gaia Gaspar et.al. 2606.13600 null
2026-06-11 Reward Modeling for Multi-Agent Orchestration King Yeung Tsang et.al. 2606.13598 null
2026-06-11 Smoothed Rank-Based Regression Estimation Using Wilcoxon Score Functions Feridun Tasdan et.al. 2606.13593 null
2026-06-11 ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages Tanmoy Kanti Halder et.al. 2606.13572 null
2026-06-11 A local Universe catalogue of structures and voids dynamically identified using Cosmic-Flows4++ZOA peculiar velocities A. M. Hollinger et.al. 2606.13538 null
2026-06-11 AgentRivet: an automated system for producing Rivet routines from journal publications Antonio J. Costa et.al. 2606.13535 null
2026-06-11 Quasi-2D trapped tilted dipoles at zero and finite temperatures in the strongly dipolar regime Juan Sánchez-Baena et.al. 2606.13502 null
2026-06-11 Population dynamics of surface-mediated autocatalytic processes Denis S. Grebenkov et.al. 2606.13498 null
2026-06-11 Reinforcement Learning for Neural Model Editing Shaivi Malik et.al. 2606.13461 null
2026-06-11 When expectation fails: stochastic MPC of linear systems with random input losses Paul Trodden et.al. 2606.13421 null
2026-06-11 The Influence of Gain and Phase Mismatches on Beam Patterns in Phased Arrays Jérémy Guichemerre et.al. 2606.13378 null
2026-06-11 IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing Tao Hu et.al. 2606.13368 null
2026-06-11 From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent Haishuo Fang et.al. 2606.13349 null
2026-06-11 Improved Runtime Bound for the $(μ+ 1)$ EA on BinVal Joris Belder et.al. 2606.13344 null
2026-06-10 ATLAS: Active Theory Learning for Automated Science Noémi Éltető et.al. 2606.12386 null
2026-06-10 APPO: Agentic Procedural Policy Optimization Xucong Wang et.al. 2606.12384 null
2026-06-10 Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization Hao Xiang et.al. 2606.12373 null
2026-06-10 UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning Haoyuan Deng et.al. 2606.12372 null
2026-06-10 Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling Yucheng Li et.al. 2606.12370 null
2026-06-10 Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics Adam Wei et.al. 2606.12365 null
2026-06-10 An Efficient Method for the Optimal Control of Microgrids Under Uncertainties using Local Reduction Edoardo Scaccia et.al. 2606.12345 null
2026-06-10 Fair Comparison of Scheduling Algorithms on Heterogeneous Edge Clusters: A Continuous Adaptive Benchmark Zihang Wang et.al. 2606.12343 null
2026-06-10 Fourier Features Let Agents Learn High Precision Policies with Imitation Learning Balázs Gyenes et.al. 2606.12334 null
2026-06-10 CCKS: Consensus-based Communication and Knowledge Sharing Jinyuan Zu et.al. 2606.12281 null
2026-06-10 Mathematical perspective on genetic algorithms with optimization guided operators Anna Brandenberger et.al. 2606.12279 null
2026-06-10 Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language Models Jia Deng et.al. 2606.12273 null
2026-06-10 Rbreak: An R Package for Estimating Structural Breaks under Linear Restrictions with Application to Linear Model Tree Cheolju Kim et.al. 2606.12261 null
2026-06-10 Reinforcement Learning Disrupts Gradient-Based Adversarial Optimization Xinhai Zou et.al. 2606.12251 null
2026-06-10 DrivingAgent: Design and Scheduling Agents for Autonomous Driving Systems Zhongyu Xia et.al. 2606.12236 null
2026-06-10 InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Ziang Yan et.al. 2606.12195 null
2026-06-10 Shared Infrastructure Investment and Pricing: Stackelberg Equilibria in Risk-Aware Take-or-Pay Contracts Amal Sakr et.al. 2606.12167 null
2026-06-10 Saturation of Nuclear Binding from Lattice Hamiltonians Maxwell Rothman et.al. 2606.12166 null
2026-06-10 Ionization-Induced Electrostatic Hose Instability in Electron-Beam-Sustained Plasmas Jia-Hong Chen et.al. 2606.12127 null
2026-06-10 The Bishop–Phelps–Bollobás Property for Extremally Disconnected Ranges: Separable and Low-Density Domains Tattwamasi Amrutam et.al. 2606.12080 null
2026-06-09 ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations Junke Wang et.al. 2606.11188 null
2026-06-09 TacForeSight: Force-Guided Tactile World Model for Contact-Rich Manipulation Yujie Zang et.al. 2606.11184 null
2026-06-09 Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models Atsumoto Ohashi et.al. 2606.11167 null
2026-06-09 Quantum Monte Carlo calculations of Zemach moments in $A\leq 9$ nuclei Garrett B. King et.al. 2606.11153 null
2026-06-09 Data assimilation for subsurface flow using latent diffusion model parameterization: performance of ensemble-Kalman and Monte Carlo techniques Guido Di Federico et.al. 2606.11140 null
2026-06-09 Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation Soham Bhattacharjee et.al. 2606.11127 null
2026-06-09 Monte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in Football Andrew Kang et.al. 2606.11120 null
2026-06-09 TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning Heming Zou et.al. 2606.11119 null
2026-06-09 Limitations of Learning Tanh Neural Networks with Finite Precision Philipp Grohs et.al. 2606.11104 null
2026-06-09 A data-driven method for measuring corner-clipping probabilities in segmented particle detectors Joaquín de Jesús et.al. 2606.11097 null
2026-06-09 RoboNaldo: Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning Yichao Zhong et.al. 2606.11092 null
2026-06-09 Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Zhiyuan Zhou et.al. 2606.11087 null
2026-06-09 On the representation for stochastic graph delay propagation Shibo Zeng et.al. 2606.11086 null
2026-06-09 Exploring the Design Space of Reward Backpropagation for Flow Matching Ruoyu Wang et.al. 2606.11075 null
2026-06-09 Range of Normalized Glandular Dose for Mammography Using Patient-Specific Glandular Fractions Lacey L. Medlock et.al. 2606.11071 null
2026-06-09 LLM-Mediated Demand Response Coordination in Smart Microgrids J. de Curtò et.al. 2606.11050 null
2026-06-09 What Fits (Into Few Tokens) Doesn’t Overfit: Compression and Generalization in ML Research Agents Martin Andres Bertran et.al. 2606.11045 null
2026-06-09 Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models Bowen Ping et.al. 2606.11025 null
2026-06-09 FairWave : A Fairness-Aware Asynchronous DAG-BFT Consensus Syariful Mujaddiq et.al. 2606.10982 null
2026-06-09 Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets Yi Chen et.al. 2606.10979 null
2026-06-08 An Agency-Transferring Model-Free Policy Enhancement Technique Anton Bolychev et.al. 2606.09825 null
2026-06-08 Rethinking the Divergence Regularization in LLM RL Jiarui Yao et.al. 2606.09821 null
2026-06-08 Finite-n Estimate of Dedekind Numbers by Layer-Ratio Monte Carlo Tian-Shun Chen et.al. 2606.09795 null
2026-06-08 Certified spectral functions from lattice Monte Carlo data Sophie Mutzel et.al. 2606.09791 null
2026-06-08 Preserving Plasticity in Continual Learning via Dynamical Isometry Andries Rosseau et.al. 2606.09762 null
2026-06-08 Difference-Aware Retrieval Policies for Imitation Learning Quinn Pfeifer et.al. 2606.09758 null
2026-06-08 Jamming-Resilient Sparse Delay-Doppler NOMA: Unitary Precoding, Randomized Active Sets, and Superincreasing Power Allocation Michel Kulhandjian et.al. 2606.09753 null
2026-06-08 The Neutral Mask: How RLHF Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model Wendy K. Tam et.al. 2606.09735 null
2026-06-08 Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO Blake Bullwinkel et.al. 2606.09701 null
2026-06-08 Gradient-Guided Reward Optimization for Inference-time Alignment Hankun Lin et.al. 2606.09635 null
2026-06-08 DexPIE: Stable Dexterous Policy Improvement from Real-World Experience Ruizhe Liao et.al. 2606.09615 null
2026-06-08 Shape Formation for the Cooperative Transportation of Arbitrary Objects Using Multi-Agent Reinforcement Learning Mohamed Sayed et.al. 2606.09610 null
2026-06-08 Path-Traced Inverse Rendering with Global Illumination in 3D Gaussian Fields Junke Zhu et.al. 2606.09606 null
2026-06-08 UXBench: Benchmarking User Experience in AI Assistants Mengze Hong et.al. 2606.09570 null
2026-06-08 Safe-RULE: Safe Reinforcement UnLEarning Shixiong Jiang et.al. 2606.09559 null
2026-06-08 Emergence of Context Characteristics Sensitivity in Large Language Models Nadya Yuki Wangsajaya et.al. 2606.09525 null
2026-06-08 Orbital Plane Geometry and Information Conditioning for Doppler-Only LEO Positioning Charles E Thornton et.al. 2606.09496 null
2026-06-08 Primordial Black Holes from Slow Phase Transitions with Delayed Reheating: A Peak-Theory Approach Indra Kumar Banerjee et.al. 2606.09482 null
2026-06-08 AliyunConsoleAgent: Training Web Agents in Real-World Cloud Environments via Distillation and Reinforcement Learning Bojie Rong et.al. 2606.09447 null
2026-06-08 PriFT: Prior-Support Guided Supervised Fine-Tuning Ke Wang et.al. 2606.09396 null
2026-06-08 CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Penghui Yang et.al. 2606.09393 null
2026-06-08 An Introduction to Measurement Uncertainty Samanta Piano et.al. 2606.09385 null
2026-06-08 ReGIL: Retrieval-Guided Imitation Learning from a Single Demonstration Yuying Zhang et.al. 2606.09381 null
2026-06-08 Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short Han Zhou et.al. 2606.09380 null
2026-06-08 Coupling Complementary Simulations for Combined Performance and Energy Optimization Adel Dabah et.al. 2606.09356 null
2026-06-08 PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Yang Tian et.al. 2606.09348 null
2026-06-08 TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation Huaihang Zheng et.al. 2606.09337 null
2026-06-08 SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling Haoran Xu et.al. 2606.09304 null
2026-06-08 One Model, Multiple Goals: Adaptive Multi-Objective Learning for E-commerce Dialogue Systems Mingzhe Li et.al. 2606.09293 null
2026-06-05 Affordance-Based Hierarchical Reinforcement Learning for Quadruped Pedipulation Tuba Girgin et.al. 2606.07506 null
2026-06-05 Modelling Opinion Dynamics at Scale with Deep MARL Lukas Seier et.al. 2606.07487 null
2026-06-05 Rapid co-design of Buoyancy-assisted robots for Challenging Locomotion using Gaussian Evolutionary Specialists Ankit Sinha et.al. 2606.07424 null
2026-06-05 Generative Modeling of Discrete Latent Structures via Dynamic Policy Gradients Stefan Ivanovic et.al. 2606.07400 null
2026-06-05 Simulation-Driven Imitation Learning for Biosignals-Free Shared-Autonomy Prosthetic Grasping Kaijie Shi et.al. 2606.07389 null
2026-06-05 Spline Policy: A Structured Representation for Robot Policies Mengze Tian et.al. 2606.07386 null
2026-06-05 Self-evolving LLM agents with in-distribution Optimization Yudi Zhang et.al. 2606.07367 null
2026-06-05 KIT’s Submission to Cross-Lingual Voice Cloning in IWSLT 2026 Seymanur Akti et.al. 2606.07240 null
2026-06-05 Learning Multi-Agent Communication Protocol: Study on Information Entropy Efficiency in MARL Xinren Zhang et.al. 2606.07200 null
2026-06-05 On the true low-energy excitations of the three-dimensional spin glass Claudio Chilin et.al. 2606.07197 null
2026-06-05 Shield-Loco: Shielding Locomotion Policies with Predictive Safety Filtering Aditya Shirwatkar et.al. 2606.07193 null
2026-06-05 From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning Yuhang Zhou et.al. 2606.07190 null
2026-06-05 VALO1.0: New real-photon parton distributions with Monte Carlo uncertainties Madhav Chithirasreemadam et.al. 2606.07189 null
2026-06-05 Phase diagram of the extended chequerboard $J-Q$ model Jiayou Yin et.al. 2606.07178 null
2026-06-05 $α$ -PFN: Fast Entropy Search via In-Context Learning Herilalaina Rakotoarison et.al. 2606.07134 null
2026-06-05 Predictive Style Matching: Natural and Robust Humanoid Locomotion Simeon Nedelchev et.al. 2606.07083 null
2026-06-05 On the Geometry of On-Policy Distillation Zhennan Shen et.al. 2606.07082 null
2026-06-05 SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating Zequn Xie et.al. 2606.07074 null
2026-06-05 Exact noise characterization of entanglement distribution in star networks Kenneth Goodenough et.al. 2606.07043 null
2026-06-05 Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling Xing Yue et.al. 2606.07040 null
2026-06-04 TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies Dong Jing et.al. 2606.06491 null
2026-06-04 RREDCoT: Segment-Level Reward Redistribution for Reasoning Models Mykyta Ielanskyi et.al. 2606.06475 null
2026-06-04 Latent Reasoning with Normalizing Flows Guancheng Tu et.al. 2606.06447 null
2026-06-04 Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation Hanxu Hu et.al. 2606.06428 null
2026-06-04 Nonreversible Gauge Fields in Fokker–Planck Dynamics: Supersymmetric Hamiltonians and Learned Finite Forces Masayuki Ohzeki et.al. 2606.06412 null
2026-06-04 A high-energy neutrino flare associated with nearby bright interacting supernova SN 2021foa Ming-Xuan Lu et.al. 2606.06409 null
2026-06-04 Emergent Language as an Approach to Conscious AI Zengqing Wu et.al. 2606.06380 null
2026-06-04 Maximising the Set-Piece Return: Optimising Football Corner Tactics with Graph Reinforcement Learning Sean Groom et.al. 2606.06353 null
2026-06-04 EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading Zhihao Wu et.al. 2606.06350 null
2026-06-04 VOLT: Vision and Language Trajectory Segmentation for Faster-than-Demonstration Policies Robert Ramirez Sanchez et.al. 2606.06323 null
2026-06-04 Discrete Causal Representations from Heterogeneous Domains: A Bayesian Approach with Social Survey Applications Ankur Garg et.al. 2606.06288 null
2026-06-04 Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation Rickmer Krohn et.al. 2606.06281 null
2026-06-04 SecRL-Prune: Structured Reinforcement Learning-Based Pruning of CodeLLMs for Preserving Adversarial Code Mutation Parsa Memarzadehsaghezi et.al. 2606.06254 null
2026-06-04 Interdependent Hitting Times Jaap H. Abbring et.al. 2606.06251 null
2026-06-04 Drag reduction or reward hacking? Recurrent multi-agent reinforcement learning that earns its reward Giorgio Maria Cavallazzi et.al. 2606.06227 null
2026-06-04 DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments Tan Zhang et.al. 2606.06217 null
2026-06-04 Learning to replenish: A hybrid deep reinforcement learning for dynamic inventory management in the pharmaceutical supply chains Amandeep Kaur et.al. 2606.06201 null
2026-06-04 Deep reinforcement learning with spatial and temporal awareness for active boundary control of buoyancy-driven convection Giorgio Maria Cavallazzi et.al. 2606.06191 null
2026-06-04 ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity Prathamjyot Singh et.al. 2606.06168 null
2026-06-04 Learning to Contest: Decentralized Robust Fairness in Cooperative MARL via Cross-Attention Can Savcı et.al. 2606.06162 null
2026-05-29 LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards Nianyi Lin et.al. 2605.31584 null
2026-05-29 Preference-Aware Rubric Learning for Personalized Evaluation Yilun Qiu et.al. 2605.31545 null
2026-05-29 Value Functions as Supermartingale Certificates Alessandro Abate et.al. 2605.31524 null
2026-05-29 Bayesian Nonparametric Clustering to Support Medical Decision-Making: A Variational Inference Approach Inga Huld Ármann et.al. 2605.31511 null
2026-05-29 Skill Reuse as Compression in Agentic RL Zhikun Xu et.al. 2605.31509 null
2026-05-29 Are Full Rollouts Necessary for On-Policy Distillation? Yaocheng Zhang et.al. 2605.31490 null
2026-05-29 Learning Controlled Separation of Small Objects Between Two Fingers with a Tactile Skin Ulf Kasolowsky et.al. 2605.31486 null
2026-05-29 Batched Differentiable Rigid Body Dynamics in PyTorch for GPU-Accelerated Robot Learning Yue Wang et.al. 2605.31481 null
2026-05-29 GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization Zaid Khan et.al. 2605.31464 null
2026-05-29 DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization Jian Mu et.al. 2605.31455 null
2026-05-29 Answer-Set-Programming-based Abstractions for Reinforcement Learning Rafael Bankosegger et.al. 2605.31444 null
2026-05-29 Astra: a generalizable report generation foundation model for 3D computed tomography Zhuhao Wang et.al. 2605.31437 null
2026-05-29 Improved Guarantees for Langevin Monte Carlo with Average Smoothness Arnak S. Dalalyan et.al. 2605.31413 null
2026-05-29 Constrained Multi-Objective Reinforcement Learning with Max-Min Criterion Giseung Park et.al. 2605.31388 null
2026-05-29 Unlocking Fine-Grained Translation Quality Estimation in LRMs through Synergistically Evolving Implicit and Explicit Reasoning Renfei Dang et.al. 2605.31378 null
2026-05-29 Dreaming Of Others: Latent Teammate Modeling In World Models For Multi-Agent Reinforcement Learning Tomas Leroy-Stone et.al. 2605.31361 null
2026-05-29 Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards Magnus Jørgenvåg et.al. 2605.31328 null
2026-05-29 Surface Constraint Policy for Learning Surface-Constrained and Dynamically Feasible Robot Skills Shuai Ke et.al. 2605.31321 null
2026-05-29 Generalized Intention Modeling in Multi-Agent Reinforcement Learning Mateusz Odrowaz-Sypniewski et.al. 2605.31318 null
2026-05-29 Model-free LQG Control with Chance Constraints Arunava Naha et.al. 2605.31310 null
2026-05-28 Reasoning with Sampling: Cutting at Decision Points Felix Zhou et.al. 2605.30327 null
2026-05-28 In-Context Reward Adaptation for Robust Preference Modeling Zhenyu Sun et.al. 2605.30323 null
2026-05-28 Enhanced Loading of a Molecular Magneto-Optical Trap Ebram Youssef et.al. 2605.30296 null
2026-05-28 Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection Yutong Wang et.al. 2605.30274 null
2026-05-28 Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning Ciara Rowles et.al. 2605.30257 null
2026-05-28 Reinforcement Learning with Robust Rubric Rewards Ya-Qi Yu et.al. 2605.30244 null
2026-05-28 Electronic correlations driving Chirality-Induced Spin Selectivity Jacek Herbrych et.al. 2605.30240 null
2026-05-28 How’s it going? Reinforcement learning in language models recruits a functional welfare axis Andy Q Han et.al. 2605.30232 null
2026-05-28 BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models Zhongxi Chen et.al. 2605.30226 null
2026-05-28 TriSearch: Learning to Optimize Triangulations via Bistellar Flips Yiran Wang et.al. 2605.30220 null
2026-05-28 When Should Models Change Their Minds? Contextual Belief Management in Large Language Models Haoming Xu et.al. 2605.30219 null
2026-05-28 HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Mohamed Sana et.al. 2605.30201 null
2026-05-28 Active Continual Learning with Metaplastic Binary Bayesian Neural Networks Kellian Cottart et.al. 2605.30198 null
2026-05-28 Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents Wenhao Li et.al. 2605.30190 null
2026-05-28 On Distributional Reinforcement Learning in Chaotic Dynamical Systems James Rudd-Jones et.al. 2605.30160 null
2026-05-28 Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents Ziyan Liu et.al. 2605.30159 null
2026-05-28 RL2ML: Finite-Rollout Surrogate Objectives from Reinforcement Learning to Maximum Likelihood Yifu Zheng et.al. 2605.30154 null
2026-05-28 Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies Kajetan Schweighofer et.al. 2605.30148 null
2026-05-28 FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection Leqi Zhu et.al. 2605.30062 null
2026-05-28 Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance Shutong Ding et.al. 2605.30056 null
2026-05-25 MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research Dingbang Wu et.al. 2605.26114 null
2026-05-25 Reinforcing Few-step Generators via Reward-Tilted Distribution Matching Yushi Huang et.al. 2605.26108 null
2026-05-25 On-Policy Adversarial Flow Distillation for Autoregressive Video Generation Yang Luo et.al. 2605.26105 null
2026-05-25 Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning Zhaoyu Zhu et.al. 2605.26078 null
2026-05-25 AI-Powered Sustainable Finance: An Integrative Taxonomy and Framework of AI Applications for Sustainable Investment Decision-Making Eduardo C. Garrido-Merchán et.al. 2605.26076 null
2026-05-25 X-ray Polarization Signatures from Comptonization by Magnetic Reconnection Plasmoids John Groger et.al. 2605.26065 null
2026-05-25 Accelerating Bayesian inverse design in computational fluid dynamics using neural operators Bipin Tiwari et.al. 2605.26059 null
2026-05-25 Quantile autoregressive moving average models for ratio-based bounded time series Helton Saulo et.al. 2605.26052 null
2026-05-25 Uncovering multi-channel magnetic hopfion annihilation via a single-node, billion-spin-scale atomistic framework Qichen Xu et.al. 2605.26016 null
2026-05-25 AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models Branislav Kveton et.al. 2605.26013 null
2026-05-25 Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning Aleksandar Todorov et.al. 2605.26012 null
2026-05-25 MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control Bin Li et.al. 2605.26006 null
2026-05-25 Causal methods for LLM development and evaluation Dennis Frauen et.al. 2605.25998 null
2026-05-25 SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation Michael Orme et.al. 2605.25984 null
2026-05-25 Reversible-jump MCMC reveals binary black hole subpopulations with distinct redshift evolution April Qiu Cheng et.al. 2605.25980 null
2026-05-25 LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation Jiabei Xiao et.al. 2605.25964 null
2026-05-25 DetMesh-Gadep: Triangulated Surface Modeling and GPU-based Monte Carlo Efficiency Calibration of High-Purity Germanium Detectors Kainan Zhang et.al. 2605.25963 null
2026-05-25 CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS Junyang Chen et.al. 2605.25930 null
2026-05-25 Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization Meshal Alamr et.al. 2605.25928 null
2026-05-25 Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning Wei Fan et.al. 2605.25920 null
2026-05-22 Geo-Align: Video Generation Alignment via Metric Geometry Reward Zizun Li et.al. 2605.23903 null
2026-05-22 Robotic Strawberry Harvesting with Robust Vision and Deep Reinforcement Learning based Sim-to-Real Control Al Bashir et.al. 2605.23863 null
2026-05-22 TCAD + Allpi $\text{x}^2$ Simulation study of MALTA2, a Depleted Monolithic Active Pixel Sensor for future tracking L. Li et.al. 2605.23860 null
2026-05-22 Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion Remko Proesmans et.al. 2605.23847 null
2026-05-22 Debiased Negative Mining Improves Out-of-distribution Detection with Pre-trained Vision-Language Models Bo Peng et.al. 2605.23797 null
2026-05-22 FPGA Acceleration of Matrix-Element Calculations for Monte Carlo Event Generation H. Gutiérrez Arance et.al. 2605.23785 null
2026-05-22 Direct Dynamic Retargeting for Humanoid Imitation Learning from Videos Constant Roux et.al. 2605.23762 null
2026-05-22 SeedER: Seed-and-Expand Retrieval from Knowledge Graphs Hamed Shirzad et.al. 2605.23753 null
2026-05-22 Vision-Based Agile Landing on Turbulent Waters Dimosthenis Angelis et.al. 2605.23717 null
2026-05-22 OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations Jiangwang Chen et.al. 2605.23668 null
2026-05-22 One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents Yoosung Hong et.al. 2605.23652 null
2026-05-22 Learning Kernel-Based MDPs from Episodic Preferential Feedback Nikola Pavlovic et.al. 2605.23650 null
2026-05-22 Less Effort, Shorter Proofs: Reinforcement Learning for Security Protocol Analysis in Tamarin Matthias Cosler et.al. 2605.23643 null
2026-05-22 Dirichlet-Based Monte Carlo Dropout for Uncertainty Estimation in Neural Networks Rouaa Hoblos et.al. 2605.23635 null
2026-05-22 Directional subset simulation method for reliability analysis Oindrila Kanjilal et.al. 2605.23631 null
2026-05-22 First-principles transition-state tensorial cluster expansion of vacancy diffusion in Ta-W beyond the kinetically-resolved activation approximation Jacob Jeffries et.al. 2605.23612 null
2026-05-22 A Markov-Chain-Monte-Carlo-based Hybrid Noise Inference for Continuous Wavelet Power Spectra: with Applications to Solar and Stellar Oscillatory Signals Song Feng et.al. 2605.23587 null
2026-05-22 Understanding Goal Generalisation in Sequential Reinforcement Learning Jason Ross Brown et.al. 2605.23565 null
2026-05-22 ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning Elie Abboud et.al. 2605.23562 null
2026-05-22 SafeSABR: Risk-Calibrated Adaptive Bitrate Streaming over Starlink Networks Hongjun Xie et.al. 2605.23560 null
2026-05-21 Vector Policy Optimization: Training for Diversity Improves Test-Time Search Ryan Bahlous-Boldi et.al. 2605.22817 null
2026-05-21 Remember to be Curious: Episodic Context and Persistent Worlds for 3D Exploration Lily Goli et.al. 2605.22814 null
2026-05-21 DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback Yunpeng Dong et.al. 2605.22781 null
2026-05-21 Deep Reinforcement Learning for Flexible Job Shop Scheduling with Random Job Arrivals Yu Tang et.al. 2605.22773 null
2026-05-21 Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning Ismail Geles et.al. 2605.22748 null
2026-05-21 Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation Dong Nie et.al. 2605.22731 null
2026-05-21 N3P: Accelerated Automated Parking via a Learning-Based Naturalistic Three-Stage Scheme Yifan Xue et.al. 2605.22722 null
2026-05-21 Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators Zachary Novack et.al. 2605.22717 null
2026-05-21 Abstraction for Offline Goal-Conditioned Reinforcement Learning Clarisse Wibault et.al. 2605.22711 null
2026-05-21 Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals Shuo Yang et.al. 2605.22703 null
2026-05-21 SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation Zhenyu Lu et.al. 2605.22658 null
2026-05-21 Directed extended-range percolation Wenbo Liu et.al. 2605.22646 null
2026-05-21 Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning Banghao Chi et.al. 2605.22642 null
2026-05-21 Whole-Blood Boundary Analysis of BioFET-Based ctDNA Detection for Intravascular Sensing in Intrabody Nanonetworks Ida Kleger-Rudomin et.al. 2605.22637 null
2026-05-21 A note on convergence of Wasserstein policy optimization David Šiška et.al. 2605.22622 null
2026-05-21 Two is better than one: A Collapse-free Multi-Reward RLIF Training Framework Shourov Joarder et.al. 2605.22620 null
2026-05-21 Upscaling DFT-trained machine-learning interatomic potential toward Quantum Monte Carlo accuracy: Sulfur-vacancy migration in monolayer MoS $_2$ as a testbed Adam Hložný et.al. 2605.22601 null
2026-05-21 LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance Yuchun Fan et.al. 2605.22567 null
2026-05-21 F-TIS: Harnessing Diverse Models in Collaborative GRPO Nikolay Blagoev et.al. 2605.22537 null
2026-05-21 Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning Zihan Liang et.al. 2605.22511 null
2026-05-20 Variance Reduction for Expectations with Diffusion Teachers Jesse Bettencourt et.al. 2605.21489 null
2026-05-20 Agent JIT Compilation for Latency-Optimizing Web Agent Planning and Scheduling Caleb Winston et.al. 2605.21470 null
2026-05-20 You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories Zhepei Wei et.al. 2605.21468 null
2026-05-20 DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Kaiyi Zhang et.al. 2605.21467 null
2026-05-20 Mem- $π$ : Adaptive Memory through Learning When and What to Generate Xiaoqiang Wang et.al. 2605.21463 null
2026-05-20 roto 2.0: The Robot Tactile Olympiad Elle Miller et.al. 2605.21429 null
2026-05-20 FedCritic: Serverless Federated Critic Learning-based Resource Allocation for Multi-Cell OFDMA in 6G Amin Farajzadeh et.al. 2605.21418 null
2026-05-20 Validating Navmesh using Geometry: Voxel-Based Analysis with Prioritized Exploration Ramesh Raghavan et.al. 2605.21397 null
2026-05-20 Learning Robust Dexterous In-Hand Manipulation from Joint Sensors with Proprioceptive Transformer Senlan Yao et.al. 2605.21330 null
2026-05-20 Smart strategies to navigate turbulent odor plumes reorienting to local wind Lorenzo Piro et.al. 2605.21329 null
2026-05-20 DeCoR: Design and Control Co-Optimization for Urban Streets Using Reinforcement Learning Bibek Poudel et.al. 2605.21311 null
2026-05-20 TimeSRL: Generalizable Time-Series Behavioral Modeling via Semantic RL-Tuned LLMs – A Case Study in Mental Health Yuang Fan et.al. 2605.21295 null
2026-05-20 \textit{Stochastic} MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent Zeyuan Wang et.al. 2605.21282 null
2026-05-20 DriveMA: Rethinking Language Interfaces in Driving VLAs with One-Step Meta-Actions Weicheng Zheng et.al. 2605.21273 null
2026-05-20 How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR Richa Verma et.al. 2605.21266 null
2026-05-20 Reinforcement Learning for Risk Adaptation via Differentiable CVaR Barrier Functions Xinyi Wang et.al. 2605.21257 null
2026-05-20 Random Matrix Spectra from Boltzmann-Weighted Lattice Ensembles Yaprak Önder et.al. 2605.21254 null
2026-05-20 LamPO: A Lambda Style Policy Optimization for Reasoning Language Models Zhe Yuan et.al. 2605.21235 null
2026-05-20 PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment Richa Verma et.al. 2605.21225 null
2026-05-20 Behavior-Consistent Deep Reinforcement Learning Marcel Hussing et.al. 2605.21214 null
2026-05-19 Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR Utkarsh Tyagi et.al. 2605.20164 null
2026-05-19 Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP Jann Pfeifer et.al. 2605.20066 null
2026-05-19 Non-equilibrium quantum dynamics of interacting integrable models by Monte Carlo sampling Lehmann representations Riccardo Senese et.al. 2605.20065 null
2026-05-19 Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Wenjie Tang et.al. 2605.20061 null
2026-05-19 Gaussian Process Eigenmodes for Statistical and Systematic Uncertainties in Template Fits Vincent Alexander Croft et.al. 2605.20048 null
2026-05-19 When Critics Disagree: Adaptive Reward Poisoning Attacks in RIS-Aided Wireless Control System Deemah H. Tashman et.al. 2605.20037 null
2026-05-19 GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards Kyeongjin Ahn et.al. 2605.20006 null
2026-05-19 CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition Hongji Yang et.al. 2605.19995 null
2026-05-19 Error Bounds for Importance Sampling with Estimated Proposal Distributions Cathrine Aeckerle-Willems et.al. 2605.19989 null
2026-05-19 A conceptual framework for learning to listen by reward: Curiosity-driven search for novel sources Andreas Triantafyllopoulos et.al. 2605.19984 null
2026-05-19 Safe Deep Reinforcement Learning for Spacecraft Reorientation with Pointing Keep-Out Constraint Juntang Yang et.al. 2605.19967 null
2026-05-19 Variance-Reduced Manifold Sampling via Polynomial-Maximization Density Estimation Serhii Zabolotnii et.al. 2605.19938 null
2026-05-19 JAXenstein: Accelerated Benchmarking for First-Person Environments Ruo Yu Tao et.al. 2605.19926 null
2026-05-19 RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations Shuoqin Zhang et.al. 2605.19924 null
2026-05-19 Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement Learning Dongjie Yu et.al. 2605.19919 null
2026-05-19 Fair-Aurora: Comparing Fairness Strategies for Reinforcement Learning-Based Congestion Control in Multi-Flow Environments Thomas Mbrice et.al. 2605.19909 null
2026-05-19 Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning Qinghe Ma et.al. 2605.19852 null
2026-05-19 Domain-wall Quintessence Nobufusa Kobayashi et.al. 2605.19841 null
2026-05-19 Radiative depolarization of high-energy electron beams in wakefield accelerators Oliver Mathiak et.al. 2605.19814 null
2026-05-19 Reliable model selection in the presence of parameter non-identifiability Yong See Foo et.al. 2605.19807 null
2026-05-18 General Preference Reinforcement Learning Muhammad Umer et.al. 2605.18721 null
2026-05-18 SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training Komal Kumar et.al. 2605.18719 null
2026-05-18 EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL Minrui Xu et.al. 2605.18703 null
2026-05-18 COOPO: Cyclic Offline-Online Policy Optimization Algorithm Qisai Liu et.al. 2605.18675 null
2026-05-18 Leveraging Latent Visual Reasoning in Silence Dongyao Zhu et.al. 2605.18641 null
2026-05-18 Emergent Thiemann coherent states in the near-kernel sector of quantum reduced loop gravity Ilkka Mäkinen et.al. 2605.18625 null
2026-05-18 ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics Ziyu Wei et.al. 2605.18617 null
2026-05-18 Unified Walking, Running, and Recovery for Humanoids via State-Dependent Adversarial Motion Priors Yidan Lu et.al. 2605.18611 null
2026-05-18 Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Yuhuan Wu et.al. 2605.18603 null
2026-05-18 AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning Peilin Wu et.al. 2605.18592 null
2026-05-18 Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation Mingfei Sun et.al. 2605.18591 null
2026-05-18 Markov chain Monte Carlo (MCMC) based Likelihood Extraction of Chiral-Odd Compton Form Factors from Deeply Virtual Exclusive Experiments Saraswati Pandey et.al. 2605.18589 null
2026-05-18 Reinforcement Learning Assisted Quantum Simulation of Many-Body Excited States and Real-Time Dynamics Jiaji Zhang et.al. 2605.18569 null
2026-05-18 HJ-Gauss: A Monte-Carlo HJ Reachability Scheme Lekan Molu et.al. 2605.18566 null
2026-05-18 STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics Tingfeng Hui et.al. 2605.18548 null
2026-05-18 Bilayer crystals in a polar-molecules system Vinicius Zampronio et.al. 2605.18546 null
2026-05-18 Discovering Data Encoding Strategies for Quantum-Classical Neural Networks Using Monte Carlo Tree Search Lena Tokuhiro et.al. 2605.18540 null
2026-05-18 AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment Zhenlin Wei et.al. 2605.18529 null
2026-05-18 Offline Contextual Bandits in the Presence of New Actions Ren Kishimoto et.al. 2605.18509 null
2026-05-18 DiPRL: Learning Discrete Programmatic Policies via Architecture Entropy Regularization Chengpeng Hu et.al. 2605.18508 null
2026-05-18 Properties of the quantum vacuum in non-abelian gauge theories Fernando Ezquerro et.al. 2605.18220 null
2026-05-18 Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation Guining Cao et.al. 2605.18191 null
2026-05-18 The Dynamics of Policy Gradient in Social Dilemmas with Partner Selection Benedict Russell et.al. 2605.18185 null
2026-05-18 Testing the reliability of magnetic field strength measurements for M dwarfs I. Amateis et.al. 2605.18151 null
2026-05-18 Taming the 3D Wilson-Fisher Fixed Point via Nonlocal Effective Action Hyeon Jung Kim et.al. 2605.18148 null
2026-05-18 Equilibrium Selection in Multi-Agent Policy Gradients via Opponent-Aware Basin Entry Yevhen Shcherbinin et.al. 2605.18078 null
2026-05-18 LLM-Guided Communication for Cooperative Multi-Agent Reinforcement Learning Sangjun Bae et.al. 2605.18077 null
2026-05-18 Chemo-mechanical coupling stabilizes mixed $\mathrm{Ag}{x}\mathrm{Cu}{1-x}\mathrm{GaSe}_{2}$ solar-cell absorbers: Insights from Monte-Carlo simulations assisted by ab initio informed machine-learning potentials Vasilios Karanikolas et.al. 2605.18057 null
2026-05-18 Interaction-Breaking Adversarial Learning Framework for Robust Multi-Agent Reinforcement Learning Sunwoo Lee et.al. 2605.18024 null
2026-05-18 Uncertainty Reliability Under Domain Shift: An Investigation for Data-Driven Blood Pressure Estimation in Photoplethysmography Mohammad Moulaeifard et.al. 2605.18008 null
2026-05-18 RL4RLA: Teaching ML to Discover Randomized Linear Algebra Algorithms Through Curriculum Design and Graph-Based Search Jinglong Xiong et.al. 2605.18004 null
2026-05-18 AutoVecCoder: Teaching LLMs to Generate Explicitly Vectorized Code Shangzhan Li et.al. 2605.17978 null
2026-05-18 Generation Navigator: A State-Aware Agentic Framework for Image Generation Jinming Liu et.al. 2605.17969 null
2026-05-18 Function graph transformers universally approximate operators between function spaces Takashi Furuya et.al. 2605.17968 null
2026-05-18 Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning Zhanyue Qin et.al. 2605.17958 null
2026-05-18 AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents Pan Wang et.al. 2605.17933 null
2026-05-18 Transfer Learning for Customized Car Racing Environments Benedict Florance Arockiaraj et.al. 2605.17928 null
2026-05-18 An Efficient Streaming Video Understanding Framework with Agentic Control Jinming Liu et.al. 2605.17921 null
2026-05-18 Assessing the Impact of Source Confusion for GREX-PLUS based on Deep JWST NIRCam Imaging Yoshiaki Ono et.al. 2605.17882 null
2026-05-18 PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization Wonjoong Kim et.al. 2605.17877 null
2026-05-14 RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO Yanzuo Lu et.al. 2605.15190 null
2026-05-14 Hand-in-the-Loop: Improving Dexterous VLA via Seamless Interventional Correction Zhuohang Li et.al. 2605.15157 null
2026-05-14 Self-Distilled Agentic Reinforcement Learning Zhengxi Lu et.al. 2605.15155 null
2026-05-14 Downlink Performance Analysis of Pinching Antenna Systems: WDMA or NOMA? Han Zhang et.al. 2605.15129 null
2026-05-14 Identification and Estimation of Staggered Difference-in-Differences with Network Spillovers Hayato Tagawa et.al. 2605.15119 null
2026-05-14 Learning from Language Feedback via Variational Policy Distillation Yang Li et.al. 2605.15113 null
2026-05-14 Two Protons, Two Positrons, and Four Electrons: Covalent Bond with van der Waals Characteristics Jorge Charry et.al. 2605.15099 null
2026-05-14 Computational Imaging Priors for Wireless Capsule Endoscopy: Monte Carlo-Guided Hemoglobin Mapping for Rare-Anomaly Detection Chengshuai Yang et.al. 2605.15062 null
2026-05-14 DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models Quanhao Li et.al. 2605.15055 null
2026-05-14 Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use Renning Pang et.al. 2605.15041 null
2026-05-14 Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance Kai Yan et.al. 2605.15012 null
2026-05-14 A Monte Carlo positronium decay source model with multiple annihilation channels in GATE Wojciech Krzemien et.al. 2605.14987 null
2026-05-14 Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition Sanjeev Manivannan et.al. 2605.14982 null
2026-05-14 Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowing Jie Jiang et.al. 2605.14978 null
2026-05-14 Multi-regime Markov-switching models with time-varying transition probabilities: An application to U.S. Treasury yields Samuel Modée et.al. 2605.14976 null
2026-05-14 Analyzing the two-dimensional doped Hubbard model with the Worldvolume HMC method Masafumi Fukuma et.al. 2605.14965 null
2026-05-14 Quantum-Secure Physical Unclonable Function enabled by Silicon Photonics Integrated Circuits G. Sarantoglou et.al. 2605.14959 null
2026-05-14 A CUBS-Compatible Ultrasound Morphology and Uncertainty-Aware Baseline for Carotid Intima-Media Segmentation and Preliminary Risk Prediction Aueaphum Aueawatthanaphisut et.al. 2605.14949 null
2026-05-14 Piece-wise linear isotonic regression Timo Kuosmanen et.al. 2605.14943 null
2026-05-14 Not All Symbols Are Equal: Importance-Aware Constellation Design for Semantic Communication Albert Shaju et.al. 2605.14940 null
2026-05-13 Quantitative Linear Logic for Neuro-Symbolic Learning and Verification Thomas Flinkow et.al. 2605.13845 null
2026-05-13 Parallel Scan Recurrent Neural Quantum States for Scalable Variational Monte Carlo Ejaaz Merali et.al. 2605.13807 null
2026-05-13 EvoGround: Self-Evolving Video Agents for Video Temporal Grounding Minjoon Jung et.al. 2605.13803 null
2026-05-13 Uniqueness of synchronized stationary equilibria in the Kuramoto mean field game Sebastian Munoz et.al. 2605.13783 null
2026-05-13 Application of exhaustive simulation flow for advanced performance prediction of monolithic active pixel sensors E. Sacchetti et.al. 2605.13760 null
2026-05-13 Do Hopfield Networks Dream of Stored Patterns? A Statistical-Mechanical Theory of Dreaming in Multidirectional Associative Memories Adriano Barra et.al. 2605.13721 null
2026-05-13 Tight Sample Complexity Bounds for Entropic Best Policy Identification Amer Essakine et.al. 2605.13717 null
2026-05-13 Scale-Sensitive Shattering: Learnability and Evaluability at Optimal Scale Shashaank Aiyer et.al. 2605.13684 null
2026-05-13 SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models Vladislav Makarov et.al. 2605.13667 null
2026-05-13 Robot Squid Game: Quadrupedal Locomotion for Traversing Narrow Tunnels Amir Hossain Raj et.al. 2605.13665 null
2026-05-13 Air-Sea Surface Modeling and Operating Link Range Evaluation for AUV-to-UAV Optical Wireless Communication Links Ikenna Chinazaekpere Ijeh et.al. 2605.13661 null
2026-05-13 Efficient simulation of chemical reaction in DSMC Hong Deng et.al. 2605.13653 null
2026-05-13 Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization Yang Bai et.al. 2605.13641 null
2026-05-13 Achieving $ε^{-2}$ Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions Ishaq Hamza et.al. 2605.13639 null
2026-05-13 CO-MAP: A Reinforcement Learning Approach to the Qubit Allocation Problem Ankit Kulshrestha et.al. 2605.13638 null
2026-05-13 A Majorization-Minimization with Monte Carlo Approach for Hyperparameter Estimation Elle Buser et.al. 2605.13620 null
2026-05-13 Adaptive time-domain simulation of optical cavities with arbitrary dynamics A. Svizzeretto et.al. 2605.13599 null
2026-05-13 On the Apparent Correlation between X-ray and Neutrino Luminosities of Active Galactic Nuclei Jian-Jun Luo et.al. 2605.13588 null
2026-05-13 Learning Local Constraints for Reinforcement-Learned Content Generators Debosmita Bhaumik et.al. 2605.13570 null
2026-05-13 Phase Ordering in a few O(n) Symmetric Models: Slow Growth, Mpemba Effect and Experimental Relevance Wasim Akram et.al. 2605.13564 null
2026-05-12 Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training Yuanda Xu et.al. 2605.12483 null
2026-05-12 OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation Guohui Zhang et.al. 2605.12480 null
2026-05-12 Reward Hacking in Rubric-Based Reinforcement Learning Anas Mahmoud et.al. 2605.12474 null
2026-05-12 Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs Jose E. Aguilar Escamilla et.al. 2605.12462 null
2026-05-12 Strongly Integrable Operator-Valued Functions, Generated Vector Measures and Compactness of Integrals Miloš Arsenović et.al. 2605.12454 null
2026-05-12 LychSim: A Controllable and Interactive Simulation Framework for Vision Research Wufei Ma et.al. 2605.12449 null
2026-05-12 ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models Chen Li et.al. 2605.12446 null
2026-05-12 Basilisk and Docker for Reproducible GN&C Simulation: A Workflow Reference Anubhav Gupta et.al. 2605.12443 null
2026-05-12 Learning Minimally Rigid Graphs with High Realization Counts Oleksandr Slyvka et.al. 2605.12427 null
2026-05-12 Aligning Flow Map Policies with Optimal Q-Guidance Christos Ziakas et.al. 2605.12416 null
2026-05-12 Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images Yuangong Chen et.al. 2605.12413 null
2026-05-12 Model-based Bootstrap of Controlled Markov Chains Ziwei Su et.al. 2605.12410 null
2026-05-12 Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems William Parris et.al. 2605.12406 null
2026-05-12 Events as Triggers for Behavioral Diversity in Multi-Agent Reinforcement Learning Hannes Büchi et.al. 2605.12388 null
2026-05-12 Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training Rasool Fakoor et.al. 2605.12380 null
2026-05-12 Discrete Flow Matching for Offline-to-Online Reinforcement Learning Fairoz Nower Khan et.al. 2605.12379 null
2026-05-12 QAP-Router: Tackling Qubit Routing as Dynamic Quadratic Assignment with Reinforcement Learning Kien X. Nguyen et.al. 2605.12365 null
2026-05-12 BSO: Safety Alignment Is Density Ratio Matching Tien-Phat Nguyen et.al. 2605.12339 null
2026-05-12 Reinforcing VLAs in Task-Agnostic World Models Yucen Wang et.al. 2605.12334 null
2026-05-12 The wave nature of a Mott insulator Xudong Yu et.al. 2605.12322 null
2026-05-11 Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Haoyuan Sun et.al. 2605.10937 null
2026-05-11 Variational Inference for Lévy Process-Driven SDEs via Neural Tilting Yaman Kindap et.al. 2605.10934 null
2026-05-11 Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning Junhao Shen et.al. 2605.10923 null
2026-05-11 gemlib.mcmc: composable kernels for Metropolis-within-Gibbs sampling schemes Alin Morariu et.al. 2605.10914 null
2026-05-11 Equivariant Reinforcement Learning for Clifford Quantum Circuit Synthesis Richie Yeung et.al. 2605.10910 null
2026-05-11 Revisiting Policy Gradients for Restricted Policy Classes: Escaping Myopic Local Optima with $k$ -step Policy Gradients Alex DeWeese et.al. 2605.10909 null
2026-05-11 RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Gaotang Li et.al. 2605.10899 null
2026-05-11 BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD Haozhe Zhang et.al. 2605.10865 null
2026-05-11 Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents Shijue Huang et.al. 2605.10832 null
2026-05-11 Unified Noise Steering for Efficient Human-Guided VLA Adaptation Junjie Lu et.al. 2605.10821 null
2026-05-11 Policy Gradient Methods for Non-Markovian Reinforcement Learning Avik Kar et.al. 2605.10816 null
2026-05-11 Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Daniel Ranard et.al. 2605.10810 null
2026-05-11 New AI-Driven Tools for Enhancing Campus Well-being: A Prevention and Intervention Approach Jinwen Tang et.al. 2605.10804 null
2026-05-11 RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology Wenxuan Li et.al. 2605.10761 null
2026-05-11 XQCfD: Accelerating Fast Actor-Critic Algorithms with Prior Data and Prior Policies Daniel Palenicek et.al. 2605.10734 null
2026-05-11 What should post-training optimize? A test-time scaling law perspective Muheng Li et.al. 2605.10716 null
2026-05-11 xApp Empowered Resource Management for Non-Terrestrial Users in 5G O-RAN Networks Mohammed M. H. Qazzaz et.al. 2605.10704 null
2026-05-11 Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework Phalguni Nanda et.al. 2605.10671 null
2026-05-11 Micro-environment of the Eu interstitial in $β$-SiAlON:Eu$^{2+}$ green phosphor Julien Bouquiaux et.al. 2605.10665 null
2026-05-11 Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents Zhiyuan Fan et.al. 2605.10663 null
2026-05-11 Partial annealing and pattern decorrelation in associative neural networks Linda Albanese et.al. 2605.10304 null
2026-05-11 Robust Probabilistic Shielding for Safe Offline Reinforcement Learning Maris F. L. Galesloot et.al. 2605.10293 null
2026-05-11 MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading Baibei Ji et.al. 2605.10268 null
2026-05-11 Towards Autonomous Railway Operations: A Semi-Hierarchical Deep Reinforcement Learning Approach to the Vehicle Rescheduling Problem Alberto Castagna et.al. 2605.10257 null
2026-05-11 Floquet-tuned superfluid-checkerboard competition in dipolar bosons Jin Yang et.al. 2605.10254 null
2026-05-11 When Does Non-Uniform Replay Matter in Reinforcement Learning? Michal Korniak et.al. 2605.10236 null
2026-05-11 Relative Score Policy Optimization for Diffusion Language Models Zichao Yu et.al. 2605.10218 null
2026-05-11 TRACE: Distilling Where It Matters via Token-Routed Self On-Policy Alignment Jiaxuan Wang et.al. 2605.10194 null
2026-05-11 MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning Guangli Chen et.al. 2605.10177 null
2026-05-11 Balancing Efficiency and Fairness in Traffic Light Control through Deep Reinforcement Learning Matteo Cederle et.al. 2605.10170 null
2026-05-11 Data-Asymmetric Latent Imagination and Reranking for 3D Robotic Imitation Learning Lianghao Luo et.al. 2605.10166 null
2026-05-11 Unsupervised Process Reward Models Artyom Gadetsky et.al. 2605.10158 null
2026-05-11 Is DRL-based MAC Ready for Underwater Acoustic Networks? Exploring Its Practicality in Real Field Experiments Jiani Guo et.al. 2605.10144 null
2026-05-11 FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models Zeynel A. Uluşan et.al. 2605.10141 null
2026-05-11 Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation Zhixuan Shen et.al. 2605.10118 null
2026-05-11 Arcane: An Assertion Reduction Framework through Semantic Clustering and MCTS-Guided Rule Exploring Hongqin Lyu et.al. 2605.10107 null
2026-05-11 EFGCL: Learning Dynamic Motion through Spotting-Inspired External Force Guided Curriculum Learning Keita Yoneda et.al. 2605.10063 null
2026-05-11 Guided Streaming Stochastic Interpolant Policy Puming Jiang et.al. 2605.10051 null
2026-05-11 Adaptive Action Chunking via Multi-Chunk Q Value Estimation Yongjae Shin et.al. 2605.10044 null
2026-05-11 Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving Aron Distelzweig et.al. 2605.10034 null
2026-05-08 123D: Unifying Multi-Modal Autonomous Driving Data at Scale Daniel Dauner et.al. 2605.08084 null
2026-05-08 Chase-like Decoding: Test Pattern Design and Performance Analysis Tim Janz et.al. 2605.08081 null
2026-05-08 Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning Manish Bhattarai et.al. 2605.08061 null
2026-05-08 Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs Gugan Thoppe et.al. 2605.08053 null
2026-05-08 Joint Beamforming and Antenna Placement Optimization in Pinching Antenna Systems with User Mobility: A Deep Reinforcement Learning Approach Ali Amhaz et.al. 2605.08039 null
2026-05-08 Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph Ning Liu et.al. 2605.08037 null
2026-05-08 Active Embodiment Identification with Reinforcement Learning for Legged Robots Nico Bohlinger et.al. 2605.08020 null
2026-05-08 Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners Botos Csaba et.al. 2605.08019 null
2026-05-08 Learning CLI Agents with Structured Action Credit under Selective Observation Haoyang Su et.al. 2605.08013 null
2026-05-08 Interpreting Reinforcement Learning Agents with Susceptibilities Chris Elliott et.al. 2605.08007 null
2026-05-08 Bayesian Sensitivity of Causal Inference Estimators under Evidence-Based Priors Nikita Dhawan et.al. 2605.07993 null
2026-05-08 Uncertainty Quantification for Cardiac Shape Reconstruction with Deep Signed Distance Functions via MCMC methods Jan Verhülsdonk et.al. 2605.07987 null
2026-05-08 TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning Giacomo Spigler et.al. 2605.07943 null
2026-05-08 Accelerating Langevin Monte Carlo via Efficient Stochastic Runge–Kutta Methods beyond Log-Concavity Bin Yang et.al. 2605.07939 null
2026-05-08 Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models Yuancheng Wei et.al. 2605.07872 null
2026-05-08 Systematic frequency-collision analysis of the cross-resonance gate outside the straddling regime Shinichi Inoue et.al. 2605.07868 null
2026-05-08 Cluster Dynamics Stay Fast-Until Tricriticality Minjun Jeon et.al. 2605.07867 null
2026-05-08 KL for a KL: On-Policy Distillation with Control Variate Baseline Minjae Oh et.al. 2605.07865 null
2026-05-08 From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data Yue Yu et.al. 2605.07861 null
2026-05-08 Actor-Critic Algorithm for Dynamic Expectile and CVaR Yudong Luo et.al. 2605.07857 null
2026-05-07 Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Mingwei Xu et.al. 2605.06650 null
2026-05-07 StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction Xiangyuan Xue et.al. 2605.06642 null
2026-05-07 Recursive Agent Optimization Apurva Gandhi et.al. 2605.06639 null
2026-05-07 Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key Tianle Wang et.al. 2605.06638 null
2026-05-07 Cross-Modal Navigation with Multi-Agent Reinforcement Learning Shuo Liu et.al. 2605.06595 null
2026-05-07 ReActor: Reinforcement Learning for Physics-Aware Motion Retargeting David Müller et.al. 2605.06593 null
2026-05-07 SNAPO: Smooth Neural Adjoint Policy Optimization for Optimal Control via Differentiable Simulation Dmitri Goloubentsev et.al. 2605.06570 null
2026-05-07 Dynamic Treatment on Networks Bengusu Nar et.al. 2605.06564 null
2026-05-07 Criticality and Saturation in Orthogonal Neural Networks Max Guillen et.al. 2605.06563 null
2026-05-07 Coordination Matters: Evaluation of Cooperative Multi-Agent Reinforcement Learning Maria Ana Cardei et.al. 2605.06557 null
2026-05-07 Sequential Design of Genetic Circuits Under Uncertainty With Reinforcement Learning Michal Kobiela et.al. 2605.06552 null
2026-05-07 Affine Subcode Ensemble Decoding for Degeneracy-Aware Quantum Error Correction Leo Wursthorn et.al. 2605.06547 null
2026-05-07 Delay-Robust Deep Reinforcement Learning for Ranging-Free Channel Access under Mobility in Underwater Acoustic Networks Huaisheng Ye et.al. 2605.06536 null
2026-05-07 ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL Wei Gao et.al. 2605.06534 null
2026-05-07 On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR Hao Ye et.al. 2605.06523 null
2026-05-07 Learning to Cut: Reinforcement Learning for Benders Decomposition Haochen Cai et.al. 2605.06516 null
2026-05-07 MARBLE: Multi-Aspect Reward Balance for Diffusion RL Canyu Zhao et.al. 2605.06507 null
2026-05-07 Operator-Guided Invariance Learning for Continuous Reinforcement Learning Zuyuan Zhang et.al. 2605.06500 null
2026-05-07 Scaling the Queue: Reinforcement Learning for Equitable Call Classification Capacity in NYC Municipal Complaint Systems Irene Aldridge et.al. 2605.06482 null
2026-05-07 Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching Xiang Li et.al. 2605.06474 null
2026-05-06 OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents Shuang Chen et.al. 2605.05185 null
2026-05-06 Estimating the expected output of wide random MLPs more efficiently than sampling Wilson Wu et.al. 2605.05179 null
2026-05-06 When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Lakshita Dodeja et.al. 2605.05172 null
2026-05-06 Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning Alper Kamil Bozkurt et.al. 2605.05123 null
2026-05-06 Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime Tianshu Zhu et.al. 2605.05112 null
2026-05-06 LineRides: Line-Guided Reinforcement Learning for Bicycle Robot Stunts Seungeun Rho et.al. 2605.05110 null
2026-05-06 Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning Harin Lee et.al. 2605.05102 null
2026-05-06 Dynamic Collateral Control for Permissionless Spot Perpetual Basis Trading Anatoly Krestenko et.al. 2605.05089 null
2026-05-06 Provable imitation learning for control of instability in partially-observed Vlasov–Poisson equations Xiaofan Xia et.al. 2605.05081 null
2026-05-06 Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization Xin Yu et.al. 2605.05040 null
2026-05-06 Study of Particle Fluence Effects on Collected Charge and Depletion Voltage of the ATLAS IBL Planar Pixel Sensors ATLAS Collaboration et.al. 2605.05030 null
2026-05-06 Graph-SND: Sparse Aggregation for Behavioral Diversity in Multi-Agent Reinforcement Learning Shawn Ray et.al. 2605.05020 null
2026-05-06 Misaligned by Reward: Socially Undesirable Preferences in LLMs Gayane Ghazaryan et.al. 2605.05003 null
2026-05-06 DualTCN: A Physics-Constrained Temporal Convolutional Network for 2 Time-Domain Marine CSEM Inversion Khaled Ahmed et.al. 2605.04997 null
2026-05-06 EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance Song Yu et.al. 2605.04960 null
2026-05-06 Modular Reinforcement Learning For Cooperative Swarms Erel Shtossel et.al. 2605.04939 null
2026-05-06 A Convolution Process for Sea Surface Temperature Hot-Spot Identification in the Mediterranean Sea Leonardo Marchesin et.al. 2605.04921 null
2026-05-06 Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization Xiyan Fu et.al. 2605.04920 null
2026-05-06 Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games Yidong He et.al. 2605.04906 null
2026-05-06 A Hierarchical Agent System with Reinforcement Learning for Multivariate Time Series Data Cleaning Yuhan Shi et.al. 2605.04902 null
2026-05-05 Sequential vs. Simultaneous Entanglement Swapping under Optimal Link-Layer Control Priyam Srivastava et.al. 2605.04047 null
2026-05-05 OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories Yuwen Du et.al. 2605.04036 null
2026-05-05 Mitigating False Positives in Static Memory Safety Analysis of Rust Programs via Reinforcement Learning P Akilesh et.al. 2605.04000 null
2026-05-05 Selecting optimal unrestricted Hartree-Fock trial wavefunctions for phaseless auxiliary-field quantum Monte Carlo: Accuracy and limitations in modeling three iron-sulfur clusters Don Danilov et.al. 2605.03981 null
2026-05-05 Magic-Informed Quantum Architecture Search Vincenzo Lipardi et.al. 2605.03932 null
2026-05-05 Measurement of the neutron shielding efficacy of magnetite for Proton Therapy Facilities and other applications Kijun Park et.al. 2605.03931 null
2026-05-05 Time-dependent variational Monte Carlo without bias Wladislaw Krinitsin et.al. 2605.03930 null
2026-05-05 More Permutations Do Not Always Increase Power: Non-monotonicity in Monte Carlo Permutation Tests Suman Cha et.al. 2605.03886 null
2026-05-05 Fiscal Aggregation and the Limits of IS–LM–BP: Derivations, Aggregation Bias and Reproducible Adversarial Simulations Ricardo Alonzo Fernandez Salguero et.al. 2605.03881 null
2026-05-05 EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics Shuyue Stella Li et.al. 2605.03871 null
2026-05-05 Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards Tianyang Han et.al. 2605.03862 null
2026-05-05 Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation Bin Wu et.al. 2605.03849 null
2026-05-05 Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligenc Munkhdegerekh Batzorig et.al. 2605.03847 null
2026-05-05 SigLoMa: Learning Open-World Quadrupedal Loco-Manipulation from Ego-Centric Vision Shiyi Chen et.al. 2605.03846 null
2026-05-05 SOAR: Real-Time Joint Optimization of Order Allocation and Robot Scheduling in Robotic Mobile Fulfillment Systems Yibang Tang et.al. 2605.03842 null
2026-05-05 RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models Hao Wu et.al. 2605.03821 null
2026-05-05 Towards accurate extreme event likelihoods from diffusion model climate emulators Peter Manshausen et.al. 2605.03802 null
2026-05-05 Natural Language Processing: A Comprehensive Practical Guide from Tokenisation to RLHF Mullosharaf K. Arabov et.al. 2605.03799 null
2026-05-05 What You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic Curiosity Haoxi Li et.al. 2605.03782 null
2026-05-05 Spontaneous Topological Locking and Symmetry Restoration of Meron Lattices in Synthetic Antiferromagnets Gülşen Doğan et.al. 2605.03755 null
2026-05-01 SAVGO: Learning State-Action Value Geometry with Cosine Similarity for Continuous Control Stavros Orfanoudakis et.al. 2605.00787 null
2026-05-01 Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values Shradha Sharma et.al. 2605.00762 null
2026-05-01 Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Indraneil Paul et.al. 2605.00754 null
2026-05-01 NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search Sizhe Tang et.al. 2605.00751 null
2026-05-01 Decentralized Proximal Stochastic Gradient Langevin Dynamics Mohammad Rafiqul Islam et.al. 2605.00723 null
2026-05-01 Beyond Bragg-Mirrors for Gravitational Wave Telescopes: A Fabrication Tolerant Hybrid Metasurface-Bragg Mirror Design Christian Kranhold et.al. 2605.00714 null
2026-05-01 Bootstrap Inference under General Two-way Clustering with Serially and Spatially Dependent Common Effects Ulrich Hounyo et.al. 2605.00709 null
2026-05-01 Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory Derong Xu et.al. 2605.00702 null
2026-05-01 STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack Xutao Mao et.al. 2605.00699 null
2026-05-01 Optimal Merton’s Problem under Multivariate Affine Volterra Models with Jumps Sigui Brice Dro et.al. 2605.00688 null
2026-05-01 Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning Jiaming Zhang et.al. 2605.00667 null
2026-05-01 Stability of parton distributions at high $x$ : impact of nuclear and power corrections C. Cocuzza et.al. 2605.00666 null
2026-05-01 Spectral Duality and Reset-Neutral Distributions in Random Walks with Multi-Site Geometric Resetting Juan Antonio Vega Coso et.al. 2605.00657 null
2026-05-01 Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation Andrzej Ruszczynski et.al. 2605.00654 null
2026-05-01 Learning Multimodal Energy-Based Model with Multimodal Variational Auto-Encoder via MCMC Revision Jiali Cui et.al. 2605.00644 null
2026-05-01 Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding Yan Zhang et.al. 2605.00642 null
2026-05-01 Monte Carlo study of the superfluid phase of $^4$ He Massimo Boninsegni et.al. 2605.00629 null
2026-05-01 Recovering Hidden Reward in Diffusion-Based Policies Yanbiao Ji et.al. 2605.00623 null
2026-05-01 Dynamic Linear Panel Regression Models with Interactive Fixed Effects Hyungsik Roger Moon et.al. 2605.00612 null
2026-05-01 Estimation of random coefficients logit demand models with interactive fixed effects Hyungsik Roger Moon et.al. 2605.00602 null
2026-04-30 LaST-R1: Reinforcing Action via Adaptive Physical Latent Reasoning for VLA Models Hao Chen et.al. 2604.28192 null
2026-04-30 Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Keming Wu et.al. 2604.28185 null
2026-04-30 Exploration Hacking: Can LLMs Learn to Resist RL Training? Eyon Jang et.al. 2604.28182 null
2026-04-30 Synthetic Computers at Scale for Long-Horizon Productivity Simulation Tao Ge et.al. 2604.28181 null
2026-04-30 Global Optimality for Constrained Exploration via Penalty Regularization Florian Wolf et.al. 2604.28144 null
2026-04-30 PRISM: Pre-alignment via Black-box On-policy Distillation for Multimodal Reinforcement Learning Sudong Wang et.al. 2604.28123 null
2026-04-30 GSDrive: Reinforcing Driving Policies by Multi-mode Trajectory Probing with 3D Gaussian Splatting Environment Ziang Guo et.al. 2604.28111 null
2026-04-30 Neural Aided Kalman Filtering for UAV State Estimation in Degraded Sensing Environments Akhil Gupta et.al. 2604.28107 null
2026-04-30 FiLMMeD: Feature-wise Linear Modulation for Cross-Problem Multi-Depot Vehicle Routing Arthur Corrêa et.al. 2604.28102 null
2026-04-30 Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles Zainab Rehan et.al. 2604.28087 null
2026-04-30 Intelligent Self-tuning Active EMI Filtering for Electrified Automotive Power Systems Using Reinforcement Learning Mahuizi Lu et.al. 2604.28084 null
2026-04-30 AesRM: Improving Video Aesthetics with Expert-Level Feedback Yujin Han et.al. 2604.28078 null
2026-04-30 RHyVE: Competence-Aware Verification and Phase-Aware Deployment for LLM-Generated Reward Hypotheses Feiyu Wu et.al. 2604.28056 null
2026-04-30 Exponential families from a single KL identity Marc Dymetman et.al. 2604.28036 null
2026-04-30 Topological Susceptibility and QCD at Finite Theta Angle Claudio Bonanno et.al. 2604.28035 null
2026-04-30 Cost-Aware Learning Clara Mohri et.al. 2604.28020 null
2026-04-30 Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation Jing Zhang et.al. 2604.28011 null
2026-04-30 Learning from Disagreement: Clinician Overrides as Implicit Preference Signals for Clinical AI in Value-Based Care Prabhjot Singh et.al. 2604.28010 null
2026-04-30 Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning Shijin Gong et.al. 2604.28005 null
2026-04-30 Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning Jingcheng Deng et.al. 2604.27998 null
2026-04-29 ClawGym: A Scalable Framework for Building Effective Claw Agents Fei Bai et.al. 2604.26904 null
2026-04-29 MLMC-qDRIFT: Multilevel Variance Reduction for Randomized Quantum Hamiltonian Simulation Pegah Mohammadipour et.al. 2604.26865 null
2026-04-29 Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics Bernd Frauenknecht et.al. 2604.26836 null
2026-04-29 Rule-based High-Level Coaching for Goal-Conditioned Reinforcement Learning in Search-and-Rescue UAV Missions Under Limited-Simulation Training Mahya Ramezani et.al. 2604.26833 null
2026-04-29 Finite-Temperature Ferromagnetic Correlations of the Kagome Lattice Hubbard Model Francisco Correia et.al. 2604.26827 null
2026-04-29 GLoop: A Monte Carlo program to construct higher-loop integrals from lower-loop structures Roberto Pittau et.al. 2604.26798 null
2026-04-29 Factorized Latent Reasoning for LLM-based Recommendation Tianqi Gao et.al. 2604.26760 null
2026-04-29 GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents GLM-V Team et.al. 2604.26752 null
2026-04-29 FutureWorld: A Live Environment for Training Predictive Agents with Real-World Outcome Rewards Zhixin Han et.al. 2604.26733 null
2026-04-29 Non-symmetrically $t$ -affine functions revisited Tibor Kiss et.al. 2604.26699 null
2026-04-29 Towards Quantum Optimised Malware Containment Matthew Sutcliffe et.al. 2604.26692 null
2026-04-29 ATLAS: An Annotation Tool for Long-horizon Robotic Action Segmentation Sergej Stanovcic et.al. 2604.26637 null
2026-04-29 SEP Analysis of Quantized SIMO Systems with M-PSK over Correlated Fading Channels Amila Ravinath et.al. 2604.26618 null
2026-04-29 Normalizing flows for density estimation in multi-detector gravitational-wave searches Sam Insley et.al. 2604.26581 null
2026-04-29 PAINT: Partial-Solution Adaptive Interpolated Training for Self-Distilled Reasoners Zhiquan Tan et.al. 2604.26573 null
2026-04-29 Learning to Route Electric Trucks Under Operational Uncertainty Stavros Orfanoudakis et.al. 2604.26566 null
2026-04-29 Random Number Generators in Advanced Optical Experiments: A Comparative Analysis of Semiclassical, Quantum, and Hybrid Architectures Daniil D. Reshetnikov et.al. 2604.26554 null
2026-04-29 Projections for handling uncertainties and enabling domain truncation in diffuse optical tomography Aada Hakula et.al. 2604.26548 null
2026-04-29 Neural and Tensor Networks in the Study of Quantum Annealing Processors Tomasz Śmierzchalski et.al. 2604.26534 null
2026-04-29 Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning Seungyub Han et.al. 2604.26516 null
2026-04-28 How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum Chu-Cheng Lin et.al. 2604.25907 null
2026-04-28 TSN-Affinity: Similarity-Driven Parameter Reuse for Continual Offline Reinforcement Learning Dominik Żurek et.al. 2604.25898 null
2026-04-28 Three Models of RLHF Annotation: Extension, Evidence, and Authority Steve Coyne et.al. 2604.25895 null
2026-04-28 Consistent Variable Selection for GARCH-X Models Adriano Zanin Zambom et.al. 2604.25894 null
2026-04-28 No Pedestrian Left Behind: Real-Time Detection and Tracking of Vulnerable Road Users for Adaptive Traffic Signal Control Anas Gamal Aly et.al. 2604.25887 null
2026-04-28 Explainable AI for Jet Tagging: A Comparative Study of GNNExplainer, GNNShap, and GradCAM for Jet Tagging in the Lund Jet Plane Pahal D. Patel et.al. 2604.25885 null
2026-04-28 When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Shuning Shang et.al. 2604.25872 null
2026-04-28 Implications of weak convergence rates of Markov transition kernels Austin Brown et.al. 2604.25867 null
2026-04-28 Semi-Markov Reinforcement Learning for City-Scale EV Ride-Hailing with Feasibility-Guaranteed Actions An Nguyen et.al. 2604.25848 null
2026-04-28 Analysis of quarkonium polarization in proton-proton (p-p) collisions at LHC using PYTHIA model Deekshit Kumar et.al. 2604.25838 null
2026-04-28 KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning Yixuan Huang et.al. 2604.25788 null
2026-04-28 EOS-Bench: A Comprehensive Benchmark for Earth Observation Satellite Scheduling Qian Yin et.al. 2604.25782 null
2026-04-28 Sensitivity-Based Tube NMPC for Cooperative Aerial Structures Under Parametric Uncertainty Giuseppe Silano et.al. 2604.25766 null
2026-04-28 QAROO: AI-Driven Online Task Offloading for Energy-Efficient and Sustainable MEC Networks Yongtao Yao et.al. 2604.25740 null
2026-04-28 Step-Audio-R1.5 Technical Report Yuxin Zhang et.al. 2604.25719 null
2026-04-28 A modelling perspective on mosquito infectiousness: time-varying transmission competence in arbovirus vector Léa Loisel et.al. 2604.25714 null
2026-04-28 Adaptive Meta-Learning Stochastic Gradient Hamiltonian Monte Carlo Simulation for Bayesian Updating of Structural Dynamic Models Xianghao Meng et.al. 2604.25710 null
2026-04-28 Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation Mehrdad Ghassabi et.al. 2604.25702 null
2026-04-28 K-CARE: Knowledge-driven Symmetrical Contextual Anchoring and Analogical Prototype Reasoning for E-commerce Relevance Chen Yifei et.al. 2604.25683 null
2026-04-28 Modeling Human-Like Color Naming Behavior in Context Yuqing Zhang et.al. 2604.25674 null
2026-04-27 An axion framework for Particle-in-Cell codes with Monte-Carlo sampling: emission, absorption, and detailed balance in plasmas Miles Radford et.al. 2604.24627 null
2026-04-27 Improving Vision-language Models with Perception-centric Process Reward Models Yingqian Min et.al. 2604.24583 null
2026-04-27 Hierarchical Behaviour Spaces Michael Tryfan Matthews et.al. 2604.24558 null
2026-04-27 GradMAP: Gradient-Based Multi-Agent Proximal Learning for Grid-Edge Flexibility Yihong Zhou et.al. 2604.24549 null
2026-04-27 Hierarchical Causal Uplift Modeling in Overlapping Customer Journeys Jorge Pellegrini et.al. 2604.24533 null
2026-04-27 A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning Ying-Tu Chen et.al. 2604.24532 null
2026-04-27 DECOFFEE: Decentralized Reinforcement Learning for Time-critical Workload Offloading and Energy Efficiency across the Computing Continuum Anastasios Giannopoulos et.al. 2604.24507 null
2026-04-27 TARMM: Scaling Delay-Critical Edge AI Offloading in 5G O-RAN via Temporal Graph Mobility Management Peihao Yan et.al. 2604.24501 null
2026-04-27 Comparative Evaluation of Modern Deep Learning Methodologies for Portfolio Optimization Samuel Ozechi et.al. 2604.24486 null
2026-04-27 Potential pof laser-driven VHEEs towards FLASH radiotherapy: Monte Carlo dosimetric study of single-field pencil beam scanning of a brain tumor Leonida A. Gizzi et.al. 2604.24417 null
2026-04-27 An Automatic Ground Collision Avoidance System with Reinforcement Learning Seyyid Osman Sevgili et.al. 2604.24403 null
2026-04-27 Beam Scheduling for Cross-Layer ISAC: A Deep Reinforcement Learning Approach Xiyu Wang et.al. 2604.24369 null
2026-04-27 DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models Dake Bu et.al. 2604.24357 null
2026-04-27 An Aircraft Upset Recovery System with Reinforcement Learning Mahir Demir et.al. 2604.24355 null
2026-04-27 See Further, Think Deeper: Advancing VLM’s Reasoning Ability with Low-level Visual Cues and Reflection Zhiheng Wu et.al. 2604.24339 null
2026-04-27 Perfecting Aircraft Maneuvers with Reinforcement Learning Atahan Cilan et.al. 2604.24338 null
2026-04-27 New non-Euclidean neural quantum states from additional types of hyperbolic recurrent neural networks H. L. Dao et.al. 2604.24337 null
2026-04-27 Pre-localization of Massive Black Hole Binaries in the Millihertz Band Xue-Ting Zhang et.al. 2604.24330 null
2026-04-27 DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents Junshuo Zhang et.al. 2604.24320 null
2026-04-27 On the complexity of quantum numerical integration: an angle-structure characterization Francisco Chinesta et.al. 2604.24289 null
2026-04-24 Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Meng Chu et.al. 2604.22748 null
2026-04-24 GCImOpt: Learning efficient goal-conditioned policies by imitating optimal trajectories Jon Goikoetxea et.al. 2604.22724 null
2026-04-24 Replica Tensor Train Miha Srdinsek et.al. 2604.22718 null
2026-04-24 ATRS: Adaptive Trajectory Re-splitting via a Shared Neural Policy for Parallel Optimization Jiajun Yu et.al. 2604.22715 null
2026-04-24 Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought Keshav Ramji et.al. 2604.22709 null
2026-04-24 Gel’fand Integration of B(E, F*)-Valued Functions With Emphasis on (q, p)-Summing Operators Matija Milović et.al. 2604.22701 null
2026-04-24 A homogeneous three-dimensional view of Molecular Cloud kinematics out to 2.5 kpc. Using Young Stellar Objects and Open Clusters as complementary tracers Xabier Pérez-Couto et.al. 2604.22573 null
2026-04-24 Adversarial Co-Evolution of Malware and Detection Models: A Bilevel Optimization Perspective Olha Jurečková et.al. 2604.22569 null
2026-04-24 Learning Evidence Highlighting for Frozen LLMs Shaoang Li et.al. 2604.22565 null
2026-04-24 SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning Jichao Wang et.al. 2604.22558 null
2026-04-24 Mean-Field Theory for the Three-State Active Lattice Gas Model Ana L. N. Dias et.al. 2604.22536 null
2026-04-24 Particle-Matter Interactions Giuseppe Lerner et.al. 2604.22508 null
2026-04-24 Objective Shaping with Hard Negatives: Windowed Partial AUC Optimization for RL-based LLM Recommenders Wentao Shi et.al. 2604.22504 null
2026-04-24 Measuring and Mitigating Persona Distortions from AI Writing Assistance Paul Röttger et.al. 2604.22503 null
2026-04-24 Inference in Tightly Identified and Large-Scale Sign-Restricted SVARs Markku Lanne et.al. 2604.22445 null
2026-04-24 Revisiting Neural Activation Coverage for Uncertainty Estimation Benedikt Franke et.al. 2604.22360 null
2026-04-24 SOC-ICNN: From Polyhedral to Conic Geometry for Learning Convex Surrogate Functions Kang Liu et.al. 2604.22355 null
2026-04-24 Spiral, target, stripe, and disordered waves in active six-state Potts models Hiroshi Noguchi et.al. 2604.22353 null
2026-04-24 Finite element model updating of building structures under seismic excitation: A parallelized latent space-based Bayesian framework Taro Yaoyama et.al. 2604.22305 null
2026-04-24 Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings Peixi Wu et.al. 2604.22280 null
2026-04-23 Nemobot Games: Crafting Strategic AI Gaming Agents for Interactive Learning with Large Language Models Chee Wei Tan et.al. 2604.21896 null
2026-04-23 Testing solitonic boson star interpretations of Sagittarius A* with near-infrared flare astrometry Xiangyu Wang et.al. 2604.21883 null
2026-04-23 Replay-buffer engineering for noise-robust quantum circuit optimization Akash Kundu et.al. 2604.21863 null
2026-04-23 Beyond Expected Information Gain: Stable Bayesian Optimal Experimental Design with Integral Probability Metrics and Plug-and-Play Extensions Di Wu et.al. 2604.21849 null
2026-04-23 Fairness under uncertainty in sequential decisions Michelle Seng Ah Lee et.al. 2604.21711 null
2026-04-23 Task-specific Subnetwork Discovery in Reinforcement Learning for Autonomous Underwater Navigation Yi-Ling Liu et.al. 2604.21640 null
2026-04-23 AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use Yuanjie Lyu et.al. 2604.21590 null
2026-04-23 Generative Learning Enhanced Intelligent Resource Management for Cell-Free Delay Deterministic Communications Shuangbo Xiong et.al. 2604.21587 null
2026-04-23 X2-N: A Transformable Wheel-legged Humanoid Robot with Dual-mode Locomotion and Manipulation Yan Ning et.al. 2604.21541 null
2026-04-23 On a class of constrained particle filters for continuous-discrete state space models Utku Erdogan et.al. 2604.21538 null
2026-04-23 Benchmarking the Utility of Privacy-Preserving Cox Regression Under Data-Driven Clipping Bounds: A Multi-Dataset Simulation Study Keita Fukuyama et.al. 2604.21491 null
2026-04-23 Efficient Agent Evaluation via Diversity-Guided User Simulation Itay Nakash et.al. 2604.21480 null
2026-04-23 Dynamical Priors as a Training Objective in Reinforcement Learning Sukesh Subaharan et.al. 2604.21464 null
2026-04-23 Tempered Sequential Monte Carlo for Trajectory and Policy Optimization with Differentiable Dynamics Heng Yang et.al. 2604.21456 null
2026-04-23 The virial expansion of the Hydrogen equation of state in comparison to PIMC simulations: the quasiparticle concept, IPD, and ionization degree Gerd Röpke et.al. 2604.21425 null
2026-04-23 S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images Qingxiao Li et.al. 2604.21409 null
2026-04-23 KD-CVG: A Knowledge-Driven Approach for Creative Video Generation Linkai Liu et.al. 2604.21362 null
2026-04-23 ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs Jian Cui et.al. 2604.21357 null
2026-04-23 RPG: Robust Policy Gating for Smooth Multi-Skill Transitions in Humanoid Fighting Yucheng Xin et.al. 2604.21355 null
2026-04-23 Learn Weightlessness: Imitate Non-Self-Stabilizing Motions on Humanoid Robot Yucheng Xin et.al. 2604.21351 null
2026-04-22 Unconventional Quantum Criticality in Long-Range Spin-1 Chains: Insights from Entanglement Entropy and Bipartite Fluctuations Justin Tim-Lok Chau et.al. 2604.20831 null
2026-04-22 ParetoSlider: Diffusion Models Post-Training for Continuous Reward Control Shelly Golan et.al. 2604.20816 null
2026-04-22 V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Yubo Jiang et.al. 2604.20755 null
2026-04-22 Fast Bayesian equipment condition monitoring via simulation based inference: applications to heat exchanger health Peter Collett et.al. 2604.20735 null
2026-04-22 Near-Future Policy Optimization Chuanyu Qin et.al. 2604.20733 null
2026-04-22 Visual-Tactile Peg-in-Hole Assembly Learning from Peg-out-of-Hole Disassembly Yongqiang Zhao et.al. 2604.20712 null
2026-04-22 SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models Jiahao Xie et.al. 2604.20705 null
2026-04-22 Divide-and-Conquer Neural Network Surrogates for Quantum Sampling: Accelerating Markov Chain Monte Carlo in Large-Scale Constrained Optimization Problems Yuya Kawamata et.al. 2604.20701 null
2026-04-22 Bayesian approach for uncertainty quantification of hybrid spectral unmixing in $γ$ -ray spectrometry Dinh Triem Phan et.al. 2604.20691 null
2026-04-22 FingerEye: Continuous and Unified Vision-Tactile Sensing for Dexterous Manipulation Zhixuan Xu et.al. 2604.20689 null
2026-04-22 MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment Andor Vári-Kakas et.al. 2604.20685 null
2026-04-22 GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning Jingyi Wang et.al. 2604.20659 null
2026-04-22 Polaron transport and Verwey transition in magnetite Nikita Fominykh et.al. 2604.20642 null
2026-04-22 Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Aravind Venugopal et.al. 2604.20627 null
2026-04-22 Self-Guided Plan Extraction for Instruction-Following Tasks with Goal-Conditional Reinforcement Learning Zoya Volovikova et.al. 2604.20601 null
2026-04-22 Predicting co-segregation in multicomponent alloys with solute-solute interactions Zuoyong Zhang et.al. 2604.20593 null
2026-04-22 A Hierarchical MARL-Based Approach for Coordinated Retail P2P Trading and Wholesale Market Participation of DERs Patrick Wilk et.al. 2604.20586 null
2026-04-22 Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents Yuxuan Cai et.al. 2604.20572 null
2026-04-22 Where Reasoning Breaks: Logic-Aware Path Selection by Controlling Logical Connectives in LLMs Reasoning Chains Seunghyun Park et.al. 2604.20564 null
2026-04-22 Conditional Monte Carlo Tree Diffusion for Designing Cell-Type-Specific and Biologically Faithful Regulatory DNA Animesh Awasthi et.al. 2604.20488 null
2026-04-21 Safe Continual Reinforcement Learning in Non-stationary Environments Austin Coursey et.al. 2604.19737 null
2026-04-21 FASTER: Value-Guided Sampling for Fast RL Perry Dong et.al. 2604.19730 null
2026-04-21 On two ways to use determinantal point processes for Monte Carlo integration Guillaume Gautier et.al. 2604.19698 null
2026-04-21 Planning in entropy-regularized Markov decision processes and games Jean-Bastien Grill et.al. 2604.19695 null
2026-04-21 Learning Hybrid-Control Policies for High-Precision In-Contact Manipulation Under Uncertainty Hunter L. Brown et.al. 2604.19677 null
2026-04-21 Pause or Fabricate? Training Language Models for Grounded Reasoning Yiwen Qiu et.al. 2604.19656 null
2026-04-21 Regulation Zero 2: A Flow-Centric Sequential Regulation Planning Framework to Counter Regulation Cascading in Pre-tactical Air Traffic Flow Management Thinh Hoang et.al. 2604.19641 null
2026-04-21 Quadrature-Enhanced Monte Carlo fPINN Method for High-Dimensional Fractional PDEs Qingkui Ma et.al. 2604.19601 null
2026-04-21 SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing Ying Zeng et.al. 2604.19587 null
2026-04-21 Regularity Analysis and Tensor Neural Network Methods for Quasiperiodic Elliptic Equations Jingze Ren et.al. 2604.19575 null
2026-04-21 Lyapunov-Certified Direct Switching Theory for Q-Learning Donghwan Lee et.al. 2604.19569 null
2026-04-21 Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic Chuou Xu et.al. 2604.19567 null
2026-04-21 DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Zhihong Zhang et.al. 2604.19544 null
2026-04-21 The swept-back multipolar magnetic field of neutron stars: Application to NICER MSP J0030+0451 Anu Kundu et.al. 2604.19534 null
2026-04-21 Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps Alankrit Chona et.al. 2604.19533 null
2026-04-21 Design and preliminary performance study of the broad-band spectrometer detector for POLAR-2 Jian-Chao Sun et.al. 2604.19497 null
2026-04-21 EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Chengjun Pan et.al. 2604.19485 null
2026-04-21 HP-Edit: A Human-Preference Post-Training Framework for Image Editing Fan Li et.al. 2604.19406 null
2026-04-21 Lost in Translation: Do LVLM Judges Generalize Across Languages? Md Tahmid Rahman Laskar et.al. 2604.19405 null
2026-04-21 LASER: Learning Active Sensing for Continuum Field Reconstruction Huayu Deng et.al. 2604.19355 null
2026-04-21 Experimental Demonstration of SDRL Controller for TS Wave Suppression with DBD Actuator Babak Mohammadikalakoo et.al. 2604.19326 null
2026-04-21 Single-shot quantum neural networks with amplitude estimation Jaemin Seo et.al. 2604.19320 null
2026-04-21 Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation Rui Li et.al. 2604.19234 null
2026-04-21 Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification Quan Zhang et.al. 2604.19218 null
2026-04-21 Reasoning-Aware AIGC Detection via Alignment and Reinforcement Zhao Wang et.al. 2604.19172 null
2026-04-21 RL-ABC: Reinforcement Learning for Accelerator Beamline Control Anwar Ibrahim et.al. 2604.19146 null
2026-04-21 ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation Kunquan Li et.al. 2604.19144 null
2026-04-21 The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models Shuai Wu et.al. 2604.19139 null
2026-04-21 GraphRAG-IRL: Personalized Recommendation with Graph-Grounded Inverse Reinforcement Learning and LLM Re-ranking Siqi Liang et.al. 2604.19128 null
2026-04-21 Reinforcement Learning Enabled Adaptive Multi-Task Control for Bipedal Soccer Robots Yulai Zhang et.al. 2604.19104 null
2026-04-21 Multi-Gait Learning for Humanoid Robots Using Reinforcement Learning with Selective Adversarial Motion Prior Yuanye Wu et.al. 2604.19102 null
2026-04-21 OLLM: Options-based Large Language Models Shashank Sharma et.al. 2604.19087 null
2026-04-21 TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only Yilun Liu et.al. 2604.19070 null
2026-04-21 Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification from Radiology Reports Yishu Wei et.al. 2604.19060 null
2026-04-21 Intentional Updates for Streaming Reinforcement Learning Arsalan Sharifnassab et.al. 2604.19033 null
2026-04-21 MonteQ: A Monte Carlo Tree Search Based Quantum Circuit Synthesis Framework Mulundano Machiya et.al. 2604.19029 null
2026-04-20 Bounded Ratio Reinforcement Learning Yunke Ao et.al. 2604.18578 null
2026-04-20 When Can LLMs Learn to Reason with Weak Supervision? Salman Rahman et.al. 2604.18574 null
2026-04-20 FUSE: Ensembling Verifiers with Zero Labeled Data Joonhyuk Lee et.al. 2604.18547 null
2026-04-20 OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Xinyu Ma et.al. 2604.18530 null
2026-04-20 UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models Jiaqi Wang et.al. 2604.18518 null
2026-04-20 Low-noise Pauli-consistent ensemble Monte Carlo for graphene with electron-electron scattering Tigran Zalinyan et.al. 2604.18517 null
2026-04-20 Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks Md Rysul Kabir et.al. 2604.18510 null
2026-04-20 Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Zhenwen Liang et.al. 2604.18493 null
2026-04-20 XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments Kangan Qian et.al. 2604.18484 null
2026-04-20 Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts Jacob Morrison et.al. 2604.18473 null
2026-04-20 Geometric Trajectory Optimization for TRACON Arrivals: An NLP Approach with ATC Vectoring Maneuver Modeling Yutian Pang et.al. 2604.18454 null
2026-04-20 Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning Hen Davidov et.al. 2604.18419 null
2026-04-20 Impact of Initial Charge Distributions on the Kinetics of Charged Particle Coagulation Gustavo Castillo et.al. 2604.18415 null
2026-04-20 StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning Daoyu Wang et.al. 2604.18401 null
2026-04-20 OpenGame: Open Agentic Coding for Games Yilei Jiang et.al. 2604.18394 null
2026-04-20 Learning from Less: Measuring the Effectiveness of RLVR in Low Data and Compute Regimes Justin Bauer et.al. 2604.18381 null
2026-04-20 Training and Agentic Inference Strategies for LLM-based Manim Animation Generation Ravidu Suien Rammuni Silva et.al. 2604.18364 null
2026-04-20 Momentum Stability and Adaptive Control in Stochastic Reconfiguration Yuyang Wang et.al. 2604.18357 null
2026-04-20 PARM: Pipeline-Adapted Reward Model Xingyu Fan et.al. 2604.18327 null
2026-04-20 Scale-free adaptive planning for deterministic dynamics & discounted rewards Peter L. Bartlett et.al. 2604.18312 null
2026-04-19 Guardrails in Logit Space: Safety Token Regularization for LLM Alignment Thong Bach et.al. 2604.17210 null
2026-04-19 Robust Resource Allocation in RIS-Assisted Wireless Networks Integrating NOMA and Over-the-Air Federated Learning Saeid Pakravan et.al. 2604.17201 null
2026-04-19 Do LLM-derived graph priors improve multi-agent coordination? Nikunj Gupta et.al. 2604.17191 null
2026-04-18 Nuclear Heterodyne Interferometry for Gravitational Spectroscopy Ralf Röhlsberger et.al. 2604.17157 null
2026-04-18 Uncertainty Quantification in PINNs for Turbulent Flows: Bayesian Inference and Repulsive Ensembles Khemraj Shukla et.al. 2604.17156 null
2026-04-18 Live LTL Progress Tracking: Towards Task-Based Exploration Noel Brindise et.al. 2604.17106 null
2026-04-18 Reference-state System Reliability method for scalable uncertainty quantification of coherent systems Ji-Eun Byun et.al. 2604.17066 null
2026-04-18 Web-Gewu: A Browser-Based Interactive Playground for Robot Reinforcement Learning Kaixuan Chen et.al. 2604.17050 null
2026-04-18 Enabling Safety-Critical Wireless Communications via Safe Reinforcement Learning Haoran Peng et.al. 2604.17032 null
2026-04-18 Small Model as Master Orchestrator: Learning Unified Agent-Tool Orchestration with Parallel Subtask Decomposition Wenzhen Yuan et.al. 2604.17009 null
2026-04-18 SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Yifu Huo et.al. 2604.16995 null
2026-04-18 MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models Zhaokang Liao et.al. 2604.16972 null
2026-04-18 NaviFormer: A Deep Reinforcement Learning Transformer-like Model to Holistically Solve the Navigation Problem Daniel Fuertes et.al. 2604.16967 null
2026-04-18 Correcting Low-Signal Sensitivity in the Deliberative Reason Index Francesco Veri et.al. 2604.16963 null
2026-04-18 Multi-stage Planning for Multi-target Surveillance using Aircrafts Equipped with Synthetic Aperture Radars Aware of Target Visibility Daniel Fuertes et.al. 2604.16962 null
2026-04-18 Noise-Adaptive Diffusion Sampling for Inverse Problems Without Task-Specific Tuning Yingzhi Xia et.al. 2604.16919 null
2026-04-18 Freshness-Aware Prioritized Experience Replay for LLM/VLM Reinforcement Learning Weiyu Ma et.al. 2604.16918 null
2026-04-18 ProtoCycle: Reflective Tool-Augmented Planning for Text-Guided Protein Design Yutang Ge et.al. 2604.16896 null
2026-04-18 EasyVideoR1: Easier RL for Video Understanding Chuanyu Qin et.al. 2604.16893 null
2026-04-18 Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation Jiang Zhou et.al. 2604.16881 null
2026-04-17 Evaluating the Progression of Large Language Model Capabilities for Small-Molecule Drug Design Shriram Chennakesavalu et.al. 2604.16279 null
2026-04-17 VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects Xiangbo Gao et.al. 2604.16272 null
2026-04-17 Improved Desalination by Polymer Grafting Mamta Yadav et.al. 2604.16267 null
2026-04-17 Beyond Distribution Sharpening: The Importance of Task Rewards Sarthak Mittal et.al. 2604.16259 null
2026-04-17 Tensor decomposition of $e^+e^-\toπ^+π^-γ$ to higher orders in the dimensional regulator Thomas Dave et.al. 2604.16251 null
2026-04-17 Find, Fix, Reason: Context Repair for Video Reasoning Haojian Huang et.al. 2604.16243 null
2026-04-17 Detecting and Suppressing Reward Hacking with Gradient Fingerprints Songtao Wang et.al. 2604.16242 null
2026-04-17 A Bayesian Updating Framework for Long-term Multi-Environment Trial Data in Plant Breeding Stephan Bark et.al. 2604.16203 null
2026-04-17 Some results on small ordered and cyclic Ramsey numbers Nino Bašić et.al. 2604.16188 null
2026-04-17 Single-Satellite Quantum Repeater Performance Analysis Cameron Paterson et.al. 2604.16165 null
2026-04-17 AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency Max Henning Höth et.al. 2604.16158 null
2026-04-17 Hopping-Mediated Charge Transport in Graphene Beyond the Ballistic Regime J. P. Dadario Pereira et.al. 2604.16152 null
2026-04-17 Prompt Gamma Timing for range verification with carbon ion irradiation: first experimental measurements and comparison with Geant4 Monte Carlo simulations Iram Barbaro Rivas Ortiz et.al. 2604.16131 null
2026-04-17 Beyond One-Size-Fits-All: Adaptive Test-Time Augmentation for Sequential Recommendation Xibo Li et.al. 2604.16121 null
2026-04-17 Convergence Time Distributions for Max-Consensus over Unreliable Networks Katharina Stich et.al. 2604.16069 null
2026-04-17 Safe Deep Reinforcement Learning for Building Heating Control and Demand-side Flexibility Colin Jüni et.al. 2604.16033 null
2026-04-17 A Comparison of Joint and Stepwise Dynamic Cognitive Diagnostic Models Yawen Ma et.al. 2604.16031 null
2026-04-17 AgentV-RL: Scaling Reward Modeling with Agentic Verifier Jiazheng Zhang et.al. 2604.16004 null
2026-04-17 $G_2$ -structures as Octonion Algebras Isak Sundelius et.al. 2604.15966 null
2026-04-17 Reweighting Estimators for Density Response in Path Integral Monte Carlo: Applications to linear, nonlinear and cross-species density response Pontus Svensson et.al. 2604.15934 null
2026-04-16 RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework Hao Gao et.al. 2604.15308 null
2026-04-16 Generalization in LLM Problem Solving: The Case of the Shortest Path Yao Tong et.al. 2604.15306 null
2026-04-16 Abstract Sim2Real through Approximate Information States Yunfu Deng et.al. 2604.15289 null
2026-04-16 R3D: Revisiting 3D Policy Learning Zhengdong Hong et.al. 2604.15281 null
2026-04-16 From Tokens to Steps: Verification-Aware Speculative Decoding for Efficient Multi-Step Reasoning Kiran Purohit et.al. 2604.15244 null
2026-04-16 A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics Fawad Javed Fateh et.al. 2604.15215 null
2026-04-16 RL-STPA: Adapting System-Theoretic Hazard Analysis for Safety-Critical Reinforcement Learning Steven A. Senczyszyn et.al. 2604.15201 null
2026-04-16 Quantum Metropolis-Hastings via Penalised Qubitized Walks: Spectral Filtering and Circuit Implementation Miguel Carrasco-Arango et.al. 2604.15179 null
2026-04-16 LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking Lukas Helff et.al. 2604.15149 null
2026-04-16 IG-Search: Step-Level Information Gain Rewards for Search-Augmented Reasoning Zihan Liang et.al. 2604.15148 null
2026-04-16 A minimal implementation of Yang-Mills theory on a digital quantum computer Georg Bergner et.al. 2604.15132 null
2026-04-16 OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis Kanzhi Cheng et.al. 2604.15093 null
2026-04-16 Static heterogeneity generates apparent universality in first-passage bursty dynamics Morten Møller et.al. 2604.15084 null
2026-04-16 High-temperature charge-4e superconductivity in SU(4) interacting fermions Shao-Hang Shi et.al. 2604.15056 null
2026-04-16 Nonlinear backstepping with saturation for low-thrust station-keeping of libration point orbits António Nunes et.al. 2604.15028 null
2026-04-16 Fully Differentiable Ultrasound Simulation Utilizing Ray-Tracing L. River Spencer et.al. 2604.15017 null
2026-04-16 Momentum-constrained Hybrid Heuristic Trajectory Optimization Framework with Residual-enhanced DRL for Visually Impaired Scenarios Yuting Zeng et.al. 2604.14986 null
2026-04-16 Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning Jean-Bastien Grill et.al. 2604.14974 null
2026-04-16 UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards Jun Wang et.al. 2604.14967 null
2026-04-16 POMDP-based Object Search with Growing State Space and Hybrid Action Domain Yongbo Chen et.al. 2604.14965 null
2026-04-16 WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training Yifu Chen et.al. 2604.14932 null
2026-04-16 Dynamic Lagrange Multipliers in a Non-concave Utility Framework Yang Liu et.al. 2604.14924 null
2026-04-16 LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning Bowen Ping et.al. 2604.14922 null
2026-04-16 Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models Yifu Chen et.al. 2604.14920 null
2026-04-16 Data-driven Linear Quadratic Integral Control: A Convex Formulation and Policy Gradient Approach Armin Gießler et.al. 2604.14905 null
2026-04-16 Beyond Importance Sampling: Rejection-Gated Policy Optimization Ziwu Sun et.al. 2604.14895 null
2026-04-16 GenRec: A Preference-Oriented Generative Framework for Large-Scale Recommendation Yanyan Zou et.al. 2604.14878 null
2026-04-16 Does RL Expand the Capability Boundary of LLM Agents? A PASS@(k,T) Analysis Zhiyuan Zhai et.al. 2604.14877 null
2026-04-15 **From $P(y x)$ to $P(y)$ : Investigating Reinforcement Learning in Pre-train Space** Yuqiao Tan et.al. 2604.14142
2026-04-15 AI-assisted modeling and Bayesian inference of unpolarized quark transverse momentum distributions from Drell-Yan data Zhong-Bo Kang et.al. 2604.14133 null
2026-04-15 Simulating the dynamics of an SU(2) matrix model on a trapped-ion quantum computer Gavin S. Hartnett et.al. 2604.14094 null
2026-04-15 Multistage Conditional Compositional Optimization Buse Şen et.al. 2604.14075 null
2026-04-15 Finding and characterising physical states of Euclidean Abelianized loop quantum gravity using neural quantum states Hanno Sahlmann et.al. 2604.14067 null
2026-04-15 A Comparative Study of Dynamic Programming and Reinforcement Learning in Finite Horizon Dynamic Pricing Lev Razumovskiy et.al. 2604.14059 null
2026-04-15 Enhancing Local Life Service Recommendation with Agentic Reasoning in Large Language Model Shiteng Cao et.al. 2604.14051 null
2026-04-15 Hierarchical Reinforcement Learning with Runtime Safety Shielding for Power Grid Operation Gitesh Malik et.al. 2604.14032 null
2026-04-15 Physics-Informed Neural Networks for Methane Sorption: Cross-Gas Transfer Learning, Ensemble Collapse Under Physics Constraints, and Monte Carlo Dropout Uncertainty Quantification Mohammad Nooraiepour et.al. 2604.13992 null
2026-04-15 Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation Shangzhe Li et.al. 2604.13966 null
2026-04-15 Three-dimensional photon transport in spinodal photocatalytic aerogels: how bicontinuous morphology controls kinetic rate constants Renaud A. L. Vallée et.al. 2604.13929 null
2026-04-15 DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off Xiaofan Li et.al. 2604.13902 null
2026-04-15 Beyond Conservative Automated Driving in Multi-Agent Scenarios via Coupled Model Predictive Control and Deep Reinforcement Learning Saeed Rahmani et.al. 2604.13891 null
2026-04-15 Drowsiness-Aware Adaptive Autonomous Braking System based on Deep Reinforcement Learning for Enhanced Road Safety Hossem Eddine Hafidi et.al. 2604.13878 null
2026-04-15 Fourier Dimension in Duffin–Schaeffer Conjecture Bo Tan et.al. 2604.13868 null
2026-04-15 Robust parameter inference for Taiji via time-frequency contrastive learning and normalizing flows Tian-Yang Sun et.al. 2604.13867 null
2026-04-15 Simulation-Based Optimisation of Batting Order and Bowling Plans in T20 Cricket Tinniam V Ganesh et.al. 2604.13861 null
2026-04-15 First Passage Times for Variable-Order Time-Fractional Diffusion Wancheng Li et.al. 2604.13852 null
2026-04-15 Frequency Response of Nonlinear Systems: Notions, Analysis, and Graphical Representation Alessio Moreschini et.al. 2604.13842 null
2026-04-15 Robust Reward Modeling for Large Language Models via Causal Decomposition Yunsheng Lu et.al. 2604.13833 null
2026-04-13 Solving Physics Olympiad via Reinforcement Learning on Physics Simulators Mihir Prabhudesai et.al. 2604.11805 null
2026-04-13 ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents Fei Tang et.al. 2604.11784 null
2026-04-13 Efficient KernelSHAP Explanations for Patch-based 3D Medical Image Segmentation Ricardo Coimbra Brioso et.al. 2604.11775 null
2026-04-13 Autonomous Diffractometry Enabled by Visual Reinforcement Learning J. Oppliger et.al. 2604.11773 null
2026-04-13 Discourse Diversity in Multi-Turn Empathic Dialogue Hongli Zhan et.al. 2604.11742 null
2026-04-13 Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games Keyang Zhong et.al. 2604.11741 null
2026-04-13 AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation Mingyang Li et.al. 2604.11674 null
2026-04-13 Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind Hanqi Xiao et.al. 2604.11666 null
2026-04-13 Back to Basics: Let Conversational Agents Remember with Just Retrieval and Generation Yuqian Wu et.al. 2604.11628 null
2026-04-13 RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time Haozhe Wang et.al. 2604.11626 null
2026-04-13 Spectrum analysis with quantum dynamical systems. II. Finite-time analysis Xinyi Sui et.al. 2604.11614 null
2026-04-13 Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation Jiashu Yao et.al. 2604.11611 null
2026-04-13 Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language Peijie Wang et.al. 2604.11600 null
2026-04-13 A game-theoretical interpretation for a doubly nonlinear parabolic equation Felix del Teso et.al. 2604.11592 null
2026-04-13 Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Liujie Zhang et.al. 2604.11554 null
2026-04-13 Eliciting Medical Reasoning with Knowledge-enhanced Data Synthesis: A Semi-Supervised Reinforcement Learning Approach Haolin Li et.al. 2604.11547 null
2026-04-13 RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience Hanbo Huang et.al. 2604.11546 null
2026-04-13 Triviality Corrected Endogenous Reward Xinda Wang et.al. 2604.11522 null
2026-04-13 The Price of Ignorance: Information-Free Quotation for Data Retention in Machine Unlearning Bin Han et.al. 2604.11511 null
2026-04-13 Policy Split: Incentivizing Dual-Mode Exploration in LLM Reinforcement with Dual-Mode Entropy Regularization Jiashu Yao et.al. 2604.11510 null
2026-04-12 Aerial IRS Deployment-Aided Secure Computation Offloading Against DISCO Jamming Attacks Minghui Min et.al. 2604.10558 null
2026-04-12 Simple but Stable, Fast and Safe: Achieve End-to-end Control by High-Fidelity Differentiable Simulation Fanxing Li et.al. 2604.10548 null
2026-04-12 Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training? Wanyi Chen et.al. 2604.10547 null
2026-04-12 Discrete-Time Backward Stochastic LQ Control Problem Hu Ligui et.al. 2604.10510 null
2026-04-12 Beyond Compliance: A Resistance-Informed Motivation Reasoning Framework for Challenging Psychological Client Simulation Danni Liu et.al. 2604.10507 null
2026-04-12 Entangled happily ever after: Wedding reception seating mapped to classical and quantum optimizers Karie A. Nicholas Vikram Khipple Mulligan et.al. 2604.10497 null
2026-04-12 SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents Mahir Labib Dihan et.al. 2604.10493 null
2026-04-12 The effect of grain boundaries on magnetic exchange interactions in iron Martin Zelený et.al. 2604.10489 null
2026-04-12 AdverMCTS: Combating Pseudo-Correctness in Code Generation via Adversarial Monte Carlo Tree Search Qingyao Li et.al. 2604.10449 null
2026-04-12 The SpinQuest Microwave System for Dynamic Nuclear Polarization Vibodha Bandara et.al. 2604.10447 null
2026-04-12 Safety Guarantees in Zero-Shot Reinforcement Learning for Cascade Dynamical Systems Shima Rabiei et.al. 2604.10429 null
2026-04-12 A Queueing-Theoretic Framework for Dynamic Attack Surfaces: Data-Integrated Risk Analysis and Adaptive Defense Jihyeon Yun et.al. 2604.10427 null
2026-04-12 Orthogonal machine learning for conditional odds and risk ratios Jiacheng Ge et.al. 2604.10412 null
2026-04-11 Beyond Whittle: exact finite-time multispectral statistics from a single Brownian trajectory in a harmonic trap Isaac Pérez Castillo et.al. 2604.10323 null
2026-04-11 Emergent Topological Universality and Marginal Replica Symmetry Breaking in Gauge-Correlated Spin Glasses Alok Yadav et.al. 2604.10309 null
2026-04-11 Algorithmic overlaps as thermodynamic variables: from local to cluster Monte Carlo dynamics in critical phenomena Ian Pilé et.al. 2604.10254 null
2026-04-11 A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets Zunnan Xu et.al. 2604.10252 null
2026-04-11 Warm-Started Reinforcement Learning for Iterative 3D/2D Liver Registration Hanyuan Zhang et.al. 2604.10245 null
2026-04-11 Scalable Generative Sampling and Multilevel Estimation for Lattice Field Theories Near Criticality A. Singha et.al. 2604.10209 null
2026-04-11 Policy Iteration for Stationary Discounted Hamilton–Jacobi–Bellman Equations: A Viscosity Approach Namkyeong Cho et.al. 2604.10191 null
2026-04-10 VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning Wenyi Xiao et.al. 2604.09529 null
2026-04-10 Event-Driven Temporal Graph Networks for Asynchronous Multi-Agent Cyber Defense in NetForge_RL Igor Jankowski et.al. 2604.09523 null
2026-04-10 RIRF: Reasoning Image Restoration Framework Wending Yan et.al. 2604.09511 null
2026-04-10 VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning Yucheng Shen et.al. 2604.09508 null
2026-04-10 Physics-Informed Reinforcement Learning of Spatial Density Velocity Potentials for Map-Free Racing Shathushan Sivashangaran et.al. 2604.09499 null
2026-04-10 Process Reward Agents for Steering Knowledge-Intensive Reasoning Jiwoong Sohn et.al. 2604.09482 null
2026-04-10 From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models Chenchen Zhang et.al. 2604.09459 null
2026-04-10 SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning Maksim Anisimov et.al. 2604.09452 null
2026-04-10 Physics-guided surrogate learning enables zero-shot control of turbulent wings Yuning Wang et.al. 2604.09434 null
2026-04-10 Musculoskeletal Motion Imitation for Learning Personalized Exoskeleton Control Policy in Impaired Gait Itak Choi et.al. 2604.09431 null
2026-04-10 Nii-body: Bayesian Inference of Multiplanet Dynamics via N-body Simulations Hong-Fei Jia et.al. 2604.09383 null
2026-04-10 Data-Efficient Non-Gaussian Semi-Nonparametric Density Estimation for Nonlinear Dynamical Systems Aaron R. Liao et.al. 2604.09375 null
2026-04-10 Constraining the Molecular Kennicutt-Schmidt Relation with Multi-Transition CO Observations of Nearby Galaxies Victoria G. G. Samboco et.al. 2604.09353 null
2026-04-10 Visually-Guided Policy Optimization for Multimodal Reasoning Zengbin Wang et.al. 2604.09349 null
2026-04-10 Optimal Annuitization Time under a Mortality Shock Matteo Buttarazzi et.al. 2604.09342 null
2026-04-10 Mind the Gap Between Spatial Reasoning and Acting! Step-by-Step Evaluation of Agents With Spatial-Gym Lars Benedikt Kaesberg et.al. 2604.09338 null
2026-04-10 SatQNet: Satellite-assisted Quantum Network Entanglement Routing Using Directed Line Graph Neural Networks Tobias Meuser et.al. 2604.09306 null
2026-04-10 Online Intention Prediction via Control-Informed Learning Tianyu Zhou et.al. 2604.09303 null
2026-04-10 Exact Bayesian Planning for Simple Step-Stress Accelerated Life Testing with Competing Risks Kiran Prajapat et.al. 2604.09259 null
2026-04-10 Statistical Properties of the King Wen Sequence: An Anti-Habituation Structure That Does Not Improve Neural Network Training Augustin Chan et.al. 2604.09234 null
2026-04-09 Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models Shilin Yan et.al. 2604.08545 null
2026-04-09 OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks Wenbo Hu et.al. 2604.08539 null
2026-04-09 Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest Addison J. Wu et.al. 2604.08525 null
2026-04-09 SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions Ashima Suvarna et.al. 2604.08477 null
2026-04-09 Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization Sai Srinivas Kancheti et.al. 2604.08476 null
2026-04-09 LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation Jingjing Wang et.al. 2604.08475 null
2026-04-09 Bayesian Semiparametric Multivariate Density Regression with Coordinate-Wise Predictor Selection Giovanni Toto et.al. 2604.08470 null
2026-04-09 TTVS: Boosting Self-Exploring Reinforcement Learning via Test-time Variational Synthesis Sikai Bai et.al. 2604.08468 null
2026-04-09 Less Approximates More: Harmonizing Performance and Confidence Faithfulness via Hybrid Post-Training for High-Stakes Tasks Haokai Ma et.al. 2604.08454 null
2026-04-09 NL-CPS: Reinforcement Learning-Based Kubernetes Control Plane Placement in Multi-Region Clusters Sajid Alam et.al. 2604.08434 null
2026-04-09 Synthetic Data for any Differentiable Target Tristan Thrush et.al. 2604.08423 null
2026-04-09 A parametric study of plasma instability cooling and its impact on intergalactic magnetic field constraints in GeV cascades Suman Dey et.al. 2604.08375 null
2026-04-09 ASPECT:Analogical Semantic Policy Execution via Language Conditioned Transfer Ajsal Shereef Palattuparambil et.al. 2604.08355 null
2026-04-09 ProMedical: Hierarchical Fine-Grained Criteria Modeling for Medical LLM Alignment via Explicit Injection He Geng et.al. 2604.08326 null
2026-04-09 Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data Yuchuan Deng et.al. 2604.08322 null
2026-04-09 Orbital-Selective $d$-wave Superconductivity in the Two-Band $t$-$J$ Model: Possible Applications to La$_3$Ni$_2$O$_7$ Zhan Wang et.al. 2604.08319 null
2026-04-09 The peculiar velocity correlation function of the Cosmicflows-4 catalog Yuyu Wang et.al. 2604.08314 null
2026-04-09 Characterization of afterpulse in SiPMs with single-cell readout as a function of bias voltage and fluence P. Parygin et.al. 2604.08308 null
2026-04-09 Chemistry and ro-vibrational excitation of CH $^+$ in the Planetary Nebula NGC 7027 Milan Sil et.al. 2604.08273 null
2026-04-09 SMC-AI: Scaling Monte Carlo Simulation to Four Trillion Atoms with AI Accelerators Xianglin Liu et.al. 2604.08250 null
2026-04-09 HiRO-Nav: Hybrid ReasOning Enables Efficient Embodied Navigation He Zhao et.al. 2604.08232 null
2026-04-09 OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Yiduo Jia et.al. 2604.08209 null
2026-04-09 MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning Zheng Jiang et.al. 2604.08203 null
2026-04-09 Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters Ayesha Afzal et.al. 2604.08182 null
2026-04-09 Aligning Agents via Planning: A Benchmark for Trajectory-Level Reward Modeling Jiaxuan Wang et.al. 2604.08178 null
2026-04-09 Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning Teng Pang et.al. 2604.08174 null
2026-04-09 ViVa: A Video-Generative Value Model for Robot Reinforcement Learning Jindi Lv et.al. 2604.08168 null
2026-04-09 Joint Range-Angle Estimation in Near-Field ISAC System using Uniform Circular Array Lorenzo Zaniboni et.al. 2604.08160 null
2026-04-09 Dual Approaches to Stochastic Control via SPDEs and the Pathwise Hopf Formula Mathieu Laurière et.al. 2604.08155 null
2026-04-09 A Multilevel Monte Carlo Virtual Element Method for Uncertainty Quantification of Elliptic Partial Differential Equations Paola F. Antonietti et.al. 2604.08135 null
2026-04-09 Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search Chuzhan Hao et.al. 2604.08124 null
2026-04-09 Reinforcement learning with reputation-based adaptive exploration promotes the evolution of cooperation An Li et.al. 2604.08103 null
2026-04-09 PriPG-RL: Privileged Planner-Guided Reinforcement Learning for Partially Observable Systems with Anytime-Feasible MPC Mohsen Amiri et.al. 2604.08036 null
2026-04-08 Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization Qiyao Ma et.al. 2604.07343 null
2026-04-08 Robots that learn to evaluate models of collective behavior Mathis Hocke et.al. 2604.07303 null
2026-04-08 Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions Guo Gan et.al. 2604.07277 null
2026-04-08 Robust Quadruped Locomotion via Evolutionary Reinforcement Learning Brian McAteer et.al. 2604.07224 null
2026-04-08 VersaVogue: Visual Expert Orchestration and Preference Alignment for Unified Fashion Synthesis Jian Yu et.al. 2604.07210 null
2026-04-08 BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment Mohamed Darwish Mounis et.al. 2604.07201 null
2026-04-08 Perpendicular electric field induced $s^\pm$-wave to $d$-wave superconducting transition in thin film La$_3$Ni$_2$O$_7$ Yongping Wei et.al. 2604.07185 null
2026-04-08 Smart Commander: A Hierarchical Reinforcement Learning Framework for Fleet-Level PHM Decision Optimization Yong Si et.al. 2604.07171 null
2026-04-08 Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization Yu Li et.al. 2604.07165 null
2026-04-08 Multi-Turn Reasoning LLMs for Task Offloading in Mobile Edge Computing Ning Yang et.al. 2604.07148 null
2026-04-08 Energy Saving for Cell-Free Massive MIMO Networks: A Multi-Agent Deep Reinforcement Learning Approach Qichen Wang et.al. 2604.07133 null
2026-04-08 Towards viable H $_2$ storage in Ca decorated low-dimensional materials with insights from reference quantum Monte Carlo Yasmine S. Al-Hamdani et.al. 2604.07110 null
2026-04-08 STRIDE-ED: A Strategy-Grounded Stepwise Reasoning Framework for Empathetic Dialogue Systems Hongru Ji et.al. 2604.07100 null
2026-04-08 Epistemic Robust Offline Reinforcement Learning Abhilash Reddy Chenreddy et.al. 2604.07072 null
2026-04-08 Production-Ready Automated ECU Calibration using Residual Reinforcement Learning Andreas Kampmeier et.al. 2604.07059 null
2026-04-08 Seasonality in Mixed Causal-Noncausal Processes Tomás del Barrio Castro et.al. 2604.07040 null
2026-04-08 Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argumentation Philipp D. Siedler et.al. 2604.07028 null
2026-04-08 Predictive Representations for Skill Transfer in Reinforcement Learning Ruben Vereecken et.al. 2604.07016 null
2026-04-08 EmoMAS: Emotion-Aware Multi-Agent System for High-Stakes Edge-Deployable Negotiation with Bayesian Orchestration Yunbo Long et.al. 2604.07003 null
2026-04-08 MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation Xiaoxiao Ma et.al. 2604.06966 null
2026-04-08 Learning-Based Strategy for Composite Robot Assembly Skill Adaptation Khalil Abuibaid et.al. 2604.06949 null
2026-04-08 Sustainable Transfer Learning for Adaptive Robot Skills Khalil Abuibaid et.al. 2604.06943 null
2026-04-08 Cosmological Dynamics of Exponential Quintessence Constrained by BAO, Cosmic Chronometers, and DES-SN5YR/Pantheon+ Data Sanjeeda Sultana et.al. 2604.06941 null
2026-04-07 Target Policy Optimization Jean Kaddour et.al. 2604.06159 null
2026-04-07 MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control Yuchi Wang et.al. 2604.06156 null
2026-04-07 Sequential Audit Sampling with Statistical Guarantees Masahiro Kato et.al. 2604.06116 null
2026-04-07 Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning Juekai Lin et.al. 2604.06079 null
2026-04-07 Beyond Black-Scholes: A Computational Framework for Option Pricing Using Heston, GARCH, and Jump Diffusion Models Karmanpartap Singh Sidhu et.al. 2604.06068 null
2026-04-07 HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning Jiyao Zhang et.al. 2604.06067 null
2026-04-07 Value Mirror Descent for Reinforcement Learning Zhichao Jia et.al. 2604.06039 null
2026-04-07 Design and Analysis of Chirp-Layered Superposition Coding for LoRa Jingxiang Huang et.al. 2604.06033 null
2026-04-07 A deep learning framework for jointly solving transient Fokker-Planck equations with arbitrary parameters and initial distributions Xiaolong Wang et.al. 2604.06001 null
2026-04-07 Force Polytope-Based Cant-Angle Selection for Tilting Hexarotor UAVs Alberto Piccina et.al. 2604.05998 null
2026-04-07 Numerically Exact Study of Flat-Band Superconductivity I. S. Tupitsyn et.al. 2604.05997 null
2026-04-07 You’re Pushing My Buttons: Instrumented Learning of Gentle Button Presses Raman Talwar et.al. 2604.05954 null
2026-04-07 MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning Maria Nesterova et.al. 2604.05943 null
2026-04-07 Monte-Carlo Event Generation for X-Ray Thomson Scattering Analysis Uwe Hernandez Acosta et.al. 2604.05935 null
2026-04-07 Saliency-Guided Representation with Consistency Policy Learning for Visual Unsupervised Reinforcement Learning Jingbo Sun et.al. 2604.05931 null
2026-04-07 Comments on “The impact of Solar magnetic field configurations on the production of gamma rays at the Solar disk’’ (arXiv:2512.01403) M. N. Mazziotta et.al. 2604.05879 null
2026-04-07 AgentGL: Towards Agentic Graph Learning with LLMs via Reinforcement Learning Yuanfu Sun et.al. 2604.05846 null
2026-04-07 Precise Aggressive Aerial Maneuvers with Sensorimotor Policies Tianyue Wu et.al. 2604.05828 null
2026-04-07 Reinforcement Learning with Negative Tests as Completeness Signal for Formal Specification Synthesis Zhechong Huang et.al. 2604.05820 null
2026-04-07 Hierarchical Reinforcement Learning with Augmented Step-Level Transitions for LLM Agents Shuai Zhen et.al. 2604.05808 null
2026-04-07 Emergent social transmission of model-based representations without inference Silja Keßler et.al. 2604.05777 null
2026-04-07 Percolation in the three-dimensional Ising model Jinhong Zhu et.al. 2604.05772 null
2026-04-07 High-dimensional reliability-based design optimization using stochastic emulators M. Moustapha et.al. 2604.05759 null
2026-04-07 Can Large Language Models Reinvent Foundational Algorithms? Jian Zhao et.al. 2604.05716 null
2026-04-07 CuraLight: Debate-Guided Data Curation for LLM-Centered Traffic Signal Control Qing Guo et.al. 2604.05663 null
2026-04-07 Causal Dynamical Triangulations: New Lattice Theory of Quantum Gravity J. Ambjørn et.al. 2604.05641 null
2026-04-07 Optimality Robustness in Koopman-Based Control Yicheng Lin et.al. 2604.05633 null
2026-04-07 Uncovering Linguistic Fragility in Vision-Language-Action Models via Diversity-Aware Red Teaming Baoshun Tong et.al. 2604.05595 null
2026-04-07 An Iterative Test-and-Repair Framework for Competitive Code Generation Lingxiao Tang et.al. 2604.05560 null
2026-04-07 COSMO-Agent: Tool-Augmented Agent for Closed-loop Optimization,Simulation,and Modeling Orchestration Liyuan Deng et.al. 2604.05547 null
2026-04-07 SignalClaw: LLM-Guided Evolutionary Synthesis of Interpretable Traffic Signal Control Skills Da Lei et.al. 2604.05535 null
2026-04-07 ActivityEditor: Learning to Synthesize Physically Valid Human Mobility Chenjie Yang et.al. 2604.05529 null
2026-04-07 Development of a 3D-CNN-based Prediction Model for Migration Barriers in Plasma-Wall Interactions Seiki Saito et.al. 2604.05521 null
2026-04-07 UniCreative: Unifying Long-form Logic and Short-form Sparkle via Reference-Free Reinforcement Learning Xiaolong Wei et.al. 2604.05517 null
2026-04-07 OmniDiagram: Advancing Unified Diagram Code Generation via Visual Interrogation Reward Haoyue Yang et.al. 2604.05514 null
2026-04-05 Comparative reversal learning reveals rigid adaptation in LLMs under non-stationary uncertainty Haomiaomiao Wang et.al. 2604.04182 null
2026-04-05 Variance Reduction Methods for Dirichlet Expectations Ayeong Lee et.al. 2604.04181 null
2026-04-05 Disentangling Flow Contributions from the Chiral Magnetic Effect in U+U Collisions with Forward-Backward Multiplicity Asymmetry Kaiser Shafi et.al. 2604.04178 null
2026-04-05 Learning Dexterous Grasping from Sparse Taxonomy Guidance Juhan Park et.al. 2604.04138 null
2026-04-05 The optical Su-Schrieffer-Heeger model on a triangular lattice Max Casebolt et.al. 2604.04123 null
2026-04-05 Gaussian-Process Emulation of the Redshift-Space Halo Power Spectrum Monopole in Cosmologies with Massive Neutrinos Jixin Gan et.al. 2604.04122 null
2026-04-05 Stellar Parameters and Orbital Period Estimates for Composite-Spectrum sdB+MS Binaries from LAMOST Jiangdan Li et.al. 2604.04110 null
2026-04-05 Interplay of Anisotropy, Dzyaloshinskii Moriya Interaction and Symmetry breaking Fields in a 2D XY Ferromagnet Rajdip Banerjee et.al. 2604.04104 null
2026-04-05 Restless Bandits with Individual Penalty Constraints: A New Near-Optimal Index Policy and How to Learn It Nida Zamir et.al. 2604.04101 null
2026-04-05 Fine-grained Analysis of Stability and Generalization for Stochastic Bilevel Optimization Xuelin Zhang et.al. 2604.04090 null
2026-04-05 Multi-AUV Trajectory Learning for Sustainable Underwater IoT with Acoustic Energy Transfer Mohamed Afouene Melki et.al. 2604.04079 null
2026-04-05 Searching for vector-like leptons decaying into an electron and missing transverse energy in e $^{+}$e$^{-}$ collisions with $\sqrt{s} = 240$ GeV at the FCC-ee S. Elgammal et.al. 2604.04023 null
2026-04-05 VA-FastNavi-MARL: Real-Time Robot Control with Multimedia-Driven Meta-Reinforcement Learning Yang Zhang et.al. 2604.03998 null
2026-04-05 Can LLMs Learn to Reason Robustly under Noisy Supervision? Shenzhi Yang et.al. 2604.03993 null
2026-04-05 Predict, Don’t React: Value-Based Safety Forecasting for LLM Streaming Pride Kavumba et.al. 2604.03962 null
2026-04-04 Provable Multi-Task Reinforcement Learning: A Representation Learning Framework with Low Rank Rewards Yaoze Guo et.al. 2604.03891 null
2026-04-04 A test for normality based on self-similarity Akin Anarat et.al. 2604.03810 null
2026-04-04 Computing Rare Probabilities of Voltage Collapse Tongtong Jin et.al. 2604.03807 null
2026-04-04 Decomposing Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning Zihong Gao et.al. 2604.03785 null
2026-04-04 RL-Driven Sustainable Land-Use Allocation for the Lake Malawi Basin Ying Yao et.al. 2604.03768 null
2026-04-03 Parametric SED Modelling of Protoplanetary Discs: Validation and Application to an Unstudied YSO Volkan Bakış et.al. 2604.03211 null
2026-04-03 Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models Gengwei Zhang et.al. 2604.03179 null
2026-04-03 Bounds on Decorated Sweep Covers in Tree Posets Blake A. Wilson et.al. 2604.03165 null
2026-04-03 From Gaussian Fading to Gilbert-Elliott: Bridging Physical and Link-Layer Channel Models in Closed Form Bhaskar Krishnamachari et.al. 2604.03160 null
2026-04-03 Chart-RL: Policy Optimization Reinforcement Learning for Enhanced Visual Reasoning in Chart Question Answering with Vision Language Models Yunfei Bai et.al. 2604.03157 null
2026-04-03 FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation Mingao Tan et.al. 2604.03139 null
2026-04-03 Self-Distilled RLVR Chenxu Yang et.al. 2604.03128 null
2026-04-03 Distributed Snitch Digital Twin-Based Anomaly Detection for Smart Voltage Source Converter-Enabled Wind Power Systems Mohammad Ashraf Hossain Sadi et.al. 2604.03123 null
2026-04-03 Nested Multilevel Monte Carlo with Preintegration for Efficient Risk Estimation Yu Xu et.al. 2604.03122 null
2026-04-03 Co-Evolution of Policy and Internal Reward for Language Agents Xinyu Wang et.al. 2604.03098 null
2026-04-03 Quantitative spectroscopy of single and multiple OB-type stars. Non-LTE spectrum analysis with machine learning P. Aschenbrenner et.al. 2604.03082 null
2026-04-03 JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency Aichen Cai et.al. 2604.03044 null
2026-04-03 ARM: Advantage Reward Modeling for Long-Horizon Manipulation Yiming Mao et.al. 2604.03037 null
2026-04-03 Hilbert space fragmentation in quantum Ising systems induced by side coupling E. S. Ma et.al. 2604.03026 null
2026-04-03 Behavior-Constrained Reinforcement Learning with Receding-Horizon Credit Assignment for High-Performance Control Siwei Ju et.al. 2604.03023 null
2026-04-03 R2-Write: Reflection and Revision for Open-Ended Writing with Deep Reasoning Wanlong Liu et.al. 2604.03004 null
2026-04-03 A semicontinuous relaxation of Saito’s criterion and freeness as angular minimization Tomás S. R. Silva et.al. 2604.02995 null
2026-04-03 Mitigating Reward Hacking in RLHF via Advantage Sign Robustness Shinnosuke Ono et.al. 2604.02986 null
2026-04-03 Digital Twin-Assisted In-Network and Edge Collaboration for Joint User Association, Task Offloading, and Resource Allocation in the Metaverse Ibrahim Aliyu et.al. 2604.02938 null
2026-04-03 Towards Near-Real-Time Telemetry-Aware Routing with Neural Routing Algorithms Andreas Boltres et.al. 2604.02927 null
2026-04-02 Beyond Referring Expressions: Scenario Comprehension Visual Grounding Ruozhen He et.al. 2604.02323 null
2026-04-02 Detecting Symmetry-Resolved Entanglement: A Quantum Monte Carlo Approach Kuangjie Chen et.al. 2604.02307 null
2026-04-02 Disentangled Deep Priors for Bayesian Inverse Problems Arkaprabha Ganguli et.al. 2604.02304 null
2026-04-02 Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing Gengsheng Li et.al. 2604.02288 null
2026-04-02 CIVIC: Cooperative Immersion Via Intelligent Credit-sharing in DRL-Powered Metaverse Amr Aboeleneen et.al. 2604.02284 null
2026-04-02 SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization Zhengxi Lu et.al. 2604.02268 null
2026-04-02 Model-Based Reinforcement Learning for Control under Time-Varying Dynamics Klemens Iten et.al. 2604.02260 null
2026-04-02 When to ASK: Uncertainty-Gated Language Assistance for Reinforcement Learning Juarez Monteiro et.al. 2604.02226 null
2026-04-02 Multi-Agent Video Recommenders: Evolution, Patterns, and Open Challenges Srivaths Ranganathan et.al. 2604.02211 null
2026-04-02 What can be computed in average anonymous networks? Joel Rybicki et.al. 2604.02192 null
2026-04-02 Entropic crystallization of geometrically frustrated magnets on 1/1 approximant Tsai-type quasicrystal Oscar Novat et.al. 2604.02180 null
2026-04-02 Dynamic resource coordination can increase grid hosting capacity to support more renewables, storage, and electrified load growth Vineet Jagadeesan Nair et.al. 2604.02170 null
2026-04-02 Auction-Based Online Policy Adaptation for Evolving Objectives Guruprerana Shabadi et.al. 2604.02151 null
2026-04-02 Stable and Efficient Algorithms for the Fermion Determinant Johann Ostmeyer et.al. 2604.02130 null
2026-04-02 The Bures metric and the quantum metric on the density space of a C*-algebra: the non-unital case Konrad Aguilar et.al. 2604.02117 null
2026-04-02 A new wavelet-based variational family with copula dependence structures Giovanni Piccirilli et.al. 2604.02116 null
2026-04-02 Lithium Droplet Transport in Tokamak Edge Plasmas A. Diaw et.al. 2604.02100 null
2026-04-02 Importance sampling for Bayesian inference: polynomial-dimension dependent error bounds Fabián González et.al. 2604.02094 null
2026-04-02 Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Yuhang Wu et.al. 2604.02091 null
2026-04-02 Ultrafast Ionization Dynamics Encoded in a Photoelectron Spin Torus Xiaodan Mao et.al. 2604.02062 null
2026-04-02 Efficient Auxiliary-Field Quantum Monte Carlo using Isometric Tensor Hypercontraction Maxine Luo et.al. 2604.02054 null
2026-04-02 Reinforcement Learning for Speculative Trading under Exploratory Framework Yun Zhao et.al. 2604.02035 null
2026-04-02 Systematic Analyses of Reinforcement Learning Controllers in Signalized Urban Corridors Xiaofei Song et.al. 2604.02025 null
2026-04-02 Bridging Discrete Planning and Continuous Execution for Redundant Robot Teng Yan et.al. 2604.02021 null
2026-04-02 Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning Rafael Pardinas et.al. 2604.02007 null
2026-04-02 ProCeedRL: Process Critic with Exploratory Demonstration Reinforcement Learning for LLM Agentic Reasoning Jingyue Gao et.al. 2604.02006 null
2026-04-02 ATLAS and CMS measurements of the $t\bar{t}$ cross section, including off-shell and near threshold Baptiste Ravina et.al. 2604.01984 null
2026-04-02 Captioning Daily Activity Images in Early Childhood Education: Benchmark and Algorithm Sixing Li et.al. 2604.01941 null
2026-04-01 Embarrassingly Simple Self-Distillation Improves Code Generation Ruixiang Zhang et.al. 2604.01193 null
2026-04-01 Variational Dynamics of Open Quantum Spin Systems in Phase Space Jacopo Tosca et.al. 2604.01165 null
2026-04-01 Markov chain Monte Carlo for Bayesian inference of the non-conducting region in intra-atrial reentrant tachycardia Maarten Volkaerts et.al. 2604.01164 null
2026-04-01 Deep Reinforcement Learning for Robotic Manipulation under Distribution Shift with Bounded Extremum Seeking Shaifalee Saxena et.al. 2604.01142 null
2026-04-01 Multi-Agent LLM Governance for Safe Two-Timescale Reinforcement Learning in SDN-IoT Defense Saeid Jamshidi et.al. 2604.01127 null
2026-04-01 A comparison of Markov Chain Monte Carlo algorithms for Bayesian inference of constitutive models Aricia Rinkens et.al. 2604.01121 null
2026-04-01 BAT: Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control Donghoon Baek et.al. 2604.01064 null
2026-04-01 Mass Hierarchies Without Mixing: Abelian Froggatt-Nielsen Models with Uncharged Left-Handed Doublets Navid Ardakanian et.al. 2604.01055 null
2026-04-01 Adversarial Attacks in AI-Driven RAN Slicing: SLA Violations and Recovery Deemah H. Tashman et.al. 2604.01049 null
2026-04-01 Geometry-induced correlated noise in qLDPC syndrome extraction Angelo Di Bella et.al. 2604.01040 null
2026-04-01 Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding Yiheng Wang et.al. 2604.01002 null
2026-04-01 Focal plane wavefront control with model-based reinforcement learning Jalo Nousiainen et.al. 2604.00993 null
2026-04-01 Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization Ruijie Hao et.al. 2604.00977 null
2026-04-01 How uncertain are model predictions for the muon content of extensive air showers Sergey Ostapchenko et.al. 2604.00975 null
2026-04-01 Quantum Statistical Bootstrap Yongkai Chen et.al. 2604.00951 null
2026-04-01 Policy Improvement Reinforcement Learning Huaiyang Wang et.al. 2604.00860 null
2026-04-01 Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation Shuang Li et.al. 2604.00849 null
2026-04-01 Bridging RL and MPC for mixed-integer optimal control with application to Formula 1 race strategies Joschua Wüthrich et.al. 2604.00826 null
2026-04-01 Stable Determinant Monte Carlo Simulations at Large Inverse Temperature $β$ Thomas Luu et.al. 2604.00815 null
2026-04-01 RefineRL: Advancing Competitive Programming with Self-Refinement Reinforcement Learning Shaopeng Fu et.al. 2604.00790 null
2026-04-01 Finding Low Star Discrepancy 3D Kronecker Point Sets Using Algorithm Configuration Techniques Imène Ait Abderrahim et.al. 2604.00786 null
2026-04-01 LangMARL: Natural Language Multi-Agent Reinforcement Learning Huaiyuan Yao et.al. 2604.00722 null
2026-04-01 Learning to Hint for Reinforcement Learning Yu Xia et.al. 2604.00698 null
2026-04-01 TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning Soumya Shamarao Jahagirdar et.al. 2604.00696 null
2026-04-01 Full-Gradient Successor Feature Representations Ritish Shrirao et.al. 2604.00686 null
2026-04-01 Analytical Probabilistic Power Flow Approximation Using Invertible Neural Networks Weijie Xia et.al. 2604.00673 null
2026-04-01 Implementation and Workflows for INLA-Based Approximate Bayesian Structural Equation Modelling Haziq Jamil et.al. 2604.00671 null
2026-04-01 Quantum Algorithms for Gibbs Expectation of Non-log-concave and Heavy-tailed Distributions Xinmiao Li et.al. 2604.00656 null
2026-04-01 A Survey of On-Policy Distillation for Large Language Models Mingyang Song et.al. 2604.00626 null
2026-04-01 A Physical Imitation Learning Pipeline for Energy-Efficient Quadruped Locomotion Assisted by Parallel Elastic Joint Huyue Ma et.al. 2604.00611 null
2026-04-01 Single-Waveguide Multiple-Pinching-Antenna Systems: OMA versus NOMA Yanyu Cheng et.al. 2604.00588 null
2026-03-31 HapCompass: A Rotational Haptic Device for Contact-Rich Robotic Teleoperation Xiangshan Tan et.al. 2603.30042 null
2026-03-31 Hybrid Framework for Robotic Manipulation: Integrating Reinforcement Learning and Large Language Models Md Saad et.al. 2603.30022 null
2026-03-31 Phyelds: A Pythonic Framework for Aggregate Computing Gianluca Aguzzi et.al. 2603.29999 null
2026-03-31 The Hollyfeld Gambit in Astrophysics Benne Holwerda et.al. 2603.29964 null
2026-03-31 GreenFLag: A Green Agentic Approach for Energy-Efficient Federated Learning Theodora Panagea et.al. 2603.29933 null
2026-03-31 C-TRAIL: A Commonsense World Framework for Trajectory Planning in Autonomous Driving Zhihong Cui et.al. 2603.29908 null
2026-03-31 Penalized GMM Framework for Inference on Functionals of Nonparametric Instrumental Variable Estimators Edvard Bakhitov et.al. 2603.29889 null
2026-03-31 ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training Rui Ai et.al. 2603.29871 null
2026-03-31 Bayesian methods for the identification of model parameters for water transport in porous media Paola Stolfi et.al. 2603.29859 null
2026-03-31 An Output Feedback Q-learning Algorithm for Optimal Control of Nonlinear Systems with Koopman Linear Embedding Victor G. Lopez et.al. 2603.29858 null
2026-03-31 Friends, Foes, and First Authors: A Game Theory Model of How Power Plays Rewrite Academic Co-Authorship Networks Amit Bengal et.al. 2603.29834 null
2026-03-31 Generalizing Output-Feedback Covariance Steering to Incorporate Non-Orthogonal Estimation Errors Daniel C. Qi et.al. 2603.29753 null
2026-03-31 Reinforced Reasoning for End-to-End Retrosynthetic Planning Chenyang Zuo et.al. 2603.29723 null
2026-03-31 6GAgentGym: Tool Use, Data Synthesis, and Agentic Learning for Network Management Jiao Chen et.al. 2603.29656 null
2026-03-31 ASI-Evolve: AI Accelerates AI Weixian Xu et.al. 2603.29640 null
2026-03-31 Learning Diagnostic Reasoning for Decision Support in Toxicology Nico Oberländer et.al. 2603.29608 null
2026-03-31 FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration Qiyao Wang et.al. 2603.29557 null
2026-03-31 GraSP-STL: A Graph-Based Framework for Zero-Shot Signal Temporal Logic Planning via Offline Goal-Conditioned Reinforcement Learning Ancheng Hou et.al. 2603.29533 null
2026-03-31 Communication Outage-Resistant UUV State Estimation: A Variational History Distillation Approach Shuyue Li et.al. 2603.29512 null
2026-03-31 Target-Aligned Reinforcement Learning Leonard S. Pleiss et.al. 2603.29501 null
2026-03-29 RTLSeek: Boosting the LLM-Based RTL Generation with Multi-Stage Diversity-Oriented Reinforcement Learning Xinyu Zhang et.al. 2603.27630 null
2026-03-29 DSevolve: Enabling Real-Time Adaptive Scheduling on Dynamic Shop Floor with LLM-Evolved Heuristic Portfolios Jin Huang et.al. 2603.27628 null
2026-03-29 Optimal resource allocation for maintaining system solvency Gaoyue Guo et.al. 2603.27622 null
2026-03-29 Secure Reinforcement Learning: On Model-Free Detection of Man in the Middle Attacks Rishi Rani et.al. 2603.27592 null
2026-03-29 Match or Replay: Self Imitating Proximal Policy Optimization Gaurav Chaudhary et.al. 2603.27515 null
2026-03-29 Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs Xuanpu Zhao et.al. 2603.27494 null
2026-03-29 Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning Feiding et.al. 2603.27482 null
2026-03-29 Driving Condition-Aware Multi-Agent Integrated Power and Thermal Management for Hybrid Electric Vehicles Hanghang Cui et.al. 2603.27471 null
2026-03-29 NeedleDB: A Generative-AI Based System for Accurate and Efficient Image Retrieval using Complex Natural Language Queries Mahdi Erfanian et.al. 2603.27464 null
2026-03-29 FlowRL: A Taxonomy and Modular Framework for Reinforcement Learning with Diffusion Policies Chenxiao Gao et.al. 2603.27450 null
2026-03-28 Agent-Driven Autonomous Reinforcement Learning Research: Iterative Policy Improvement for Quadruped Locomotion Nimesh Khandelwal et.al. 2603.27416 null
2026-03-28 Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning Dwait Bhatt et.al. 2603.27400 null
2026-03-28 Diagnosing Non-Markovian Observations in Reinforcement Learning via Prediction-Based Violation Scoring Naveen Mysore et.al. 2603.27389 null
2026-03-28 Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models Yuhang Han et.al. 2603.27375 null
2026-03-28 DRASTIC: A Dynamic Resource Allocation Framework over 6G Network Slicing in Task-aware Closed-Loop Tactile Internet Applications Narges Golmohammadi et.al. 2603.27364 null
2026-03-28 Online Inertia Tensor Identification for Non-Cooperative Spacecraft via Augmented UKF Batu Candan et.al. 2603.27361 null
2026-03-28 D-SPEAR: Dual-Stream Prioritized Experience Adaptive Replay for Stable Reinforcement Learninging Robotic Manipulation Yu Zhang et.al. 2603.27346 null
2026-03-28 Where-to-Learn: Analytical Policy Gradient Directed Exploration for On-Policy Robotic Reinforcement Learning Leixin Chang et.al. 2603.27317 null
2026-03-28 PyINLA: Fast Bayesian Inference for Latent Gaussian Models in Python Esmail Abdul Fattah et.al. 2603.27276 null
2026-03-28 Emergent Competition Between Dynamical Channels in Nonequilibrium Systems R. A. Dumer et.al. 2603.27256 null
2026-03-26 R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning Zirui Zhang et.al. 2603.25720 null
2026-03-26 Critical curve of two-matrix models $ABBA$, $A{B,A}B$ and $ABAB$ , Part I: Monte Carlo Carlos I. Pérez Sánchez et.al. 2603.25715 null
2026-03-26 Approximate Bayesian Inference for Structural Equation Models using Integrated Nested Laplace Approximations Haziq Jamil et.al. 2603.25690 null
2026-03-26 Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning Jai Bardhan et.al. 2603.25685 null
2026-03-26 LanteRn: Latent Visual Structured Reasoning André G. Viveiros et.al. 2603.25629 null
2026-03-26 A unified quantum computing quantum Monte Carlo framework through structured state preparation Giuseppe Buonaiuto et.al. 2603.25582 null
2026-03-26 Cooperative Deep Reinforcement Learning for Fair RIS Allocation Martin Mark Zan et.al. 2603.25572 null
2026-03-26 Multi-User Covert Communication in Spatially Heterogeneous Wireless Networks Jinyoung Lee et.al. 2603.25549 null
2026-03-26 Towards Embodied AI with MuscleMimic: Unlocking full-body musculoskeletal motor learning at scale Chengkun Li et.al. 2603.25544 null
2026-03-26 Sensitivity Analysis for Instrumental Variables Under Joint Relaxations of Monotonicity and Independence Pedro Picchetti et.al. 2603.25529 null
2026-03-26 A Representation Optimization Dichotomy, Lie-Algebraic Policy Optimization Sooraj KC et.al. 2603.25525 null
2026-03-26 Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning Jiajun Hu et.al. 2603.25464 null
2026-03-26 The Reward Function and the Least Cost Principle for Gravitation and other Laws of Physics Rubén Moreno-Bote et.al. 2603.25444 null
2026-03-26 The Symmetric Perceptron: a Teacher-Student Scenario Giovanni Catania et.al. 2603.25440 null
2026-03-26 TAPO: Translation Augmented Policy Optimization for Multilingual Mathematical Reasoning Xu Huang et.al. 2603.25419 null
2026-03-26 Modernising Reinforcement Learning-Based Navigation for Embodied Semantic Scene Graph Generation Roman Kueble et.al. 2603.25415 null
2026-03-26 Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models Xunguang Wang et.al. 2603.25412 null
2026-03-26 UMBRELLA: Uncertainty-aware Multi-robot Reactive Coordination under Dynamic Temporal Logic Tasks Qisheng Zhao et.al. 2603.25395 null
2026-03-26 Enabling ab initio geometry optimization of strongly correlated systems with transferable deep quantum Monte Carlo P. Bernát Szabó et.al. 2603.25381 null
2026-03-26 Integrating Deep RL and Bayesian Inference for ObjectNav in Mobile Robotics João Castelo-Branco et.al. 2603.25366 null
2026-03-25 Polynomial Speedup in Diffusion Models with the Multilevel Euler-Maruyama Method Arthur Jacot et.al. 2603.24594 null
2026-03-25 DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving Pengxuan Yang et.al. 2603.24587 null
2026-03-25 MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination Zhuo Li et.al. 2603.24579 null
2026-03-25 VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models Qijia He et.al. 2603.24575 null
2026-03-25 Completeness of Unbounded Best-First Minimax and Descent Minimax Quentin Cohen-Solal et.al. 2603.24572 null
2026-03-25 Probing Interacting Dark Sectors with upcoming Post-Reionization and Galaxy Surveys Rahul Shah et.al. 2603.24554 null
2026-03-25 Many-body perturbation theory for the nuclear equation of state up to fifth order C. Drischler et.al. 2603.24532 null
2026-03-25 No Single Metric Tells the Whole Story: A Multi-Dimensional Evaluation Framework for Uncertainty Attributions Emily Schiller et.al. 2603.24524 null
2026-03-25 Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations for Exascale Computing Systems Jeremy J. Williams et.al. 2603.24508 null
2026-03-25 Optimal control of infinite-dimensional dissipative systems Anthony Hastir et.al. 2603.24507 null
2026-03-25 Composer 2 Technical Report Cursor Reseach et.al. 2603.24477 null
2026-03-25 RKKY-dipolar Interactions and 3D Spin Supersolid on Stacked Triangular Lattice Ning Xi et.al. 2603.24446 null
2026-03-25 CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents Xiangru Jian et.al. 2603.24440 null
2026-03-25 A first study of strong isospin breaking effects in lattice QCD using truncated polynomials David Albandea et.al. 2603.24420 null
2026-03-25 Opinion-Driven Vaccination and Epidemic Dynamics on Heterogeneous Networks Anika Roy et.al. 2603.24403 null
2026-03-25 MolEvolve: LLM-Guided Evolutionary Search for Interpretable Molecular Optimization Xiangsen Chen et.al. 2603.24382 null
2026-03-25 Towards Reward Modeling for AI Tutors in Math Mistake Remediation Kseniia Petukhova et.al. 2603.24375 null
2026-03-25 Improving Lean4 Autoformalization via Cycle Consistency Fine-tuning Arsen Shebzukhov et.al. 2603.24372 null
2026-03-25 CoordLight: Learning Decentralized Coordination for Network-Wide Traffic Signal Control Yifeng Zhang et.al. 2603.24366 null
2026-03-25 LATS: Large Language Model Assisted Teacher-Student Framework for Multi-Agent Reinforcement Learning in Traffic Signal Control Yifeng Zhang et.al. 2603.24361 null
2026-03-25 Effects of the initial-state geometry on D-meson production in pp and pPb collisions R. Terra et.al. 2603.24344 null
2026-03-25 Strong-to-Weak Spontaneous Symmetry Breaking in a $(2+1)$ D Transverse-Field Ising Model under Decoherence Yi-Ming Ding et.al. 2603.24342 null
2026-03-25 Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning Dogan Urgun et.al. 2603.24324 null
2026-03-25 Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions Shiqin Wang et.al. 2603.24322 null
2026-03-25 Universal Quantum Suppression in Frustrated Ising Magnets across the Quasi-1D to 2D Crossover via Quantum Annealing Kumar Ghosh et.al. 2603.24311 null
2026-03-25 C-STEP: Continuous Space-Time Empowerment for Physics-informed Safe Reinforcement Learning of Mobile Agents Guihlerme Daubt et.al. 2603.24241 null
2026-03-25 Decentralized End-to-End Multi-AAV Pursuit Using Predictive Spatio-Temporal Observation via Deep Reinforcement Learning Yude Li et.al. 2603.24238 null
2026-03-25 Diffusion coefficients of multi-principal element alloys from first principles Damien K. J. Lee et.al. 2603.24228 null
2026-03-25 SumRank: Aligning Summarization Models for Long-Document Listwise Reranking Jincheng Feng et.al. 2603.24204 null
2026-03-25 A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula Cansu Sancaktar et.al. 2603.24202 null
2026-03-25 RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution Yushuai Song et.al. 2603.24198 null
2026-03-25 Optimized control protocols for stable skyrmion creation using deep reinforcement learning Ji Seok Song et.al. 2603.24177 null
2026-03-25 CarePilot: A Multi-Agent Framework for Long-Horizon Computer Task Automation in Healthcare Akash Ghosh et.al. 2603.24157 null
2026-03-25 A Longitudinal Analysis of the CEC Single-Objective Competitions (2010-2024) and Implications for Variational Quantum Optimization Vojtěch Novák et.al. 2603.24140 null
2026-03-25 Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection Zhanhe Lei et.al. 2603.24139 null
2026-03-25 Equivariant Filter Transformations for Consistent and Efficient Visual–Inertial Navigation Chungeng Tian et.al. 2603.24130 null
2026-03-25 Likelihood hacking in probabilistic program synthesis Jacek Karwowski et.al. 2603.24126 null
2026-03-24 UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation Jie Liu et.al. 2603.23500 null
2026-03-24 WildWorld: A Large-Scale Dataset for Dynamic World Modeling with Actions and Explicit State toward Generative ARPG Zhen Li et.al. 2603.23497 null
2026-03-24 End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions Zakaria Mhammedi et.al. 2603.23461 null
2026-03-24 Exact analytical PGSE signal for diffusion confined to a cylindrical surface using a spectral Laplacian formalism Erick J Canales-Rodríguez et.al. 2603.23421 null
2026-03-24 SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling Yiqi Zhang et.al. 2603.23414 null
2026-03-24 Kinetic Langevin Splitting Schemes for Constrained Sampling Neil K. Chada et.al. 2603.23397 null
2026-03-24 A Joint Reinforcement Learning Scheduling and Compression Framework for Teleoperated Driving Giacomo Avanzi et.al. 2603.23387 null
2026-03-24 Off-Policy Value-Based Reinforcement Learning for Large Language Models Peng-Yuan Wang et.al. 2603.23355 null
2026-03-24 Orbit-Level Stretching in Cubic Fourier-Galerkin Navier-Stokes: Sharp Incidence, Spectral Decay, and a Continuation Criterion Oleg Kiriukhin et.al. 2603.23293 null
2026-03-24 Learning Multi-Agent Local Collision-Avoidance for Collaborative Carrying tasks with Coupled Quadrupedal Robots Francesca Bray et.al. 2603.23278 null
2026-03-24 A Learning Method with Gap-Aware Generation for Heterogeneous DAG Scheduling Ruisong Zhou et.al. 2603.23249 null
2026-03-24 Semi-cosmographic constraints on decaying dark matter and dynamical dark energy: DESI DR2 BAO and 21\, cm intensity-mapping forecasts Mohit Yadav et.al. 2603.23247 null
2026-03-24 Neural ODE and SDE Models for Adaptation and Planning in Model-Based Reinforcement Learning Chao Han et.al. 2603.23245 null
2026-03-24 GEM: Guided Expectation-Maximization for Behavior-Normalized Candidate Action Selection in Offline RL Haoyu Wang et.al. 2603.23232 null
2026-03-24 Joint Task Orchestration and Resource Optimization for SC3 Closed Loop in 6G Networks Xinran Fang et.al. 2603.23217 null
2026-03-24 Between Resolution Collapse and Variance Inflation: Weighted Conformal Anomaly Detection in Low-Data Regimes Oliver Hennhöfer et.al. 2603.23205 null
2026-03-24 ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment Hao Wang et.al. 2603.23184 null
2026-03-24 Path Planning and Reinforcement Learning-Driven Control of On-Orbit Free-Flying Multi-Arm Robots Álvaro Belmonte-Baeza et.al. 2603.23182 null
2026-03-24 Optimal Control of Switched Systems Governed by Logical Switching Dynamics Xiao Zhang et.al. 2603.23131 null
2026-03-24 Fault-Tolerant Design and Multi-Objective Model Checking for Real-Time Deep Reinforcement Learning Systems Guoxin Su et.al. 2603.23113 null
2026-03-24 SpecXMaster Technical Report Yutang Ge et.al. 2603.23101 null
2026-03-24 Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards Orhun Buğra Baran et.al. 2603.23086 null
2026-03-24 MedCausalX: Adaptive Causal Reasoning with Self-Reflection for Trustworthy Medical Vision-Language Models Jianxin Lin et.al. 2603.23085 null
2026-03-24 Minimizing Material Waste in Additive Manufacturing through Online Reel Assignment Ilayda Celenk et.al. 2603.23042 null
2026-03-24 Fine-tuning of universal machine-learning interatomic potentials for 2D high-entropy alloys Chun Zhou et.al. 2603.23029 null
2026-03-24 A Top-Down Scale Approach for Multiscale Geographically and Temporally Weighted Regression Ghislain Geniaux et.al. 2603.22990 null
2026-03-24 Cooperation in Public Goods Games over Uniform Random Hypergraphs with Game Transitions Nankun Wei et.al. 2603.22967 null
2026-03-24 From Morality Installation in LLMs to LLMs in Morality-as-a-System Gunter Bombaerts et.al. 2603.22944 null
2026-03-24 Quality Over Clicks: Intrinsic Quality-Driven Iterative Reinforcement Learning for Cold-Start E-Commerce Query Suggestion Qi Sun et.al. 2603.22922 null
2026-03-24 EVA: Efficient Reinforcement Learning for End-to-End Video Agent Yaolun Zhang et.al. 2603.22918 null
2026-03-24 VLGOR: Visual-Language Knowledge Guided Offline Reinforcement Learning for Generalizable Agents Pengsen Liu et.al. 2603.22892 null
2026-03-24 Portfolio Optimization under Recursive Utility via Reinforcement Learning Minkey Chang et.al. 2603.22880 null
2026-03-24 Confidence Calibration under Ambiguous Ground Truth Linwei Tao et.al. 2603.22879 null
2026-03-24 Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models Ruixing Jin et.al. 2603.22876 null
2026-03-24 DecompGrind: A Decomposition Framework for Robotic Grinding via Cutting-Surface Planning and Contact-Force Adaptation Shunsuke Araki et.al. 2603.22859 null
2026-03-24 Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought Yunheng Li et.al. 2603.22847 null
2026-03-24 CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action Models Youzhi Liu et.al. 2603.22846 null
2026-03-24 Approximating the Shapley Value of Minimum Cost Spanning Tree Games: An FPRAS for Saving Games Takumi Jimbo et.al. 2603.22843 null
2026-03-23 Decoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration Zakaria Mhammedi et.al. 2603.22273 null
2026-03-23 TiCo: Time-Controllable Training for Spoken Dialogue Models Kai-Wei Chang et.al. 2603.22267 null
2026-03-23 DexDrummer: In-Hand, Contact-Rich, and Long-Horizon Dexterous Robot Drumming Hung-Chieh Fang et.al. 2603.22263 null
2026-03-23 Emergent relativistic symmetry from interacting fermions on the honeycomb bilayer Zi Hong Liu et.al. 2603.22259 null
2026-03-23 SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation Sashuai Zhou et.al. 2603.22228 null
2026-03-23 Make Tracking Easy: Neural Motion Retargeting for Humanoid Whole-body Control Qingrui Zhao et.al. 2603.22201 null
2026-03-23 Generalized Sequential Monte Carlo Sampling for Redistricting Simulation Philip O’Sullivan et.al. 2603.22188 null
2026-03-23 Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement Junrong Guo et.al. 2603.22187 null
2026-03-23 Cross-Modal Reinforcement Learning for Navigation with Degraded Depth Measurements Omkar Sawant et.al. 2603.22182 null
2026-03-23 Closed-Loop Verbal Reinforcement Learning for Task-Level Robotic Planning Dmitrii Plotnikov et.al. 2603.22169 null
2026-03-23 From Singleton Obstacles to Clutter: Translation Invariant Compositional Avoid Sets Prashant Solanki et.al. 2603.22146 null
2026-03-23 Adsorption energies and decomposition barrier heights for ethylene carbonate on the surface of lithium from cluster-based quantum chemistry Ethan A. Vo et.al. 2603.22139 null
2026-03-23 On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation Kexin Huang et.al. 2603.22117 null
2026-03-23 Lattice study of the critical bubble in $\mathrm{SU(8)}$ deconfinement transition Kari Rummukainen et.al. 2603.22088 null
2026-03-23 A Context Engineering Framework for Improving Enterprise AI Agents based on Digital-Twin MDP Xi Yang et.al. 2603.22083 null
2026-03-23 MEVIUS2: Practical Open-Source Quadruped Robot with Sheet Metal Welding and Multimodal Perception Kento Kawaharazuka et.al. 2603.22031 null
2026-03-23 Tuning Real-World Image Restoration at Inference: A Test-Time Scaling Paradigm for Flow Matching Models Purui Bai et.al. 2603.22027 null
2026-03-23 Here, there and everywhere: state-dependent time-inconsistent stochastic control Dylan Possamaï et.al. 2603.22022 null
2026-03-23 TREX: Trajectory Explanations for Multi-Objective Reinforcement Learning Dilina Rajapakse et.al. 2603.21988 null
2026-03-23 Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe Xixi Wu et.al. 2603.21972 null
2026-03-22 Learning to Optimize Joint Source and RIS-assisted Channel Encoding for Multi-User Semantic Communication Systems Haidong Wang et.al. 2603.21097 null
2026-03-22 DRL-driven Online Optimization for Joint Traffic Reshaping and Channel Reconfiguration in RIS-assisted Semantic NOMA Communications Songhan Zhao et.al. 2603.21093 null
2026-03-22 Approximate Dynamic Programming for Degradation-aware Market Participation of Battery Energy Storage Systems: Bridging Market and Degradation Timescales Flemming Holtorf et.al. 2603.21089 null
2026-03-22 Neural Inference Functions for Margins for Time Series Copula Models Daniel Fynn et.al. 2603.21075 null
2026-03-22 LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning Jianing Wang et.al. 2603.21065 null
2026-03-22 OrbitStream: Training-Free Adaptive 360-degree Video Streaming via Semantic Potential Fields Aizierjiang Aiersilan et.al. 2603.20999 null
2026-03-22 The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes Benedikt Hornig et.al. 2603.20994 null
2026-03-21 Cyber Deception for Mission Surveillance via Hypergame-Theoretic Deep Reinforcement Learning Zelin Wan et.al. 2603.20981 null
2026-03-21 Adjoint DSMC Method for Spatially Inhomogeneous Boltzmann Equation with General Boundary Conditions Russel Caflisch et.al. 2603.20946 null
2026-03-21 Deep Adaptive Rate Allocation in Volatile Heterogeneous Wireless Networks Gregorio Maglione et.al. 2603.20926 null
2026-03-21 Modeling surface radiation of rotating neutron stars with Monk-NS Wenda Zhang et.al. 2603.20870 null
2026-03-21 From Photons to Electrons: Accelerated Materials Discovery via Random Libraries and Automated Scanning Transmission Electron Microscopy Boris Slautin et.al. 2603.20858 null
2026-03-21 Karhunen-Loève Expansion for Fluid Antenna Systems: Information-Theoretic Optimal Channel Compression and Outage Analysis Tuo Wu et.al. 2603.20841 null
2026-03-21 EruDiff: Refactoring Knowledge in Diffusion Models for Advanced Text-to-Image Synthesis Xiefan Guo et.al. 2603.20828 null
2026-03-21 RLVR Training of LLMs Does Not Improve Thinking Ability for General QA: Evaluation Method and a Simple Solution Kaiyuan Li et.al. 2603.20799 null
2026-03-21 Enhanced Direction-Sensing Methods and Performance Analysis in Low-Altitude Wireless Network via a Rotation Antenna Array Jinbing Jiang et.al. 2603.20784 null
2026-03-21 Ordinal Patterns Based Testing of Spatial Independence in Irregular Spatial Structures Giorgio Micali et.al. 2603.20783 null
2026-03-21 A reliability-aware randomized simheuristic for the team orienteering problem with stochastic travel times Michele Circelli et.al. 2603.20766 null
2026-03-21 4D Fresnel Space-Time Modulation for Near-Field ELAA: Kinematic Multiplexing and O(N log N) Precoding at Sub-THz Frequencies Rahul Gulia et.al. 2603.20762 null
2026-03-21 Role of interstitial $s$ orbital in a model of infinite-layer nickelates Yan Peng et.al. 2603.20705 null
2026-03-20 Detecting the 3D Ising model phase transition with a ground-state-trained autoencoder Ahmed Abuali et.al. 2603.20157 null
2026-03-20 AGILE: A Comprehensive Workflow for Humanoid Loco-Manipulation Learning Huihua Zhao et.al. 2603.20147 null
2026-03-20 Triple/Double-Debiased Lasso Denis Chetverikov et.al. 2603.20134 null
2026-03-20 Analyzing Decoders for Quantum Error Correction Abtin Molavi et.al. 2603.20127 null
2026-03-20 Chain-of-Adaptation: Surgical Vision-Language Adaptation with Reinforcement Learning Jiajie Li et.al. 2603.20116 null
2026-03-20 Spectral Alignment in Forward-Backward Representations via Temporal Abstraction Seyed Mahdi B. Azad et.al. 2603.20103 null
2026-03-20 Fine-tuning Timeseries Predictors Using Reinforcement Learning Hugo Cazaux et.al. 2603.20063 null
2026-03-20 Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs Wenjian Zhang et.al. 2603.20046 null
2026-03-20 Q-approximation of operating characteristics of clinical trial designs Susanna Gentile et.al. 2603.20022 null
2026-03-20 ReViSQL: Achieving Human-Level Text-to-SQL Yuxuan Zhu et.al. 2603.20004 null
2026-03-20 Monte Carlo conformal prediction for quantifying uncertainty in radio galaxy classification under ambiguous ground truth Alex Walls et.al. 2603.20000 null
2026-03-20 Goal-Oriented Framework for Optical Flow-based Multi-User Multi-Task Video Transmission Yujie Xu et.al. 2603.19995 null
2026-03-20 Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States Yurun Yuan et.al. 2603.19987 null
2026-03-20 Interpreting Reinforcement Learning Model Behavior via Koopman with Control William T. Redman et.al. 2603.19968 null
2026-03-20 GustPilot: A Hierarchical DRL-INDI Framework for Wind-Resilient Quadrotor Navigation Amir Atef Habel et.al. 2603.19966 null
2026-03-20 Quantum Fisher Information as a Probe of Critical Scaling in Frustrated Magnets: Signatures from Kagome Quantum Spin Liquid Zhengbang Zhou et.al. 2603.19951 null
2026-03-20 Coupled cluster theory for positron binding in anions and polyatomic molecules Rosario R. Riso et.al. 2603.19948 null
2026-03-20 SAGE: Sustainable Agent-Guided Expert-tuning for Culturally Attuned Translation in Low-Resource Southeast Asia Zhixiang Lu et.al. 2603.19931 null
2026-03-20 Robust Beam Codebooks for mmWave/THz Systems: Toward a Stochastic RL Approach Anouar Nechi et.al. 2603.19930 null
2026-03-20 Learning Adaptive Parameter Policies for Nonlinear Bayesian Filtering Ondrej Straka et.al. 2603.19910 null
2026-03-20 Infinite-dimensional spherical-radial decomposition for probabilistic functions, with application to constrained optimal control and Gaussian process regression Kewei Wang et.al. 2603.19907 null
2026-03-20 What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time Dong Yan et.al. 2603.19880 null
2026-03-20 NASimJax: GPU-Accelerated Policy Learning Framework for Penetration Testing Raphael Simon et.al. 2603.19864 null
2026-03-20 FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization Chiyu Ma et.al. 2603.19835 null
2026-03-20 Generalized Task-Driven Design of Soft Robots via Reduced-Order FEM-based Surrogate Modeling Yao Yao et.al. 2603.19794 null
2026-03-20 Uniform Maximum Projection Designs for Computer Experiments Miroslav Vořechovský et.al. 2603.19778 null
2026-03-20 ReLi3D: Relightable Multi-view 3D Reconstruction with Disentangled Illumination Jan-Niklas Dihlmann et.al. 2603.19753 null
2026-03-19 OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards Zehao Li et.al. 2603.19191 null
2026-03-19 Markov Potential Game and Multi-Agent Reinforcement Learning for Autonomous Driving Huiwen Yan et.al. 2603.19188 null
2026-03-19 Box Maze: A Process-Control Architecture for Reliable LLM Reasoning Zou Qiang et.al. 2603.19182 null
2026-03-19 VEPO: Variable Entropy Policy Optimization for Low-Resource Language Foundation Models Chonghan Liu et.al. 2603.19152 null
2026-03-19 Hierarchical Latent Structure Learning through Online Inference Ines Aitsahalia et.al. 2603.19139 null
2026-03-19 Adaptive Regime-Aware Stock Price Prediction Using Autoencoder-Gated Dual Node Transformers with Reinforcement Learning Control Mohammad Al Ridhawi et.al. 2603.19136 null
2026-03-19 Variational and Annealing-Based Approaches to Quantum Combinatorial Optimization Hala Hawashin et.al. 2603.19117 null
2026-03-19 Articulated-Body Dynamics Network: Dynamics-Grounded Prior for Robot Learning Sangwoo Shin et.al. 2603.19078 null
2026-03-19 Probabilistic multivariate statistical process control via kernel parameter uncertainty propagation Zina-Sabrina Duma et.al. 2603.19055 null
2026-03-19 MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language Models Chenyang Gu et.al. 2603.19044 null
2026-03-19 Computation of thermal entropy for the doped Hubbard Model Yu-Feng Song et.al. 2603.18998 null
2026-03-19 CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Think Zening Sun et.al. 2603.18991 null
2026-03-19 Distributed lag non-linear models with spatial effect modification using Laplacian P-splines Sara Rutten et.al. 2603.18990 null
2026-03-19 Maximum-Entropy Exploration with Future State-Action Visitation Measures Adrien Bolland et.al. 2603.18965 null
2026-03-19 Context Bootstrapped Reinforcement Learning Saaket Agashe et.al. 2603.18953 null
2026-03-19 Navigating complex phase diagrams in soft matter systems Michael Wassermair et.al. 2603.18918 null
2026-03-19 Safety-Guaranteed Imitation Learning from Nonlinear Model Predictive Control for Spacecraft Close Proximity Operations Alexander Meinert et.al. 2603.18910 null
2026-03-19 MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model Youngwan Lee et.al. 2603.18892 null
2026-03-19 Reasoning over mathematical objects: on-policy reward modeling and test time aggregation Pranjal Aggarwal et.al. 2603.18886 null
2026-03-19 Bridging Network Fragmentation: A Semantic-Augmented DRL Framework for UAV-aided VANETs Gaoxiang Cao et.al. 2603.18871 null
2026-03-19 RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models Xiao Feng et.al. 2603.18859 null
2026-03-19 Learn for Variation: Variationally Guided AAV Trajectory Learning in Differentiable Environments Xiucheng Wang et.al. 2603.18853 null
2026-03-19 Preconditioning Hamiltonian Monte Carlo by minimizing Fisher Divergence Adrian Seyboldt et.al. 2603.18845 null
2026-03-19 ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents Hao Zhang et.al. 2603.18815 null
2026-03-19 V-Dreamer: Automating Robotic Simulation and Trajectory Synthesis via Video Generation Priors Songjia He et.al. 2603.18811 null
2026-03-19 Mi:dm K 2.5 Pro KT Tech innovation Group et.al. 2603.18788 null
2026-03-19 ViTac-Tracing: Visual-Tactile Imitation Learning of Deformable Object Tracing Yongqiang Zhao et.al. 2603.18784 null
2026-03-19 Automatic Configuration of LLM Post-Training Pipelines Channe Chwa et.al. 2603.18773 null
2026-03-19 NeuroGame Transformer: Gibbs-Inspired Attention Driven by Game Theory and Statistical Physics Djamel Bouchaffra et.al. 2603.18761 null
2026-03-19 Memento-Skills: Let Agents Design Agents Huichi Zhou et.al. 2603.18743 null
2026-03-19 CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks Hao Wang et.al. 2603.18736 null
2026-03-19 Tuning polymer architecture for quasicrystal self-assembly D. J. Ratliff et.al. 2603.18694 null
2026-03-19 HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning Zhicong Lu et.al. 2603.18683 null
2026-03-19 Complexity of Auctions with Interdependence Patrick Loiseau et.al. 2603.18668 null
2026-03-19 Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning Haokun Zhao et.al. 2603.18662 null
2026-03-19 Balanced Thinking: Improving Chain of Thought Training in Vision Language Models Shaked Perek et.al. 2603.18656 null
2026-03-19 Learning to Self-Evolve Xiaoyin Chen et.al. 2603.18620 null
2026-03-18 Observational Signatures of Exact Black Hole Solutions in a Dark Matter Halo Azalbek Boltaev et.al. 2603.17986 null
2026-03-18 Unified Policy Value Decomposition for Rapid Adaptation Cristiano Capone et.al. 2603.17947 null
2026-03-18 Training Diffusion Language Models for Black-Box Optimization Zipeng Sun et.al. 2603.17919 null
2026-03-18 RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference Arpit Singh Gautam et.al. 2603.17891 null
2026-03-18 Operator-Theoretic Foundations and Policy Gradient Methods for General MDPs with Unbounded Costs Abhishek Gupta et.al. 2603.17875 null
2026-03-18 Two stroke Pumping Technique for Many-Body Systems Serge Galam et.al. 2603.17873 null
2026-03-18 Procedural Generation of Algorithm Discovery Tasks in Machine Learning Alexander D. Goldie et.al. 2603.17863 null
2026-03-18 Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control Zunzhe Zhang et.al. 2603.17834 null
2026-03-18 CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents Lintang Sutawika et.al. 2603.17829 null
2026-03-18 Federated Distributional Reinforcement Learning with Distributional Critic Regularization David Millard et.al. 2603.17820 null
2026-03-18 Process Supervision for Chain-of-Thought Reasoning via Monte Carlo Net Information Gain Corentin Royer et.al. 2603.17815 null
2026-03-18 Dropout Robustness and Cognitive Profiling of Transformer Models via Stochastic Inference Antônio Junior Alves Caiado et.al. 2603.17811 null
2026-03-18 EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards Ruixiang Wang et.al. 2603.17808 null
2026-03-18 Hamiltonian Monte Carlo enhanced by Exact Diagonalization Finn L. Temmen et.al. 2603.17788 null
2026-03-18 CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution Teng Pan et.al. 2603.17775 null
2026-03-18 Fast stabilizer state preparation via AI-optimized graph decimation Michael Doherty et.al. 2603.17743 null
2026-03-18 VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning Tianxing Zhou et.al. 2603.17720 null
2026-03-18 Machine Learning for Network Attacks Classification and Statistical Evaluation of Machine Learning for Network Attacks Classification and Adversarial Learning Methodologies for Synthetic Data Generation Iakovos-Christos Zarkadis et.al. 2603.17717 null
2026-03-18 Dielectric response and structural properties of finite-temperature electron liquids Chengliang Lin et.al. 2603.17699 null
2026-03-18 Flow Matching Policy with Entropy Regularization Ting Gao et.al. 2603.17685 null
2026-03-18 Interpreting Context-Aware Human Preferences for Multi-Objective Robot Navigation Tharun Sethuraman et.al. 2603.17510 null
2026-03-18 Efficient Soft Actor-Critic with LLM-Based Action-Level Guidance for Continuous Control Hao Ma et.al. 2603.17468 null
2026-03-18 A Full-Density Approach to Simulating Random Iteration Equations with Applications Wolfgang Hoegele et.al. 2603.17466 null
2026-03-18 AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization Dailan He et.al. 2603.17461 null
2026-03-18 CRE-T1 Preview Technical Report: Beyond Contrastive Learning for Reasoning-Intensive Retrieval Guangzhi Wang et.al. 2603.17387 null
2026-03-18 Efficient Exploration at Scale Seyed Mohammad Asghari et.al. 2603.17378 null
2026-03-18 EvoGuard: An Extensible Agentic RL-based Framework for Practical and Evolving AI-Generated Image Detection Chenyang Zhu et.al. 2603.17343 null
2026-03-18 A Progressive Visual-Logic-Aligned Framework for Ride-Hailing Adjudication Weiming Wu et.al. 2603.17328 null
2026-03-18 Empirical Likelihood Inference for Sen and Sen–Shorrocks–Thon Indices Sreelakshmi N et.al. 2603.17327 null
2026-03-18 ShuttleEnv: An Interactive Data-Driven RL Environment for Badminton Strategy Modeling Ang Li et.al. 2603.17324 null
2026-03-18 Physics-informed offline reinforcement learning eliminates catastrophic fuel waste in maritime routing Aniruddha Bora et.al. 2603.17319 null
2026-03-18 Virtual Polarization Modulation: Enabling CSI-Free DCO-OFDM over Dynamic OWC Channels Tian Cao et.al. 2603.17316 null
2026-03-18 Mean first escape times of Brownian motion on asymptotically hyperbolic and gas giant metric surfaces Jesse Gell-Redman et.al. 2603.17313 null
2026-03-18 Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress Yuelin Zhang et.al. 2603.17312 null
2026-03-18 Ruyi2.5 Technical Report Huan Song et.al. 2603.17311 null
2026-03-18 InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning Chengwei Wei et.al. 2603.17310 null
2026-03-18 ReLMXEL: Adaptive RL-Based Memory Controller with Explainable Energy and Latency Optimization Panuganti Chirag Sai et.al. 2603.17309 null
2026-03-18 Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations Haozheng Luo et.al. 2603.17305 null
2026-03-18 WINFlowNets: Warm-up Integrated Networks Training of Generative Flow Networks for Robotics and Machine Fault Adaptation Zahin Sufiyan et.al. 2603.17301 null
2026-03-18 Bayesian Scalar-on-Tensor Quantile Regression for Longitudinal Data on Alzheimer’s Disease Rongke Lyu et.al. 2603.17294 null
2026-03-17 Efficient Reasoning on the Edge Yelysei Bondarenko et.al. 2603.16867 null
2026-03-17 DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models Emily Yue-Ting Jia et.al. 2603.16860 null
2026-03-17 Long-Horizon Traffic Forecasting via Incident-Aware Conformal Spatio-Temporal Transformers Mayur Patil et.al. 2603.16857 null
2026-03-17 Lifting the fog - a case for non-reversible “lifted” Markov chains Gabriele Tartero et.al. 2603.16855 null
2026-03-17 Unifying Optimization and Dynamics to Parallelize Sequential Computation: A Guide to Parallel Newton Methods for Breaking Sequential Bottlenecks Xavier Gonzalez et.al. 2603.16850 null
2026-03-17 Stochastic Resetting Accelerates Policy Convergence in Reinforcement Learning Jello Zhou et.al. 2603.16842 null
2026-03-17 Typical models of the distribution system restoration process Arslan Ahmad et.al. 2603.16841 null
2026-03-17 Learning to Present: Inverse Specification Rewards for Agentic Slide Generation Karthik Ragunath Ananda Kumar et.al. 2603.16839 null
2026-03-17 Deep Reinforcement Learning-driven Edge Offloading for Latency-constrained XR pipelines Sourya Saha et.al. 2603.16823 null
2026-03-17 Anticipatory Planning for Multimodal AI Agents Yongyuan Liang et.al. 2603.16777 null
2026-03-17 Learning Whole-Body Control for a Salamander Robot Mengze Tian et.al. 2603.16683 null
2026-03-17 When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making Jun Liu et.al. 2603.16673 null
2026-03-17 What if Pinocchio Were a Reinforcement Learning Agent: A Normative End-to-End Pipeline Benoît Alcaraz et.al. 2603.16651 null
2026-03-17 Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLMReward Models Weijie Qiu et.al. 2603.16600 null
2026-03-17 When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective Zelin Zhang et.al. 2603.16578 null
2026-03-17 EmoLLM: Appraisal-Grounded Cognitive-Emotional Co-Reasoning in Large Language Models Yifei Zhang et.al. 2603.16553 null
2026-03-17 Rethinking Pose Refinement in 3D Gaussian Splatting under Pose Prior and Geometric Uncertainty Mangyu Kong et.al. 2603.16538 null
2026-03-17 Kamino: GPU-based Massively Parallel Simulation of Multi-Body Systems with Challenging Topologies Vassilios Tsounis et.al. 2603.16536 null
2026-03-17 Monte Carlo sampling from a projected entangled-pair state in simulations of quantum annealing in the three dimensional random Ising model Jacek Dziarmaga et.al. 2603.16509 null
2026-03-17 From the Inside Out: Progressive Distribution Refinement for Confidence Calibration Xizhong Yang et.al. 2603.16500 null
2026-03-17 TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas Ai Jian et.al. 2603.16448 null
2026-03-17 Via Negativa for AI Alignment: Why Negative Constraints Are Structurally Superior to Positive Preferences Quan Cheng et.al. 2603.16417 null
2026-03-17 PlotTwist: A Creative Plot Generation Framework with Small Language Models Abhinav Thorat et.al. 2603.16410 null
2026-03-17 Onboard MuJoCo-based Model Predictive Control for Shipboard Crane with Double-Pendulum Sway Suppression Oscar Pang et.al. 2603.16407 null
2026-03-17 Deep Reinforcement Learning-Assisted Automated Operator Portfolio for Constrained Multi-objective Optimization Shuai Shao et.al. 2603.16401 null
2026-03-17 Controlling Fish Schools via Reinforcement Learning of Virtual Fish Movement Yusuke Nishii et.al. 2603.16384 null
2026-03-17 Groups of invertible ideals of one-dimensional Prüfer domains as groups of integer-valued functions Dario Spirito et.al. 2603.16322 null
2026-03-17 Agile Interception of a Flying Target using Competitive Reinforcement Learning Timothée Gavin et.al. 2603.16279 null
2026-03-17 VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment Tengjiao Yin et.al. 2603.16271 null
2026-03-17 Resonant scattering at the center of the galaxy cluster PKS 0745-191 with XRISM Keita Tanaka et.al. 2603.16263 null
2026-03-17 Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models Junxin Wang et.al. 2603.16253 null
2026-03-17 Quantum Brownian Motion: proving that the Schmid transition belongs to the Berezinskii-Kosterlitz-Thouless universality class Francesco G. Capone et.al. 2603.16227 null
2026-03-17 Dual Consensus: Escaping from Spurious Majority in Unsupervised RLVR via Two-Stage Vote Mechanism Kaixuan Du et.al. 2603.16223 null
2026-03-17 Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning Yongyu Mu et.al. 2603.16206 null
2026-03-17 Reliable Reasoning in SVG-LLMs via Multi-Task Multi-Reward Reinforcement Learning Haomin Wang et.al. 2603.16189 null
2026-03-17 ECHO: Edge-Cloud Humanoid Orchestration for Language-to-Motion Control Haozhe Jia et.al. 2603.16188 null
2026-03-17 Enforcing Task-Specified Compliance Bounds for Humanoids via Anisotropic Lipschitz-Constrained Policies Zewen He et.al. 2603.16180 null
2026-03-17 Optimizing Density Functional Theory for Strain-Dependent Magnetic Properties of Monolayer MnBi $_2$Te$_4$ with Diffusion Monte Carlo Jeonghwan Ahn et.al. 2603.16162 null
2026-03-17 SQL-ASTRA: Alleviating Sparse Feedback in Agentic SQL via Column-Set Matching and Trajectory Aggregation Long Li et.al. 2603.16161 null
2026-03-17 Execution-Grounded Credit Assignment for GRPO in Code Generation Abhijit Kumar et.al. 2603.16158 null
2026-03-16 GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering Xincheng Shuai et.al. 2603.15616 null
2026-03-16 HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions Yukang Cao et.al. 2603.15612 null
2026-03-16 Code-A1: Adversarial Evolving of Code LLM and Test LLM via Reinforcement Learning Aozhe Wang et.al. 2603.15611 null
2026-03-16 From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation Yibin Liu et.al. 2603.15600 null
2026-03-16 Simulating the Open System Dynamics of Multiple Exchange-Only Qubits using Subspace Monte Carlo Tameem Albash et.al. 2603.15577 null
2026-03-16 Unbiased and Biased Variance-Reduced Forward-Reflected-Backward Splitting Methods for Stochastic Composite Inclusions Quoc Tran-Dinh et.al. 2603.15576 null
2026-03-16 High-Throughput Computational Exploration of MOFs for Short-Chain PFAS Removal Mengru Zhang et.al. 2603.15503 null
2026-03-16 Introduction to the artificial neural network-based variational Monte Carlo method William Freitas et.al. 2603.15460 null
2026-03-16 Why Quarks and Leptons Demand Different Symmetries: A Systematic Z3 Froggatt-Nielsen Analysis Navid Ardakanian et.al. 2603.15455 null
2026-03-16 A flexible method for estimating luminosity functions via Kernel Density Estimation - III. Extending to Multiple Flux-Limited Samples Zunli Yuan et.al. 2603.15441 null
2026-03-16 Deep Reinforcement Learning for Fano Hypersurfaces Marc Truter et.al. 2603.15437 null
2026-03-16 Listening to the Echo: User-Reaction Aware Policy Optimization via Scalar-Verbal Hybrid Reinforcement Learning Jing Ye et.al. 2603.15434 null
2026-03-16 Gym-V: A Unified Vision Environment System for Agentic Vision Research Fanqing Meng Lingxiao Du Jiawei Gu Jiaqi Liao Linjie Li Zijian Wu Xiangyan Liu Ziqi Zhao Mengkang Hu Yue Zhang Zichen Liu Jiaheng Zhang Michael Qizhe Shieh et.al. 2603.15432 null
2026-03-16 MA-VLCM: A Vision Language Critic Model for Value Estimation of Policies in Multi-Agent Team Settings Shahil Shaik et.al. 2603.15418 null
2026-03-16 Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities Vanshaj Khattar et.al. 2603.15417 null
2026-03-16 Fusian: Multi-LoRA Fusion for Fine-Grained Continuous MBTI Personality Control in Large Language Models Zehao Chen et.al. 2603.15405 null
2026-03-16 Using an SU(3)/U(2) Wigner Function to Represent Noisy Spin Ensembles Andrew Kolmer Forbes et.al. 2603.15387 null
2026-03-16 Trajectory-Diversity-Driven Robust Vision-and-Language Navigation Jiangyang Li et.al. 2603.15370 null
2026-03-16 Controlled Langevin Dynamics for Sampling of Feedforward Neural Networks Trained with Minibatches Alessandro Zambon et.al. 2603.15367 null
2026-03-16 NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation Tianshuai Hu et.al. 2603.15359 null
2026-03-14 Diffusion Reinforcement Learning via Centered Reward Distillation Yuanzhi Zhu et.al. 2603.14128 null
2026-03-14 Improving Visual Reasoning with Iterative Evidence Refinement Zeru Shi et.al. 2603.14117 null
2026-03-14 Maximin Robust Bayesian Experimental Design Hany Abdulsamad et.al. 2603.14094 null
2026-03-14 A Benchmark for Multi-Party Negotiation Games from Real Negotiation Data Leo Benac et.al. 2603.14066 null
2026-03-14 Amortizing Trajectory Diffusion with Keyed Drift Fields Gokul Puthumanaillam et.al. 2603.14056 null
2026-03-14 GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models Zhijie Wang et.al. 2603.14041 null
2026-03-14 Traffic and weather driven hybrid digital twin for bridge monitoring Phani Raja Bharath Balijepalli et.al. 2603.14028 null
2026-03-14 LLM-Guided Safe Reinforcement Learning for Energy System Topology Reconfiguration Zongyan Zhang et.al. 2603.14018 null
2026-03-14 Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models Haitao Jiang et.al. 2603.13985 null
2026-03-14 Chunk-Guided Q-Learning Gwanwoo Song et.al. 2603.13971 null
2026-03-14 LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement Chih-Ning Chen et.al. 2603.13952 null
2026-03-14 SmoothVLA: Aligning Vision-Language-Action Models with Physical Constraints via Intrinsic Smoothness Optimization Jiashun Li et.al. 2603.13925 null
2026-03-14 ATCC: Adaptive Concurrency Control for Unforeseen Agentic Transactions Weixing Zhou et.al. 2603.13906 null
2026-03-14 A note on the invariants of the $L$ -functions Jerzy Kaczorowski et.al. 2603.13889 null
2026-03-14 Path-conditioned Reinforcement Learning-based Local Planning for Long-Range Navigation Mateo Haro et.al. 2603.13888 null
2026-03-14 Mass spectrometry of $^{75}$Zn ground and isomeric states from in-trap decay of $^{75}$ Cu M. Müller et.al. 2603.13868 null
2026-03-14 APEX-Searcher: Augmenting LLMs’ Search Capabilities through Agentic Planning and Execution Kun Chen et.al. 2603.13853 null
2026-03-14 Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving Zhexi Lian et.al. 2603.13842 null
2026-03-14 NEP_MiniMax: An Approach for NEPs Based on Matrix-valued Minimax Approximations Chenkun Zhang et.al. 2603.13794 null
2026-03-14 Your Vision-Language-Action Model Already Has Attention Heads For Path Deviation Detection Jaehwan Jeong et.al. 2603.13782 null
2026-03-13 Visual-ERM: Reward Modeling for Visual Equivalence Ziyu Liu et.al. 2603.13224 null
2026-03-13 $π$, K, and p production in high-multiplicity pp collisions at $\sqrt{s} = 13$ TeV ALICE Collaboration et.al. 2603.13203 null
2026-03-13 Radiative return meets GVMD Pau Petit Rosàs et.al. 2603.13171 null
2026-03-13 PhaseJumps: fast computation of zeros from planar grid samples Antti Haimi et.al. 2603.13158 null
2026-03-13 Reinforcement Learning for Discounted and Ergodic Control of Diffusion Processes Erhan Bayraktar et.al. 2603.13155 null
2026-03-13 Determination of Nuclear PDFs using Markov Chain Monte Carlo Methods N. Derakhshanian et.al. 2603.13150 null
2026-03-13 Structural Parameters of the Globular Cluster M 15 M. V. Petkova et.al. 2603.13140 null
2026-03-13 Reweighted information inequalities Jonathan Niles-Weed et.al. 2603.13135 null
2026-03-13 Topo-R1: Detecting Topological Anomalies via Vision-Language Models Meilong Xu et.al. 2603.13054 null
2026-03-13 Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation Yifeng Liu et.al. 2603.13045 null
2026-03-13 OpenACMv2: An Accuracy-Constrained Co-Optimization Framework for Approximate DCiM Yiqi Zhou et.al. 2603.13042 null
2026-03-13 PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses Chenlong Yin et.al. 2603.13026 null
2026-03-13 ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning Bangjun Xiao et.al. 2603.13019 null
2026-03-13 Long-form RewardBench: Evaluating Reward Models for Long-form Generation Hui Huang et.al. 2603.12963 null
2026-03-13 Efficient Real-World Autonomous Racing via Attenuated Residual Policy Optimization Raphael Trumpp et.al. 2603.12960 null
2026-03-13 Thinking in Streaming Video Zikang Liu et.al. 2603.12938 null
2026-03-13 Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models David McAllister et.al. 2603.12893 null
2026-03-13 Enhanced Drug-drug Interaction Prediction Using Adaptive Knowledge Integration Pengfei Liu et.al. 2603.12885 null
2026-03-13 Test-time RL alignment exposes task familiarity artifacts in LLM benchmarks Kun Wang et.al. 2603.12875 null
2026-03-13 Weak Adversarial Neural Pushforward Method for Fractional Fokker-Planck Equations Andrew Qing He et.al. 2603.12869 null
2026-03-13 Beyond Imitation: Reinforcement Learning Fine-Tuning for Adaptive Diffusion Navigation Policies Junhe Sheng et.al. 2603.12868 null
2026-03-13 A basic model for high energy cosmic ray interactions Sergey Ostapchenko et.al. 2603.12863 null
2026-03-13 Optimal Stopping for Systems Driven by the Brownian Sheet Nacira Agram et.al. 2603.12853 null
2026-03-12 HumDex:Humanoid Dexterous Manipulation Made Easy Liang Heng et.al. 2603.12260 null
2026-03-12 DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning Yujie Wei et.al. 2603.12257 null
2026-03-12 Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing Baifeng Shi et.al. 2603.12254 null
2026-03-12 Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models Samy Jelassi et.al. 2603.12248 null
2026-03-12 Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation Xiangyu Zhao et.al. 2603.12247 null
2026-03-12 Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training Yixin Liu et.al. 2603.12246 null
2026-03-12 Separable neural architectures as a primitive for unified predictive and generative intelligence Reza T. Batley et.al. 2603.12244 null
2026-03-12 HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies Amber Xie et.al. 2603.12243 null
2026-03-12 Integrated Online Monitoring and Adaption of Process Model Predictive Controllers Samuel Mallick et.al. 2603.12187 null
2026-03-12 Operator Splitting, Policy Iteration, and Machine Learning for Stochastic Optimal Control Alain Bensoussan et.al. 2603.12167 null
2026-03-12 LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning Haiying Xu et.al. 2603.12166 null
2026-03-12 IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL Zhoujun Cheng et.al. 2603.12151 null
2026-03-12 Linking Perception, Confidence and Accuracy in MLLMs Yuetian Du et.al. 2603.12149 null
2026-03-12 EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next Ye Pan et.al. 2603.12147 null
2026-03-12 Automatic Generation of High-Performance RL Environments Seth Karten et.al. 2603.12145 null
2026-03-12 Increasing intelligence in AI agents can worsen collective outcomes Neil F. Johnson et.al. 2603.12129 null
2026-03-12 Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives Taeho Lee et.al. 2603.12110 null
2026-03-12 On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents Deyu Zou et.al. 2603.12109 null
2026-03-12 Wasserstein Gradient Flows for Batch Bayesian Optimal Experimental Design Louis Sharrock et.al. 2603.12102 null
2026-03-12 A Robust and Efficient Multi-Agent Reinforcement Learning Framework for Traffic Signal Control Sheng-You Huang et.al. 2603.12096 null
2026-03-12 Cross-Domain Policy Optimization via Bellman Consistency and Hybrid Critics Ming-Hong Chen et.al. 2603.12087 null
2026-03-12 Direct Boltzmann inversion method from particle configurations at arbitrary state points Olivier Coquand et.al. 2603.12081 null
2026-03-12 Topological Enhancement of Protein Kinetic Stability João NC Especial et.al. 2603.12053 null
2026-03-12 AGMARL-DKS: An Adaptive Graph-Enhanced Multi-Agent Reinforcement Learning for Dynamic Kubernetes Scheduling Hamed Hamzeh et.al. 2603.12031 null
2026-03-12 Sim-to-reality adaptation for Deep Reinforcement Learning applied to an underwater docking application Alaaeddine Chaarani et.al. 2603.12020 null
2026-03-12 Deep Learning-Based Metamodeling of Nonlinear Stochastic Dynamic Systems under Parametric and Predictive Uncertainty Haimiti Atila et.al. 2603.12012 null
2026-03-12 Learning Visuomotor Policy for Multi-Robot Laser Tag Game Kai Li et.al. 2603.11980 null
2026-03-12 FlexRec: Adapting LLM-based Recommenders for Flexible Needs via Reinforcement Learning Yijun Pan et.al. 2603.11901 null
2026-03-12 The price of decentralization in managing engineering systems through multi-agent reinforcement learning Prateek Bhustali et.al. 2603.11884 null
2026-03-12 Bielik-Minitron-7B: Compressing Large Language Models via Structured Pruning and Knowledge Distillation for the Polish Language Remigiusz Kinas et.al. 2603.11881 null
2026-03-12 Hybrid Human-Agent Social Dilemmas in Energy Markets Isuri Perera et.al. 2603.11834 null
2026-03-12 Towards High-Fidelity CAD Generation via LLM-Driven Program Generation and Text-Based B-Rep Primitive Grounding Jiahao Li et.al. 2603.11831 null
2026-03-12 RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset Yongzhong Wang et.al. 2603.11811 null
2026-03-12 BBN to Late-Time Acceleration in $f(T,\mathcal{L}_m)$ Gravity Sai Swagat Mishra et.al. 2603.11760 null
2026-03-12 Exploiting Expertise of Non-Expert and Diverse Agents in Social Bandit Learning: A Free Energy Approach Erfan Mirzaei et.al. 2603.11757 null
2026-03-12 Merger-driven buildup of the $M_{\rm BH}$ - $M_*$ relation bridging high-$z$ overmassive black holes with the local relation Takumi S. Tanaka et.al. 2603.11747 null
2026-03-11 Uncovering statistical structure in large-scale neural activity with Restricted Boltzmann Machines Nicolas Béreux et.al. 2603.11032 null
2026-03-11 Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge Mingyang Song et.al. 2603.11027 null
2026-03-11 Don’t Disregard the Data for Lack of a Likelihood: Bayesian Synthetic Likelihood for Enhanced Multilevel Network Meta-Regression Harlan Campbell et.al. 2603.11019 null
2026-03-11 MCMC Informed Neural Emulators for Uncertainty Quantification in Dynamical Systems Heikki Haario et.al. 2603.10987 null
2026-03-11 Learning Adaptive Force Control for Contact-Rich Sample Scraping with Heterogeneous Materials Cenk Cetin et.al. 2603.10979 null
2026-03-11 Contact Coverage-Guided Exploration for General-Purpose Dexterous Manipulation Zixuan Liu et.al. 2603.10971 null
2026-03-11 Generalized Reduced-Density-Matrix Quantum Monte Carlo Gives Access to More Zhiyan Wang et.al. 2603.10948 null
2026-03-11 Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control Yaswanth Chittepu et.al. 2603.10938 null
2026-03-11 Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment Fanqi Yu et.al. 2603.10929 null
2026-03-11 Ergodicity in reinforcement learning Dominik Baumann et.al. 2603.10895 null
2026-03-11 Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models Yixiu Mao et.al. 2603.10887 null
2026-03-11 Continuous Diffusion Transformers for Designing Synthetic Regulatory Elements Jonathan Liu et.al. 2603.10885 null
2026-03-11 RL-Augmented MPC for Non-Gaited Legged and Hybrid Locomotion Andrea Patrizi et.al. 2603.10878 null
2026-03-11 A Physics-Informed, Global-in-Time Neural Particle Method for the Spatially Homogeneous Landau Equation Minseok Kim et.al. 2603.10874 null
2026-03-11 $V_{0.5}$ : Generalist Value Model as a Prior for Sparse RL Rollouts Yi-Kai Zhang et.al. 2603.10848 null
2026-03-11 Towards Cold-Start Drafting and Continual Refining: A Value-Driven Memory Approach with Application to NPU Kernel Synthesis Yujie Zheng et.al. 2603.10846 null
2026-03-11 ReTabSyn: Realistic Tabular Data Synthesis via Reinforcement Learning Xiaofeng Lin et.al. 2603.10823 null
2026-03-11 Multilingual Reasoning Gym: Multilingual Scaling of Procedural Reasoning Environments Konstantin Dobler et.al. 2603.10793 null
2026-03-11 mAceReason-Math: A Dataset of High-Quality Multilingual Math Problems Ready For RLVR Konstantin Dobler et.al. 2603.10767 null
2026-03-11 Beyond Accuracy: Reliability and Uncertainty Estimation in Convolutional Neural Networks Sanne Ruijs et.al. 2603.10731 null
2026-03-11 ASTER: Attitude-aware Suspended-payload Quadrotor Traversal via Efficient Reinforcement Learning Dongcheng Cao et.al. 2603.10715 null
2026-03-11 MAVEN: A Meta-Reinforcement Learning Framework for Varying-Dynamics Expertise in Agile Quadrotor Maneuvers Jin Zhou et.al. 2603.10714 null
2026-03-11 Splat2Real: Novel-view Scaling for Physical AI with 3D Gaussian Splatting Hansol Lim et.al. 2603.10638 null
2026-03-11 Reinforcement Learning with Conditional Expectation Reward Changyi Xiao et.al. 2603.10624 null
2026-03-11 AdaClearGrasp: Learning Adaptive Clearing for Zero-Shot Robust Dexterous Grasping in Densely Cluttered Environments Zixuan Chen et.al. 2603.10616 null
2026-03-11 Maximum Inverse Sum Indeg Index of Trees and Unicyclic Graphs with Fixed Diameter Sunilkumar M. Hosamani et.al. 2603.10603 null
2026-03-11 Does LLM Alignment Really Need Diversity? An Empirical Study of Adapting RLVR Methods for Moral Reasoning Zhaowei Zhang et.al. 2603.10588 null
2026-03-11 Safety-critical Control Under Partial Observability: Reach-Avoid POMDP meets Belief Space Control Matti Vahs et.al. 2603.10572 null
2026-03-11 Adaptive RAN Slicing Control via Reward-Free Self-Finetuning Agents Yuanhao Li et.al. 2603.10564 null
2026-03-10 Probing Physics Beyond the Standard Model through Combined Analyses of Next-Generation Type Ia Supernova, CMB, and BAO Surveys Srinivasan Raghunathan et.al. 2603.09973 null
2026-03-10 Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization Xiaoyu Zhang et.al. 2603.09956 null
2026-03-10 When Learning Rates Go Wrong: Early Structural Signals in PPO Actor-Critic Alberto Fernández-Hernández et.al. 2603.09950 null
2026-03-10 Subspace decomposition with defect diffusion coefficient Dilini Kolombage et.al. 2603.09924 null
2026-03-10 Critical behavior of the thermal phase transition of U(1) lattice gauge systems Greta Sophie Reese et.al. 2603.09895 null
2026-03-10 Influencing LLM Multi-Agent Dialogue via Policy-Parameterized Prompts Hongbo Bo et.al. 2603.09890 null
2026-03-10 Robust Cooperative Localization in Featureless Environments: A Comparative Study of DCL, StCL, CCL, CI, and Standard-CL Nivand Khosravi et.al. 2603.09886 null
2026-03-10 Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning Yixin Zheng et.al. 2603.09882 null
2026-03-10 RecThinker: An Agentic Framework for Tool-Augmented Reasoning in Recommendation Haobo Zhang et.al. 2603.09843 null
2026-03-10 Good Reasoning Makes Good Demonstrations: Implicit Reasoning Quality Supervision via In-Context Reinforcement Learning Tiehua Mei et.al. 2603.09803 null
2026-03-10 Long-Run Conditional Value-at-Risk Reinforcement Learning Qixin Wang et.al. 2603.09734 null
2026-03-10 Event-by-Event Multiplicity Fluctuations in Heavy-Ion Collisions Using Modified HIJING Monte Carlo Generator Y. A. Rusak et.al. 2603.09732 null
2026-03-10 GSStream: 3D Gaussian Splatting based Volumetric Scene Streaming System Zhiye Tang et.al. 2603.09718 null
2026-03-10 ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning Davit Melikidze et.al. 2603.09692 null
2026-03-10 Analytic treatment of a polaron in a nonparabolic conduction band S. N. Klimin et.al. 2603.09609 null
2026-03-10 Phase diagram of 4D SU(3) Yang-Mills theory at $θ=π$ via imaginary theta simulations Akira Matsumoto et.al. 2603.09604 null
2026-03-10 Comprehensive structural and optical analysis of differently oriented Yb-implanted $β$-Ga$_2$O$_3$ Joanna Matulewicz et.al. 2603.09592 null
2026-03-10 Joint Bayesian analysis of soft and high- $p_\perp$ probes yields tighter constraints on QGP properties Marko Djordjevic et.al. 2603.09584 null
2026-03-10 An Optimal Control Approach To Transformer Training Kağan Akman et.al. 2603.09571 null
2026-03-10 ReTac-ACT: A State-Gated Vision-Tactile Fusion Transformer for Precision Assembly Minchi Ruan et.al. 2603.09565 null
2026-03-10 GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision Lang Sun et.al. 2603.09551 null
2026-03-10 NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models Ziyue Zhu et.al. 2603.09542 null
2026-03-10 Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization Ming Nie et.al. 2603.09538 null
2026-03-10 MORE-R1: Guiding LVLM for Multimodal Object-Entity Relation Extraction via Stepwise Reasoning with Reinforcement Learning Xiang Yuan et.al. 2603.09478 null
2026-03-10 Phase diagram and Ashkin-Teller universality in the classical square-lattice Heisenberg-compass model Yuchen Fan et.al. 2603.09469 null
2026-03-10 EvoDriveVLA: Evolving Autonomous Driving Vision-Language-Action Model via Collaborative Perception-Planning Distillation Jiajun Cao et.al. 2603.09465 null
2026-03-10 SEA-Nav: Efficient Policy Learning for Safe and Agile Quadruped Navigation in Cluttered Environments Shiyi Chen et.al. 2603.09460 null
2026-03-10 Beyond QED: Electroweak and hadronic extensions of McMule Sophie Kollatzsch et.al. 2603.09443 null
2026-03-10 Quantum spin ladder with ferromagnetic rungs in Bi $_2$CuO$_3$(SO$_4$ ) Rodolfo A. Rangel Hernandez et.al. 2603.09442 null
2026-03-10 From Weighting to Modeling: A Nonparametric Estimator for Off-Policy Evaluation Rong J. B. Zhu et.al. 2603.09436 null
2026-03-10 Impact of Markov Decision Process Design on Sim-to-Real Reinforcement Learning Tatjana Krau et.al. 2603.09427 null
2026-03-10 The framework to unify all complexity dichotomy theorems for Boolean tensor networks Mingji Xia et.al. 2603.09417 null
2026-03-10 Reward Prediction with Factorized World States Yijun Shen et.al. 2603.09400 null
2026-03-10 Beyond Fermi-II: Intermittent Particle Acceleration by Relativistic Turbulence in Astrophysical Plasmas Anton Dmytriiev et.al. 2603.09394 null
2026-03-09 Understanding thermal and quantum fluctuations in extended Kitaev-Yao-Lee spin-orbital model Jiefu Cen et.al. 2603.08710 null
2026-03-09 Agentic Critical Training Weize Liu et.al. 2603.08706 null
2026-03-09 How Far Can Unsupervised RLVR Scale LLM Training? Bingxiang He et.al. 2603.08660 null
2026-03-09 Embedding Classical Balance Control Principles in Reinforcement Learning for Humanoid Recovery Nehar Poddar et.al. 2603.08619 null
2026-03-09 Diff-Muscle: Efficient Learning for Musculoskeletal Robotic Table Tennis Wentao Zhao et.al. 2603.08617 null
2026-03-09 Online Learning in Semiparametric Econometric Models Xiaohong Chen et.al. 2603.08614 null
2026-03-09 Towards Batch-to-Streaming Deep Reinforcement Learning for Continuous Control Riccardo De Monte et.al. 2603.08588 null
2026-03-09 MetaWorld-X: Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation Yutong Shen et.al. 2603.08572 null
2026-03-09 RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback Xiaoying Zhang et.al. 2603.08561 null
2026-03-09 Impact of Connectivity on Laplacian Representations in Reinforcement Learning Tommaso Giorgi et.al. 2603.08558 null
2026-03-09 Evaluation of EMF Exposure to Throughput Ratio for Sustainable 5G Networks Dinh Long Trinh et.al. 2603.08549 null
2026-03-09 EquiBim: Learning Symmetry-Equivariant Policy for Bimanual Manipulation Zhiyuan Zhang et.al. 2603.08541 null
2026-03-09 Rethinking Strict Dissipativity for Economic MPC Mario Zanon et.al. 2603.08535 null
2026-03-09 Breaking the Bias Barrier in Concave Multi-Objective Reinforcement Learning Swetha Ganesh et.al. 2603.08518 null
2026-03-09 Oracle-Guided Soft Shielding for Safe Move Prediction in Chess Prajit T Rajendran et.al. 2603.08506 null
2026-03-09 LAR-MoE: Latent-Aligned Routing for Mixture of Experts in Robotic Imitation Learning Ariel Rodriguez et.al. 2603.08476 null
2026-03-09 Integrating Lagrangian Neural Networks into the Dyna Framework for Reinforcement Learning Shreya Das et.al. 2603.08468 null
2026-03-09 Reasoning as Compression: Unifying Budget Forcing via the Conditional Information Bottleneck Fabio Valerio Massoli et.al. 2603.08462 null
2026-03-09 Stochastic Loop Corrections to Belief Propagation for Tensor Network Contraction Gi Beom Sim et.al. 2603.08427 null
2026-03-09 Meta-RL with Shared Representations Enables Fast Adaptation in Energy Systems Théo Zangato et.al. 2603.08418 null
2026-03-08 Underwater Embodied Intelligence for Autonomous Robots: A Constraint-Coupled Perspective on Planning, Control, and Deployment Jingzehua Xu et.al. 2603.07393 null
2026-03-07 Learning to Reflect: Hierarchical Multi-Agent Reinforcement Learning for CSI-Free mmWave Beam-Focusing Hieu Le et.al. 2603.07370 null
2026-03-07 Neural Control and Learning of Simulated Hand Movements With an EMG-Based Closed-Loop Interface Balint K. Hodossy et.al. 2603.07364 null
2026-03-07 To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise Nouran Khallaf et.al. 2603.07330 null
2026-03-07 Extending gPET for Multi-Layer PET Simulation Satzhan Sitmukhambetov et.al. 2603.07327 null
2026-03-07 Adversarial Latent-State Training for Robust Policies in Partially Observable Domains Angad Singh Ahuja et.al. 2603.07313 null
2026-03-07 AutoResearch-RL: Perpetual Self-Evaluating Reinforcement Learning Agents for Autonomous Neural Architecture Discovery Nilesh Jain et.al. 2603.07300 null
2026-03-07 Adaptive Double-Booking Strategy for Outpatient Scheduling Using Multi-Objective Reinforcement Learning Ninda Nurseha Amalina et.al. 2603.07270 null
2026-03-07 Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving Jiazhuo Li et.al. 2603.07264 null
2026-03-07 Learning When to Cooperate Under Heterogeneous Goals Max Taylor-Davies et.al. 2603.07253 null
2026-03-07 Multi-parameter determination in the semilinear Helmholtz equation Long-Ling Du et.al. 2603.07247 null
2026-03-07 Reinforcement Learning for Vehicle-to-Grid Voltage Regulation: Single-Hub to Multi-Hub Coordination with Battery-Aware Constraints Jingbo Wang et.al. 2603.07237 null
2026-03-07 wDPO: Winsorized Direct Preference Optimization for Robust LLM Alignment Jilong Liu et.al. 2603.07211 null
2026-03-07 $\textbf{Re}^{2}$ : Unlocking LLM Reasoning via Reinforcement Learning with Re-solving Pinzheng Wang et.al. 2603.07197 null
2026-03-07 RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation Zixuan Chen et.al. 2603.07165 null
2026-03-07 Agentic Planning with Reasoning for Image Styling via Offline RL Subhojyoti Mukherjee et.al. 2603.07148 null
2026-03-07 Coexistence Regime and Thermal Crystallization in the cavity-mediated extended Bose-Hubbard Model Wei-Wei Wang et.al. 2603.07121 null
2026-03-07 Learning From Failures: Efficient Reinforcement Learning Control with Episodic Memory Chenyang Miao et.al. 2603.07110 null
2026-03-07 Ultra-Sharp Upright Photon Radiotherapy via Low Energy Extended Distance: An Alternative to FLASH for high flux Sources Lloyd E Kamole Ghomsi et.al. 2603.07103 null
2026-03-07 Parametric modal regression for right-censored positive responses Christian E. Galarza et.al. 2603.07099 null
2026-03-06 Boosting deep Reinforcement Learning using pretraining with Logical Options Zihan Ye et.al. 2603.06565 null
2026-03-06 EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking Fangrui Zhu et.al. 2603.06561 null
2026-03-06 Kinematically Coherent Multiphase Galactic Winds in Star-Forming Galaxies Revealed by Unified Radiative Transfer Modeling of UV Emission and Absorption Lines Zhihui Li et.al. 2603.06546 null
2026-03-06 Comment on: “Third-order corrections to the slow-roll expansion: Calculation and constraints with Planck, ACT, SPT, and BICEP/Keck [2025 PDU 47 101813]” Pierre Auclair et.al. 2603.06521 null
2026-03-06 On a PDE model for Learning in Stochastic Market Entry Games Esther Bou Dagher et.al. 2603.06514 null
2026-03-06 Balancing Efficiency and Feasibility: A Sensitivity Analysis of the Augmentation Parameter in the Finite Selection Model Safaa K. Kadhem et.al. 2603.06493 null
2026-03-06 Minimizers for boundary reactions: renormalized energy, location of singularities, and applications Xavier Cabre et.al. 2603.06435 null
2026-03-06 A Reference Architecture of Reinforcement Learning Frameworks Xiaoran Liu et.al. 2603.06413 null
2026-03-06 Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion Pengcheng Jiang et.al. 2603.06397 null
2026-03-06 OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis Yuxuan Fan et.al. 2603.06366 null
2026-03-06 From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty Azza Jenane et.al. 2603.06317 null
2026-03-06 Sparse probabilistic evaluation for treatment planning: a feasibility study in IMPT head & neck patients Jenneke I. de Jong et.al. 2603.06314 null
2026-03-06 Large Wave Direction Data Modeling Using Wrapped Spatial Gaussian Markov Random Fields Arnab Hazra et.al. 2603.06293 null
2026-03-06 Artificial Intelligence for Climate Adaptation: Reinforcement Learning for Climate Change-Resilient Transport Miguel Costa et.al. 2603.06278 null
2026-03-06 Synthetic Monitoring Environments for Reinforcement Learning Leonard Pleiss et.al. 2603.06252 null
2026-03-06 Star-based Navigation in the Outer Solar System Vittorio Franzese et.al. 2603.06247 null
2026-03-06 MAPO: Mixed Advantage Policy Optimization for Long-Horizon Multi-Turn Dialogue Naifan Zhang et.al. 2603.06194 null
2026-03-06 Optimizing 3D Diffusion Models for Medical Imaging via Multi-Scale Reward Learning Yueying Tian et.al. 2603.06173 null
2026-03-06 Dual-Agent Multiple-Model Reinforcement Learning for Event-Triggered Human-Robot Co-Adaptation in Decoupled Task Spaces Yaqi Li et.al. 2603.06163 null
2026-03-06 Policy Iteration Achieves Regularized Equilibrium under Time Inconsistency Yu-Jui Huang et.al. 2603.06145 null
2026-03-06 Partial Policy Gradients for RL in LLMs Puneet Mathur et.al. 2603.06138 null
2026-03-06 ChatShopBuddy: Towards Reliable Conversational Shopping Agents via Reinforcement Learning Yiruo Cheng et.al. 2603.06065 null
2026-03-06 Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models Canyu Chen et.al. 2603.06049 null
2026-03-05 RoboPocket: Improve Robot Policies Instantly with Your Phone Junjie Fang et.al. 2603.05504 null
2026-03-05 Neural Wavefunction Calculations of μSR Spectra with Quantum Muons and Protons Jamie Carr et.al. 2603.05453 null
2026-03-05 Latent Wasserstein Adversarial Imitation Learning Siqi Yang et.al. 2603.05440 null
2026-03-05 Ensembling Language Models with Sequential Monte Carlo Robin Shing Moon Chan et.al. 2603.05432 null
2026-03-05 Optimal Decoding with the Worm Zac Tobias et.al. 2603.05428 null
2026-03-05 SpiderCat: Optimal Fault-Tolerant Cat State Preparation Andrey Boris Khesin et.al. 2603.05391 null
2026-03-05 Finite-size scaling in quasi-3D stick percolation Ryan K. Daniels et.al. 2603.05374 null
2026-03-05 Detection of C3 in Titan with VLT-ESPRESSO Rafael Rianço-Silva et.al. 2603.05365 null
2026-03-05 A Comprehensive Approach to Directly Addressing Estimation Delays in Stochastic Guidance Liraz Mudrik et.al. 2603.05363 null
2026-03-05 DiSCTT: Consensus-Guided Self-Curriculum for Efficient Test-Time Adaptation in Reasoning Mohammad Mahdi Moradi et.al. 2603.05357 null
2026-03-05 Evaluation of Feynman integrals via numerical integration of differential equations Pau Petit Rosàs et.al. 2603.05336 null
2026-03-05 Latent Policy Steering through One-Step Flow Policies Hokyun Im et.al. 2603.05296 null
2026-03-05 Knowledge Divergence and the Value of Debate for Scalable Oversight Robin Young et.al. 2603.05293 null
2026-03-05 Whispering to a Blackbox: Bootstrapping Frozen OCR with Visual Prompts Samandar Samandarov et.al. 2603.05276 null
2026-03-05 SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning Zhu Li et.al. 2603.05275 null
2026-03-05 Monitoring Covariance in Multichannel Profiles via Functional Graphical Models Christian Capezza et.al. 2603.05274 null
2026-03-05 Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum Shan Ning et.al. 2603.05256 null
2026-03-05 Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards Linghan Fang et.al. 2603.05231 null
2026-03-05 KARL: Knowledge Agents via Reinforcement Learning Jonathan D. Chang et.al. 2603.05218 null
2026-03-05 LBM: Hierarchical Large Auto-Bidding Model via Reasoning and Acting Yewen Li et.al. 2603.05134 null
2026-03-04 Dynamic properties in a collisional model for confined granular fluids. A review Ricardo Brito et.al. 2603.04388 null
2026-03-04 TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoning Maximilian von Klinski et.al. 2603.04380 null
2026-03-04 Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks Haoyu Liu et.al. 2603.04364 null
2026-03-04 A Constrained RL Approach for Cost-Efficient Delivery of Latency-Sensitive Applications Ozan Aygün et.al. 2603.04353 null
2026-03-04 Tendon Force Modeling for Sim2Real Transfer of Reinforcement Learning Policies for Tendon-Driven Robots Valentin Yuryev et.al. 2603.04351 null
2026-03-04 What Does Flow Matching Bring To TD Learning? Bhavya Agrawalla et.al. 2603.04333 null
2026-03-04 IPD: Boosting Sequential Policy with Imaginary Planning Distillation in Offline Reinforcement Learning Yihao Qin et.al. 2603.04289 null
2026-03-04 SSR: A Generic Framework for Text-Aided Map Compression for Localization Mohammad Omama et.al. 2603.04272 null
2026-03-04 Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory Zhenting Wang et.al. 2603.04257 null
2026-03-04 Emergent dimensional reduction in a distorted kagome magnet $\mathrm{YCa_3(CrO)_3(BO_3)_4}$ driven by exchange hierarchy Umashankar Jena et.al. 2603.04253 null
2026-03-04 OptiQKD: A Machine Learning-Optimized Framework for Real-Time Parameter Tuning in Quantum Key Distribution Noureldin Mohamed et.al. 2603.04192 null
2026-03-04 Learning Hip Exoskeleton Control Policy via Predictive Neuromusculoskeletal Simulation Ilseung Park et.al. 2603.04166 null
2026-03-04 Data-Aware Random Feature Kernel for Transformers Amirhossein Farzam et.al. 2603.04127 null
2026-03-04 BeamPERL: Parameter-Efficient RL with Verifiable Rewards Specializes Compact LLMs for Structured Beam Mechanics Reasoning Tarjei Paule Hage et.al. 2603.04124 null
2026-03-04 Real Eyes Realize Faster: Gaze Stability and Pupil Novelty for Efficient Egocentric Learning Ajan Subramanian et.al. 2603.04098 null
2026-03-04 Swimming Under Constraints: A Safe Reinforcement Learning Framework for Quadrupedal Bio-Inspired Propulsion Xinyu Cui et.al. 2603.04073 null
2026-03-04 SaFeR: Safety-Critical Scenario Generation for Autonomous Driving Test via Feasibility-Constrained Token Resampling Jinlong Cui et.al. 2603.04071 null
2026-03-04 Force-Aware Residual DAgger via Trajectory Editing for Precision Insertion with Impedance Control Yiou Huang et.al. 2603.04038 null
2026-03-04 Self-adapting Robotic Agents through Online Continual Reinforcement Learning with World Model Feedback Fabian Domberg et.al. 2603.04029 null
2026-03-04 Rethinking the Efficiency and Effectiveness of Reinforcement Learning for Radiology Report Generation Zilin Lu et.al. 2603.04022 null
2026-03-03 How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference Toru Lin et.al. 2603.03280 null
2026-03-03 ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation Xialin He et.al. 2603.03279 null
2026-03-03 Valet: A Standardized Testbed of Traditional Imperfect-Information Card Games Mark Goadrich et.al. 2603.03252 null
2026-03-03 The Extended Real Line with Reentry: A Compact Quotient Space Separating US from KC Damian Rafael Lattenero et.al. 2603.03228 null
2026-03-03 Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use Aradhye Agarwal et.al. 2603.03205 null
2026-03-03 Specificity-aware reinforcement learning for fine-grained open-world classification Samuele Angheben et.al. 2603.03197 null
2026-03-03 A Covering Framework for Offline POMDPs Learning using Belief Space Metric Youheng Zhu et.al. 2603.03191 null
2026-03-03 Heavy-quark box-loop corrections to $q\bar q \to Zγ$ at two loops in QCD Dario Kermanschah et.al. 2603.03169 null
2026-03-03 Stable solutions to reaction-diffusion elliptic problems Xavier Cabre et.al. 2603.03161 null
2026-03-03 Geometry-Guided Reinforcement Learning for Multi-view Consistent 3D Scene Editing Jiyuan Wang et.al. 2603.03143 null
2026-03-03 RL-Based Coverage Path Planning for Deformable Objects on 3D Surfaces Yuhang Zhang et.al. 2603.03137 null
2026-03-03 Deep Q-Learning-Based Gain Scheduling for Nonlinear Quadcopter Dynamics Hossein Rastgoftar et.al. 2603.03127 null
2026-03-03 Proactive Guiding Strategy for Item-side Fairness in Interactive Recommendation Chongjun Xia et.al. 2603.03094 null
2026-03-03 Safe and Robust Domains of Attraction for Discrete-Time Systems: A Set-Based Characterization and Certifiable Neural Network Estimation Mohamed Serry et.al. 2603.03082 null
2026-03-03 RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy Optimization Siwei Zhang et.al. 2603.03078 null
2026-03-03 TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning Christian Greisinger et.al. 2603.03072 null
2026-03-03 Reinforcement Learning with Symbolic Reward Machines Thomas Krug et.al. 2603.03068 null
2026-03-03 CMoE: Contrastive Mixture of Experts for Motion Control and Terrain Adaptation of Humanoid Robots Shihao Ma et.al. 2603.03067 null
2026-03-03 PrivMedChat: End-to-End Differentially Private RLHF for Medical Dialogue Systems Sudip Bhujel et.al. 2603.03054 null
2026-03-03 QFlowNet: Fast, Diverse, and Efficient Unitary Synthesis with Generative Flow Networks Inhoe Koo et.al. 2603.03045 null
2026-03-02 Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Training Valentin Lacombe et.al. 2603.02208 null
2026-03-02 Tool Verification for Test-Time Reinforcement Learning Ruotong Liao et.al. 2603.02203 null
2026-03-02 Kinetic energy fluctuations and specific heat in generalized ensembles Sergio Davis et.al. 2603.02168 null
2026-03-02 Near-Optimal Regret for KL-Regularized Multi-Armed Bandits Kaixuan Ji et.al. 2603.02155 null
2026-03-02 Boltzmann-based Exploration for Robust Decentralized Multi-Agent Planning Nhat Nguyen et.al. 2603.02154 null
2026-03-02 LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards Guanzheng Chen et.al. 2603.02146 null
2026-03-02 Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation Han Xue et.al. 2603.02139 null
2026-03-02 Pencil Puzzle Bench: A Benchmark for Multi-Step Verifiable Reasoning Justin Waugh et.al. 2603.02119 null
2026-03-02 Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons Anthony Liang et.al. 2603.02115 null
2026-03-02 ACDC: Adaptive Curriculum Planning with Dynamic Contrastive Control for Goal-Conditioned Reinforcement Learning in Robotic Manipulation Xuerui Wang et.al. 2603.02104 null
2026-03-02 Learning from Synthetic Data Improves Multi-hop Reasoning Anmol Kabra et.al. 2603.02091 null
2026-03-02 Reinforcement Learning-Based Filters for Convection-Dominated Flows: Reference-Free and Reference-Guided Training Anna Ivagnes et.al. 2603.02086 null
2026-03-02 $π$ -StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs Siting Wang et.al. 2603.02083 null
2026-03-02 Accelerating PDE Surrogates via RL-Guided Mesh Optimization Yang Meng et.al. 2603.02066 null
2026-03-02 Expanding LLM Agent Boundaries with Strategy-Guided Exploration Andrew Szot et.al. 2603.02045 null
2026-03-02 Temporal Representations for Exploration: Learning Complex Exploratory Behavior without Extrinsic Rewards Faisal Mohamed et.al. 2603.02008 null
2026-03-02 Process Over Outcome: Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection Yuchen Zhang et.al. 2603.01993 null
2026-03-02 CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production Yixin Nie et.al. 2603.01973 null
2026-03-02 CoVe: Training Interactive Tool-Use Agents via Constraint-Guided Verification Jinpeng Chen et.al. 2603.01940 null
2026-03-02 LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving Yuechen Luo et.al. 2603.01928 null
2026-02-27 DARE-bench: Evaluating Modeling and Instruction Fidelity of LLMs in Data Science Fan Shu et.al. 2602.24288 null
2026-02-27 CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation Weinan Dai et.al. 2602.24286 null
2026-02-27 Geometric Resilience of Quantum LiDAR in Turbulent Media: A Wasserstein Distance Approach Arnaud Coatanhay et.al. 2602.24280 null
2026-02-27 Curriculum-Based Soft Actor-Critic for Multi-Section R2R Tension Control Shihao Li et.al. 2602.24259 null
2026-02-27 On Hamiltonian Monte Carlo for Gaussian Random Variables with Random Hamiltonians Yingdong Lu et.al. 2602.24256 null
2026-02-27 SafeGen-LLM: Enhancing Safety Generalization in Task Planning for Robotic Systems Jialiang Fan et.al. 2602.24235 null
2026-02-27 Enhancing Spatial Understanding in Image Generation via Reward Modeling Zhenyu Tang et.al. 2602.24233 null
2026-02-27 Empirical Challenges with Peers-of-Peers Instruments in the Linear-In-Means Model Nathan Canen et.al. 2602.24215 null
2026-02-27 Vacancy-induced local moments in quantum paramagnetic phases: An SU( $N$ ) designer Hamiltonian study Md Zahid Ansari et.al. 2602.24203 null
2026-02-27 Multi-Objective Reinforcement Learning for Large-Scale Tote Allocation in Human-Robot Collaborative Fulfillment Centers Sikata Sengupta et.al. 2602.24182 null
2026-02-27 Learning Flexible Job Shop Scheduling under Limited Buffers and Material Kitting Constraints Shishun Zhang et.al. 2602.24180 null
2026-02-27 Theoretical Studies of alpha Clustering in Nuclei and Beyond Takaharu Otsuka et.al. 2602.24175 null
2026-02-27 Advanced Scheduling Strategies for Distributed Quantum Computing Jobs Gongyu Ni et.al. 2602.24152 null
2026-02-27 Planning from Observation and Interaction Tyler Han et.al. 2602.24121 null
2026-02-27 Microwave response of fractional quantum Hall droplets with quasiparticle tunneling Fumihiro Murabayashi et.al. 2602.24116 null
2026-02-27 Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidance Yanwei Ren et.al. 2602.24110 null
2026-02-27 Bi-level RL-Heuristic Optimization for Real-world Winter Road Maintenance Yue Xie et.al. 2602.24097 null
2026-02-27 An $ε$ -Optimal Sequential Approach for Solving zs-POSGs Jilles S. Dibangoye et.al. 2602.24092 null
2026-02-27 Preference Packing: Efficient Preference Optimization for Large Language Models Jaekyung Cho et.al. 2602.24082 null
2026-02-27 Adaptive Correlation-Weighted Intrinsic Rewards for Reinforcement Learning Viet Bac Nguyen et.al. 2602.24081 null
2026-02-26 MediX-R1: Open Ended Medical Reinforcement Learning Sahal Shaji Mullappilly et.al. 2602.23363 null
2026-02-26 Generalized Rapid Action Value Estimation in Memory-Constrained Environments Aloïs Rautureau et.al. 2602.23318 null
2026-02-26 Impacts of Aggregation on Model Diversity and Consumer Utility Kate Donahue et.al. 2602.23293 null
2026-02-26 Simple Models, Real Swimming: Digital Twins for Tendon-Driven Underwater Robots Mike Y. Michelis et.al. 2602.23283 null
2026-02-26 Physics Informed Viscous Value Representations Hrishikesh Viswanath et.al. 2602.23280 null
2026-02-26 Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving Jiangxin Sun et.al. 2602.23259 null
2026-02-26 SPARR: Simulation-based Policies with Asymmetric Real-world Residuals for Assembly Yijie Guo et.al. 2602.23253 null
2026-02-26 A Model-Free Universal AI Yegon Kim et.al. 2602.23242 null
2026-02-26 Agency and Architectural Limits: Why Optimization-Based Systems Cannot Be Norm-Responsive Radha Sarma et.al. 2602.23239 null
2026-02-26 Symmetric Mass Generation via Multicriticality in a 3D Lattice Gross-Neveu Model Sandip Maiti et.al. 2602.23158 null
2026-02-26 Towards Intelligible Human-Robot Interaction: An Active Inference Approach to Occluded Pedestrian Scenarios Kai Chen et.al. 2602.23109 null
2026-02-26 Accelerated Online Risk-Averse Policy Evaluation in POMDPs with Theoretical Guarantees and Novel CVaR Bounds Yaacov Pariente et.al. 2602.23073 null
2026-02-26 GeoWorld: Geometric World Models Zeyu Zhang et.al. 2602.23058 null
2026-02-26 Learning-based Multi-agent Race Strategies in Formula 1 Giona Fieni et.al. 2602.23056 null
2026-02-26 Tracking the Lithiation State of Li $_x$ Si from Machine-Learned XPS Binding Energies Michael Alejandro Hernandez Bertran et.al. 2602.23028 null
2026-02-26 Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization Zeyuan Liu et.al. 2602.23008 null
2026-02-26 Regular Fourier Features for Nonstationary Gaussian Processes Arsalan Jawaid et.al. 2602.23006 null
2026-02-26 A Perspective on Open Challenges in Deformable Object Manipulation Ryan Paul McKennaa et.al. 2602.22998 null
2026-02-26 Influence of Hydrogen on Dislocation Relaxation in BCC Iron: Atomistic Mechanisms and Implications Sanjay Manda et.al. 2602.22995 null
2026-02-26 Energy Deposition by Galactic Cosmic Rays and Implications for Ozone Chemistry Luiz Augusto Stuani Pereia et.al. 2602.22987 null
2026-02-25 Improving Parametric Knowledge Access in Reasoning Language Models Melody Ma et.al. 2602.22193 null
2026-02-25 Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual Yining Li et.al. 2602.22146 null
2026-02-25 Tempered Christoffel-Weighted Polynomial Chaos Expansion for Resilience-Oriented Uncertainty Quantification Mahsa Ebadat-Parast et.al. 2602.22133 null
2026-02-25 SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents Patrick Tser Jern Kon et.al. 2602.22124 null
2026-02-25 System Design of the Ultra Mobility Vehicle: A Driving, Balancing, and Jumping Bicycle Robot Benjamin Bokser et.al. 2602.22118 null
2026-02-25 Stochastic Optimal Control with Side Information and Bayesian Learning Johannes Milz et.al. 2602.22047 null
2026-02-25 IGR J12580+0134: A Candidate for Repeating Partial Tidal Disruption Events Supported by Multi-Wavelength Observations Po Ma et.al. 2602.22040 null
2026-02-25 Function-Space Empirical Bayes Regularisation with Student’s t Priors Pengcheng Hao et.al. 2602.22015 null
2026-02-25 Intrusive and Non-Intrusive Model Order Reduction for Airborne Contaminant Transport: Comparative Analysis and Uncertainty Quantification Lisa Kühn et.al. 2602.21996 null
2026-02-25 PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement Learning Zekai Lin et.al. 2602.21992 null
2026-02-25 RADAR: Reasoning as Discrimination with Aligned Representations for LLM-based Knowledge Graph Reasoning Bo Xue et.al. 2602.21951 null
2026-02-25 Bayesian Generative Adversarial Networks via Gaussian Approximation for Tabular Data Synthesis Bahrul Ilmi Nasution et.al. 2602.21948 null
2026-02-25 ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selection Changjiang Gao et.al. 2602.21887 null
2026-02-25 Distill and Align Decomposition for Enhanced Claim Verification Jabez Magomere et.al. 2602.21857 null
2026-02-25 LightSim: A Lightweight Cell Transmission Model Simulator for Traffic Signal Control Research Haoran Su et.al. 2602.21852 null
2026-02-25 Self-Curriculum Model-based Reinforcement Learning for Shape Control of Deformable Linear Objects Zhaowei Liang et.al. 2602.21816 null
2026-02-25 DexRepNet++: Learning Dexterous Robotic Manipulation with Geometric and Spatial Hand-Object Representations Qingtao Liu et.al. 2602.21811 null
2026-02-25 RAMSeS: Robust and Adaptive Model Selection for Time-Series Anomaly Detection Algorithms Mohamed Abdelmaksoud et.al. 2602.21766 null
2026-02-25 Generalisation of RLHF under Reward Shift and Clipped KL Regularisation Kenton Tang et.al. 2602.21765 null
2026-02-25 Pilot-Free Optimal Control over Wireless Networks: A Control-Aided Channel Prediction Approach Minjie Tang et.al. 2602.21752 null
2026-02-24 Minimal loop currents in doped Mott insulators Can Cui et.al. 2602.21206 null
2026-02-24 Squint: Fast Visual Reinforcement Learning for Sim-to-Real Robotics Abdulaziz Almuzairee et.al. 2602.21203 null
2026-02-24 Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training Anas Barakat et.al. 2602.21189 null
2026-02-24 SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards Dengjia Zhang et.al. 2602.21158 null
2026-02-24 Equivariant Floer cohomology for contactomorphisms of quotient spaces Dylan Cant et.al. 2602.21152 null
2026-02-24 Cooperative-Competitive Team Play of Real-World Craft Robots Rui Zhao et.al. 2602.21119 null
2026-02-24 Density Functional Theory Predictions of Derivative Thermodynamic Properties of a Confined Fluid Gennady Y. Gor et.al. 2602.21111 null
2026-02-24 Singular Arrange and Traverse Algorithm for Computing Reeb Spaces of Bivariate PL Maps Petar Hristov et.al. 2602.21087 null
2026-02-24 Localized Dynamics-Aware Domain Adaption for Off-Dynamics Offline Reinforcement Learning Zhangjie Xia et.al. 2602.21072 null
2026-02-24 PIME: Prototype-based Interpretable MCTS-Enhanced Brain Network Analysis for Disorder Diagnosis Kunyu Zhang et.al. 2602.21046 null
2026-02-24 Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning Antoine Bergerault et.al. 2602.21020 null
2026-02-24 Cell-Free Massive MIMO-Assisted SWIPT Using Stacked Intelligent Metasurfaces Thien Duc Hua et.al. 2602.20983 null
2026-02-24 The Art of Efficient Reasoning: Data, Reward, and Optimization Taiqiang Wu et.al. 2602.20945 null
2026-02-24 Empathy Modeling in Active Inference Agents for Perspective-Taking and Alignment Albarracin Mahault et.al. 2602.20936 null
2026-02-24 Task-oriented grasping for dexterous robots using postural synergies and reinforcement learning Dimitrios Dimou et.al. 2602.20915 null
2026-02-24 LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Jihao Qiu et.al. 2602.20913 null
2026-02-24 Regret-Guided Search Control for Efficient Learning in AlphaZero Yun-Jui Tsai et.al. 2602.20809 null
2026-02-24 Probing Dec-POMDP Reasoning in Cooperative MARL Kale-ab Tessera et.al. 2602.20804 null
2026-02-24 Overton Pluralistic Reinforcement Learning for Large Language Models Yu Fu et.al. 2602.20759 null
2026-02-24 Deep unfolding of MCMC kernels: scalable, modular & explainable GANs for high-dimensional posterior sampling Jonathan Spence et.al. 2602.20758 null
2026-02-23 Recurrent Structural Policy Gradient for Partially Observable Mean Field Games Clarisse Wibault et.al. 2602.20141 null
2026-02-23 PackFlow: Generative Molecular Crystal Structure Prediction via Reinforcement Learning Alignment Akshay Subramanian et.al. 2602.20140 null
2026-02-23 Development of a Cherenkov-Based Time-of-Flight Detector Using Silicon Photomultipliers Liliana Congedo et.al. 2602.20139 null
2026-02-23 LAD: Learning Advantage Distribution for Reasoning Wendi Li et.al. 2602.20132 null
2026-02-23 Enormous Fluid Antenna Systems (E-FAS)–Part II: Channel Estimation Farshad Rostami Ghadi et.al. 2602.20127 null
2026-02-23 ReSyn: Autonomously Scaling Synthetic Environments for Reasoning Models Andre He et.al. 2602.20117 null
2026-02-23 Energy gap of quantum spin glasses: a projection quantum Monte Carlo study L. Brodoloni et.al. 2602.20108 null
2026-02-23 Adaptive Underwater Acoustic Communications with Limited Feedback: An AoI-Aware Hierarchical Bandit Approach Fabio Busacca et.al. 2602.20105 null
2026-02-23 Descent-Guided Policy Gradient for Scalable Cooperative Multi-Agent Learning Shan Yang et.al. 2602.20078 null
2026-02-23 Cosmic strings and domain walls: the impact of CMB $B$ -mode data Luca Caloni et.al. 2602.20050 null
2026-02-23 noDice: Inference for Discrete Probabilistic Programs with Nondeterminism and Conditioning Tobias Gürtler et.al. 2602.20049 null
2026-02-23 A Secure and Private Distributed Bayesian Federated Learning Design Nuocheng Yang et.al. 2602.20003 null
2026-02-23 The interplay of cation/anion and monovalent/divalent selectivity in negatively charged nanopores: local charge inversion and anion leakage Eszter Lakics et.al. 2602.19992 null
2026-02-23 RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection Tianyu Wang et.al. 2602.19974 null
2026-02-23 Sparse Masked Attention Policies for Reliable Generalization Caroline Horsch et.al. 2602.19956 null
2026-02-23 Beyond Mimicry: Toward Lifelong Adaptability in Imitation Learning Nathan Gavenski et.al. 2602.19930 null
2026-02-23 Investigating polarization signatures from GRB models Pieter vd Merwe et.al. 2602.19921 null
2026-02-23 Janus-Q: End-to-End Event-Driven Trading via Hierarchical-Gated Reward Modeling Xiang Li et.al. 2602.19919 null
2026-02-23 Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning Thanh Nguyen et.al. 2602.19917 null
2026-02-23 DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning Zhongwei Wan et.al. 2602.19895 null
2026-02-20 Convex Block-Cholesky Approach to Risk-Constrained Low-thrust Trajectory Design under Operational Uncertainty Kenshiro Oguri et.al. 2602.18416 null
2026-02-20 Learning to Tune Pure Pursuit in Autonomous Racing: Joint Lookahead and Steering-Gain Control with PPO Mohamed Elgouhary et.al. 2602.18386 null
2026-02-20 On the simulated kinematic distributions of semileptonic $B$ decays Florian Herren et.al. 2602.18378 null
2026-02-20 Phase diagram of a lattice fermion model with symmetric mass generation Sandip Maiti et.al. 2602.18360 null
2026-02-20 Learning Smooth Time-Varying Linear Policies with an Action Jacobian Penalty Zhaoming Xie et.al. 2602.18312 null
2026-02-20 Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies Zhuoran Li et.al. 2602.18291 null
2026-02-20 PRISM: Parallel Reward Integration with Symmetry for MORL Finn van der Knaap et.al. 2602.18277 null
2026-02-20 CMB anisotropies from cosmic (super)strings in light of ACT DR6 Juhan Raidal et.al. 2602.18272 null
2026-02-20 A Probabilistic Framework for LLM-Based Model Discovery Stefan Wahl et.al. 2602.18266 null
2026-02-20 BLM-Guard: Explainable Multimodal Ad Moderation with Chain-of-Thought and Policy-Aligned Rewards Yiran Yang et.al. 2602.18193 null
2026-02-20 Kolmogorov-Type Maximal Inequalities for Independent and Dependent Negative Binomial Random Variables: Sharp Bounds, Sub-Exponential Refinements, and Applications to Overdispersed Count Data Aristides V. Doumas et.al. 2602.18184 null
2026-02-20 Unifying Formal Explanations: A Complexity-Theoretic Perspective Shahaf Bassan et.al. 2602.18160 null
2026-02-20 Time consistent portfolio strategies for a general utility function Oumar Mbodji et.al. 2602.18157 null
2026-02-20 Inclusive Ranking of Indian States via Bayesian Bradley-Terry Model Arshi Rizvi et.al. 2602.18150 null
2026-02-20 Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning Yongjae Shin et.al. 2602.18117 null
2026-02-20 TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs Rong Fu et.al. 2602.18109 null
2026-02-20 Interacting safely with cyclists using Hamilton-Jacobi reachability and reinforcement learning Aarati Andrea Noronha et.al. 2602.18097 null
2026-02-20 Entropy-regularized penalization schemes for American options and reflected BSDEs with singular generators Daniel Chee et.al. 2602.18078 null
2026-02-20 Hidden-charm (uds\,c\bar c) pentaquarks as flavor eigenstates in a constituent quark model M. C. Gordillo et.al. 2602.18075 null
2026-02-20 EgoPush: Learning End-to-End Egocentric Multi-Object Rearrangement for Mobile Robots Boyuan An et.al. 2602.18071 null
2026-02-19 MARS: Margin-Aware Reward-Modeling with Self-Refinement Payel Bhattacharjee et.al. 2602.17658 null
2026-02-19 SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer Nathan S. de Lara et.al. 2602.17632 null
2026-02-19 Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMs Luke Huang et.al. 2602.17616 null
2026-02-19 Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts for Geospatial Discovery Jowaria Khan et.al. 2602.17605 null
2026-02-19 Asymptotically Optimal Sequential Testing with Markovian Data Alhad Sethi et.al. 2602.17587 null
2026-02-19 Momentum Measurement of Charged Particles in FASER’s Emulsion Detector at the LHC FASER Collaboration et.al. 2602.17575 null
2026-02-19 Hybrid Monte Carlo for Fractional Quantum Hall States Ting-Tung Wang et.al. 2602.17564 null
2026-02-19 RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward Qiucheng Wu et.al. 2602.17558 null
2026-02-19 MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning Xiaoliang Fu et.al. 2602.17550 null
2026-02-19 IRIS: Learning-Driven Task-Specific Cinema Robot Arm for Visuomotor Motion Control Qilong Cheng et.al. 2602.17537 null
2026-02-19 An extension to reversible jump Markov chain Monte Carlo for change point problems with heterogeneous temporal dynamics Emily Gribbin et.al. 2602.17503 null
2026-02-19 Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Models Wen-Tse Chen et.al. 2602.17497 null
2026-02-19 Modeling of Relativistic Plasmas with a Conservative Discontinuous Galerkin Method James Juno et.al. 2602.17487 null
2026-02-19 Hartree shift and pairing gap in ultracold Fermi gases in the framework of low-momentum interactions Michael Urban et.al. 2602.17420 null
2026-02-19 Stackelberg Dynamic Location Planning under Cumulative Demand Warley Almeida Silva et.al. 2602.17392 null
2026-02-19 MDP Planning as Policy Inference David Tolpin et.al. 2602.17375 null
2026-02-19 Computer-Using World Model Yiming Guan et.al. 2602.17365 null
2026-02-19 g4chargeit: Geant4-based kinetic Monte Carlo simulations of charging in dielectric materials Kush P. Gandhi et.al. 2602.17332 null
2026-02-19 LexiSafe: Offline Safe Reinforcement Learning with Lexicographic Safety-Reward Hierarchy Hsin-Jung Yang et.al. 2602.17312 null
2026-02-19 Lepton energy scale and resolution corrections based on the minimization of an analytical likelihood: IJazZ2.0 F. Couderc et.al. 2602.17300 null
2026-02-18 Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation Runpei Dong et.al. 2602.16705 null
2026-02-18 Reinforced Fast Weights with Next-Sequence Prediction Hee Seung Hwang et.al. 2602.16704 null
2026-02-18 Ab Initio Auxiliary-Field Quantum Monte Carlo in the Thermodynamic Limit Jinghong Zhang et.al. 2602.16679 null
2026-02-18 Learning to unfold cloth: Scaling up world models to deformable object manipulation Jack Rome et.al. 2602.16675 null
2026-02-18 Optimizing p-spin models through hypergraph neural networks and deep reinforcement learning Li Zeng et.al. 2602.16665 null
2026-02-18 Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes Ethan Blaser et.al. 2602.16629 null
2026-02-18 A Rough Functional Breuer-Major Theorem Henri Elad Altman et.al. 2602.16615 null
2026-02-18 Optimal Placement and Sizing of PV-Based DG Units in a Distribution Network Considering Loading Capacity Abhinav Sharma et.al. 2602.16565 null
2026-02-18 A Scalable Approach to Solving Simulation-Based Network Security Games Michael Lanier et.al. 2602.16564 null
2026-02-18 Learning Distributed Equilibria in Linear-Quadratic Stochastic Differential Games: An $α$ -Potential Approach Philipp Plank et.al. 2602.16555 null
2026-02-18 RIDER: 3D RNA Inverse Design with Reinforcement Learning-Guided Diffusion Tianmeng Hu et.al. 2602.16548 null
2026-02-18 Vulnerability Analysis of Safe Reinforcement Learning via Inverse Constrained Reinforcement Learning Jialiang Fan et.al. 2602.16543 null
2026-02-18 Capacity-constrained demand response in smart grids using deep reinforcement learning Shafagh Abband Pashaki et.al. 2602.16525 null
2026-02-18 Reinforcement Learning for Parameterized Quantum State Preparation: A Comparative Study Gerhard Stenzel et.al. 2602.16523 null
2026-02-18 VIGOR: Visual Goal-In-Context Inference for Unified Humanoid Fall Safety Osher Azulay et.al. 2602.16511 null
2026-02-18 Quantum-classical correspondence for spins at finite temperatures with application to Monte Carlo simulations A. El Mendili et.al. 2602.16501 null
2026-02-18 Certifying Hamilton-Jacobi Reachability Learned via Reinforcement Learning Prashant Solanki et.al. 2602.16475 null
2026-02-18 Monte Carlo study of the classical antiferromagnetic $J_1$-$J_2$-$J_3$ Heisenberg model on a simple cubic lattice A. N. Ignatenko et.al. 2602.16471 null
2026-02-18 Causally-Guided Automated Feature Engineering with Multi-Agent Reinforcement Learning Arun Vignesh Malarkkan et.al. 2602.16435 null
2026-02-18 Computation of thermal conductivity based on Path Integral Monte Carlo methods Vladislav Efremkin et.al. 2602.16405 null
2026-02-17 Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching Zhen Wu et.al. 2602.15827 null
2026-02-17 Solving Parameter-Robust Avoid Problems with Unknown Feasibility using Reinforcement Learning Oswin So et.al. 2602.15817 null
2026-02-17 GLM-5: from Vibe Coding to Agentic Engineering GLM-5 Team et.al. 2602.15763 null
2026-02-17 MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction Qiang Zhang et.al. 2602.15733 null
2026-02-17 Recursive Concept Evolution for Compositional Reasoning in Large Language Models Sarim Chaudhry et.al. 2602.15725 null
2026-02-17 Learning to Retrieve Navigable Candidates for Efficient Vision-and-Language Navigation Shutian Gu et.al. 2602.15724 null
2026-02-17 Pricing Discrete and Nonlinear Markets With Semidefinite Relaxations Cheng Guo et.al. 2602.15722 null
2026-02-17 Bayesian parameter study of the Seyfert-starburst composite galaxies NGC 1068 and NGC 7469 Björn Eichmann et.al. 2602.15644 null
2026-02-17 Reinforcement Learning in Real Option Models Jodi Dianetti et.al. 2602.15643 null
2026-02-17 Latency-aware Human-in-the-Loop Reinforcement Learning for Semantic Communications Peizheng Li et.al. 2602.15640 null
2026-02-17 STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens Shiqi Liu et.al. 2602.15620 null
2026-02-17 Physics-Informed Anomaly Detection of Terrain Material Change in Radar Imagery Abdel Hakiem Mohamed Abbas Mohamed Ahmed et.al. 2602.15618 null
2026-02-17 Stability of Bose-Fermi mixtures in two dimensions: a lowest-order constrained variational approach Pietro Cordioli et.al. 2602.15598 null
2026-02-17 Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQL Yihan Wang et.al. 2602.15564 null
2026-02-17 Some phenomenological aspects of a quantum-corrected Reissner-Nordström black hole: quasi-periodic oscillations, scalar perturbations and thermal fluctuations Faizuddin Ahmed et.al. 2602.15551 null
2026-02-17 Efficient Knowledge Transfer for Jump-Starting Control Policy Learning of Multirotors through Physics-Aware Neural Architectures Welf Rehberg et.al. 2602.15533 null
2026-02-17 Tight Communication Bounds for Distributed Algorithms in the Quantum Routing Model Fabien Dufoulon et.al. 2602.15529 null
2026-02-17 The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes Mohammad Taufeeque et.al. 2602.15515 null
2026-02-17 The First Instrumentally Documented Fall of an Iron Meteorite: atmospheric trajectory and ground impact Jarmo Moilanen et.al. 2602.15440 null
2026-02-17 Fairness over Equality: Correcting Social Incentives in Asymmetric Sequential Social Dilemmas Alper Demir et.al. 2602.15407 null
2026-02-16 Modeling isolated magnetar spin-down evolution and implications for long-period radio transients Jon Kwong et.al. 2602.15024 null
2026-02-16 Cold-Start Personalization via Training-Free Priors from Structured World Models Avinandan Bose et.al. 2602.15012 null
2026-02-16 BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames Max Sobol Mark et.al. 2602.15010 null
2026-02-16 Learning User Interests via Reasoning and Distillation for Cross-Domain News Recommendation Mengdan Zhu et.al. 2602.15005 null
2026-02-16 Activation-Space Uncertainty Quantification for Pretrained Networks Richard Bergna et.al. 2602.14934 null
2026-02-16 MAC-AMP: A Closed-Loop Multi-Agent Collaboration System for Multi-Objective Antimicrobial Peptide Design Gen Zhou et.al. 2602.14926 null
2026-02-16 Auxiliary field quantum Monte Carlo at the basis set limit: application to lattice constants Moritz Humer et.al. 2602.14923 null
2026-02-16 BFS-PO: Best-First Search for Large Reasoning Models Fiorenzo Parascandolo et.al. 2602.14917 null
2026-02-16 On the Learning Dynamics of RLVR at the Edge of Competence Yu Huang et.al. 2602.14872 null
2026-02-16 Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning Ilia Mahrooghi et.al. 2602.14868 null
2026-02-16 Gauge-independent gravitational waves from a minimal dark $U(1)$ sector with viable dark matter candidates Wan-Zhe Feng et.al. 2602.14866 null
2026-02-16 Bias analysis of a linear order-statistic inequality index estimator: Unbiasedness under gamma populations Roberto Vila et.al. 2602.14861 null
2026-02-16 Lower Estimates for $L_1$ -Distortion of Transportation Cost Spaces Chris Gartland et.al. 2602.14852 null
2026-02-16 Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment Elias Malomgré et.al. 2602.14844 null
2026-02-16 Constructions of linear codes from vectorial plateaued functions and their subfield codes with applications to quantum CSS codes Virginio Fratianni et.al. 2602.14832 null
2026-02-16 The empirical distribution of sequential LS factors in Multi-level Dynamic Factor Models Gian Pietro Bellocca et.al. 2602.14813 null
2026-02-16 ManeuverNet: A Soft Actor-Critic Framework for Precise Maneuvering of Double-Ackermann-Steering Robots with Optimized Reward Functions Kohio Deflesselle et.al. 2602.14726 null
2026-02-16 Evolutionary System Prompt Learning can Facilitate Reinforcement Learning for LLMs Lunjun Zhang et.al. 2602.14697 null
2026-02-16 Weak Poincaré inequalities for Deterministic-scan Metropolis-within-Gibbs samplers Mengxi Gao et.al. 2602.14692 null
2026-02-16 GREAT-EER: Graph Edge Attention Network for Emergency Evacuation Responses Attila Lischka et.al. 2602.14676 null
2026-02-13 $\texttt{GPUmonty}$ : A GPU-accelerated relativistic Monte Carlo radiative transfer code Pedro Naethe Motta et.al. 2602.13198 null
2026-02-13 Nuclear gradients from auxiliary-field quantum Monte Carlo and their application in geometry optimization and transition state search Jo S. Kurian et.al. 2602.13187 null
2026-02-13 Operator Learning for Families of Finite-State Mean-Field Games William Hofgard et.al. 2602.13169 null
2026-02-13 A Data-Driven Algorithm for Model-Free Control Synthesis Sean Bowerfind et.al. 2602.13157 null
2026-02-13 In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach Yiran Gao et.al. 2602.13156 null
2026-02-13 Learning to Approximate Uniform Facility Location via Graph Neural Networks Chendi Qian et.al. 2602.13155 null
2026-02-13 Peaceful Anarcho-Accelerationism: Decentralized Full Automation for a Society of Universal Care Eduardo C. Garrido-Merchán et.al. 2602.13154 null
2026-02-13 Deconfinement from Thermal Tensor Networks: Universal CFT signature in (2+1)-dimensional $\mathbb{Z}_N$ lattice gauge theory Adwait Naravane et.al. 2602.13124 null
2026-02-13 Curriculum-DPO++: Direct Preference Optimization via Data and Model Curricula for Text-to-Image Generation Florinel-Alin Croitoru et.al. 2602.13055 null
2026-02-13 TCRL: Temporal-Coupled Adversarial Training for Robust Constrained Reinforcement Learning in Worst-Case Scenarios Wentao Xu et.al. 2602.13040 null
2026-02-13 Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL Yixiao Zhou et.al. 2602.13035 null
2026-02-13 How Swarms Differ: Challenges in Collective Behaviour Comparison André Fialho Jesus et.al. 2602.13016 null
2026-02-13 Neural Quantum States Based on Selected Configurations Marco Julian Solanki et.al. 2602.12993 null
2026-02-13 ProbeLLM: Automating Principled Diagnosis of LLM Failures Yue Huang et.al. 2602.12966 null
2026-02-13 TFTF: Training-Free Targeted Flow for Conditional Sampling Qianqian Qu et.al. 2602.12932 null
2026-02-13 Hierarchical Reinforcement Learning for Cooperative Air-Ground Delivery in Urban System Songxin Lei et.al. 2602.12913 null
2026-02-13 A unified testing approach for log-symmetry using Fourier methods Ganesh Vishnu Avhad et.al. 2602.12900 null
2026-02-13 BSN-VI: Multiband Light Curve Modeling of Four W UMa-Type Contact Binaries I. Revisiting Energy Transfer Mechanisms and Luminosity Behavior Elham Sarvari et.al. 2602.12864 null
2026-02-13 The Refractive Index of Gallium Antimonide Ulrich Galander et.al. 2602.12862 null
2026-02-13 DPUConfig: Optimizing ML Inference in FPGAs Using Reinforcement Learning Alexandros Patras et.al. 2602.12847 null
2026-02-12 CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use Zhen Zhang et.al. 2602.12268 null
2026-02-12 Intrinsic-Energy Joint Embedding Predictive Architectures Induce Quasimetric Spaces Anthony Kobanda et.al. 2602.12245 null
2026-02-12 Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks Zhihong Liu et.al. 2602.12244 null
2026-02-12 Diffusion Alignment Beyond KL: Variance Minimisation as Effective Policy Optimiser Zijing Ou et.al. 2602.12229 null
2026-02-12 Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training Miaosen Zhang et.al. 2602.12222 null
2026-02-12 DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing Dianyi Wang et.al. 2602.12205 null
2026-02-12 Convex Markov Games and Beyond: New Proof of Existence, Characterization and Learning Algorithms for Nash Equilibria Anas Barakat et.al. 2602.12181 null
2026-02-12 FAIL: Flow Matching Adversarial Imitation Learning for Image Generation Yeyao Ma et.al. 2602.12155 null
2026-02-12 Seq2Seq2Seq: Lossless Data Compression via Discrete Latent Transformers and Reinforcement Learning Mahdi Khodabandeh et.al. 2602.12146 null
2026-02-12 Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation Wenkai Yang et.al. 2602.12125 null
2026-02-12 Capability-Oriented Training Induced Alignment Risk Yujun Zhou et.al. 2602.12124 null
2026-02-12 Meta-Sel: Efficient Demonstration Selection for In-Context Learning via Supervised Meta-Learning Xubin Wang et.al. 2602.12123 null
2026-02-12 P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling Pinyi Zhang et.al. 2602.12116 null
2026-02-12 Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty Zewei Yu et.al. 2602.12113 null
2026-02-12 On the Complexity of Offline Reinforcement Learning with $Q^\star$ -Approximation and Partial Coverage Haolin Liu et.al. 2602.12107 null
2026-02-12 GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning GigaBrain Team et.al. 2602.12099 null
2026-02-12 GAN-based data augmentation for rare and exotic hadron searches in Pb–Pb collisions in ALICE Anisa Khatun et.al. 2602.12088 null
2026-02-12 Geometry of Uncertainty: Learning Metric Spaces for Multimodal State Estimation in RL Alfredo Reichlin et.al. 2602.12087 null
2026-02-12 Improving HPC Code Generation Capability of LLMs via Online Reinforcement Learning with Real-Machine Benchmark Rewards Ryo Mikasa et.al. 2602.12049 null
2026-02-12 Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models Xin Xu et.al. 2602.12036 null
2026-02-11 Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling Gongye Liu et.al. 2602.11146 null
2026-02-11 APEX: Learning Adaptive High-Platform Traversal for Humanoid Robots Yikai Wang et.al. 2602.11143 null
2026-02-11 Data-Efficient Hierarchical Goal-Conditioned Reinforcement Learning via Normalizing Flows Shaswat Garg et.al. 2602.11142 null
2026-02-11 Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards Reinhard Heckel et.al. 2602.11128 null
2026-02-11 Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away Soumya Suvra Ghosal et.al. 2602.11096 null
2026-02-11 DataChef: Cooking Up Optimal Data Recipes for LLM Adaptation via Reinforcement Learning Yicheng Chen et.al. 2602.11089 null
2026-02-11 General Flexible $f$ -divergence for Challenging Offline RL Datasets with Low Stochasticity and Diverse Behavior Policies Jianxun Wang et.al. 2602.11087 null
2026-02-11 Constrained Fiducial Inference for Gaussian Models Hank Flury et.al. 2602.11080 null
2026-02-11 Interpretable Attention-Based Multi-Agent PPO for Latency Spike Resolution in 6G RAN Slicing Kavan Fatehi et.al. 2602.11076 null
2026-02-11 RISE: Self-Improving Robot Policy with Compositional World Model Jiazhi Yang et.al. 2602.11075 null
2026-02-11 Chatting with Images for Introspective Visual Thinking Junfei Wu et.al. 2602.11073 null
2026-02-11 Simultaneous Speech-to-Speech Translation Without Aligned Data Tom Labiausse et.al. 2602.11072 null
2026-02-11 Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language Models Xinyu Yuan et.al. 2602.11057 null
2026-02-11 OSIL: Learning Offline Safe Imitation Policies with Safety Inferred from Non-preferred Trajectories Returaj Burnwal et.al. 2602.11018 null
2026-02-11 Noise-balanced multilevel on-the-fly sparse grid surrogates for coupling Monte Carlo models into continuum models with application to heterogeneous catalysis Tobias Hülser et.al. 2602.11006 null
2026-02-11 Fine-Tuning GPT-5 for GPU Kernel Generation Ali Tehrani et.al. 2602.11000 null
2026-02-11 Multi-Task Reinforcement Learning of Drone Aerobatics by Exploiting Geometric Symmetries Zhanyu Guo et.al. 2602.10997 null
2026-02-11 Non-centred Bayesian inference for discrete-valued state-transition models: the Rippler algorithm James Neill et.al. 2602.10924 null
2026-02-11 Near-Constant Strong Violation and Last-Iterate Convergence for Online CMDPs via Decaying Safety Margins Qian Zuo et.al. 2602.10917 null
2026-02-11 Resource-Efficient Model-Free Reinforcement Learning for Board Games Kazuki Ota et.al. 2602.10894 null
2026-02-10 Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction Sizhe Yang et.al. 2602.10101 null
2026-02-10 Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning Zhaoyang Wang et.al. 2602.10090 null
2026-02-10 CODE-SHARP: Continuous Open-ended Discovery and Evolution of Skills as Hierarchical Reward Programs Richard Bornemann et.al. 2602.10085 null
2026-02-10 Anagent For Enhancing Scientific Table & Figure Analysis Xuehang Guo et.al. 2602.10081 null
2026-02-10 Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability Aaditya Vikram Prasad et.al. 2602.10067 null
2026-02-10 Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization Xinchen Han et.al. 2602.10048 null
2026-02-10 Optimistic World Models: Efficient Exploration in Model-Based Deep Reinforcement Learning Akshay Mete et.al. 2602.10044 null
2026-02-10 Fake-HR1: Rethinking reasoning of vision language model for synthetic image detection Changjiang Jiang et.al. 2602.10042 null
2026-02-10 Resilient Topology-Aware Coordination for Dynamic 3D UAV Networks under Node Failure Chuan-Chi Lai et.al. 2602.10029 null
2026-02-10 ADORA: Training Reasoning Models with Dynamic Advantage Estimation on Reinforcement Learning Qingnan Ren et.al. 2602.10019 null
2026-02-10 Online Selective Conformal Prediction with Asymmetric Rules: A Permutation Test Approach Mingyi Zheng et.al. 2602.10018 null
2026-02-10 A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Chenruo Liu et.al. 2602.10014 null
2026-02-10 A Collaborative Safety Shield for Safe and Efficient CAV Lane Changes in Congested On-Ramp Merging Bharathkumar Hegde et.al. 2602.10007 null
2026-02-10 Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning Shijie Zhang et.al. 2602.10006 null
2026-02-10 ESTAR: Early-Stopping Token-Aware Reasoning For Efficient Inference Junda Wang et.al. 2602.10004 null
2026-02-10 ORCHID: Fairness-Aware Orchestration in Mission-Critical Air-Ground Integrated Networks Chuan-Chi Lai et.al. 2602.09994 null
2026-02-10 SCOPE: A Training-Free Online 3D Deployment for UAV-BSs with Theoretical Analysis and Comparative Study Chuan-Chi Lai et.al. 2602.09971 null
2026-02-10 L’Hopital rules for complex-valued functions in higher dimensions Albert Chern et.al. 2602.09958 null
2026-02-10 ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning Shuaiyi Nie et.al. 2602.09953 null
2026-02-10 Geometric Analysis of Blind User Identification for Massive MIMO Networks Levi Bohnacker et.al. 2602.09910 null
2026-02-09 TwinRL-VLA: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation Qinwen Xu et.al. 2602.09023 null
2026-02-09 WorldCompass: Reinforcement Learning for Long-Horizon World Models Zehan Wang et.al. 2602.09022 null
2026-02-09 iGRPO: Self-Feedback-Driven LLM Reasoning Ali Hatamizadeh et.al. 2602.09000 null
2026-02-09 Contraction Metric Based Safe Reinforcement Learning Force Control for a Hydraulic Actuator with Real-World Training Lucca Maitan et.al. 2602.08977 null
2026-02-09 Learning to Coordinate via Quantum Entanglement in Multi-Agent Reinforcement Learning John Gardiner et.al. 2602.08965 null
2026-02-09 Two Robust Interstellar Meteor Candidates in the Post-2018 CNEOS Fireball Database Richard Cloete et.al. 2602.08956 null
2026-02-09 StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors Suraj Ranganath et.al. 2602.08934 null
2026-02-09 Efficient and Stable Reinforcement Learning for Diffusion Language Models Jiawei Liu et.al. 2602.08905 null
2026-02-09 Learning Potentials for Dynamic Matching and Application to Heart Transplantation Itai Zilberstein et.al. 2602.08878 null
2026-02-09 AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection Junru Zhang et.al. 2602.08868 null
2026-02-09 Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems Lang Feng et.al. 2602.08847 null
2026-02-09 Learning the Value Systems of Societies with Preference-based Multi-objective Reinforcement Learning Andrés Holgado-Sánchez et.al. 2602.08835 null
2026-02-09 WildReward: Learning Reward Models from In-the-Wild Human Interactions Hao Peng et.al. 2602.08829 null
2026-02-09 VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning Hao Tan et.al. 2602.08828 null
2026-02-09 Bayesian Preference Learning for Test-Time Steerable Reward Models Jiwoo Hong et.al. 2602.08819 null
2026-02-09 Finite-State Controllers for (Hidden-Model) POMDPs using Deep Reinforcement Learning David Hudák et.al. 2602.08734 null
2026-02-09 Friedkin-Johnsen Social Influence Dynamics on Networks: A Boundary-Value Formulation and Influenceability Measures Moses Boudourides et.al. 2602.08704 null
2026-02-09 SoK: The Pitfalls of Deep Reinforcement Learning for Cybersecurity Shae McFadden et.al. 2602.08690 null
2026-02-09 Learning To Sample From Diffusion Models Via Inverse Reinforcement Learning Constant Bourdrez et.al. 2602.08689 null
2026-02-09 LLaDA2.1: Speeding Up Text Diffusion via Token Editing Tiwei Bie et.al. 2602.08676 null
2026-02-08 Direct Soft-Policy Sampling via Langevin Dynamics Donghyeon Ki et.al. 2602.07873 null
2026-02-08 MARTI-MARS $^2$ : Scaling Multi-Agent Self-Search via Reinforcement Learning for Code Generation Shijie Wang et.al. 2602.07848 null
2026-02-08 System-Level Error Propagation and Tail-Risk Amplification in Reference-Based Robotic Navigation Ning Hu et.al. 2602.07846 null
2026-02-08 TodoEvolve: Learning to Architect Agent Planning Systems Jiaxi Liu et.al. 2602.07839 null
2026-02-08 RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI Hongzhi Zang et.al. 2602.07837 null
2026-02-08 rePIRL: Learn PRM with Inverse RL for LLM Reasoning Xian Wu et.al. 2602.07832 null
2026-02-08 Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning Jiahui Zhou et.al. 2602.07830 null
2026-02-08 Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models Xuan Ding et.al. 2602.07804 null
2026-02-08 VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos Wenqi Liu et.al. 2602.07801 null
2026-02-08 Fairness Aware Reward Optimization Ching Lam Choi et.al. 2602.07799 null
2026-02-08 Optimal Control of Unbounded Stochastic Evolution Systems in Hilbert Spaces Shanjian Tang et.al. 2602.07793 null
2026-02-08 Uncertainty-Aware Counterfactual Traffic Signal Control with Predictive Safety and Starvation-Avoidance Constraints Using Vision-Based Sensing Jayawant Bodagala et.al. 2602.07784 null
2026-02-08 CoLF: Learning Consistent Leader-Follower Policies for Vision-Language-Guided Multi-Robot Cooperative Transport Joachim Yann Despature et.al. 2602.07776 null
2026-02-08 Generative Reasoning Re-ranker Mingfu Liang et.al. 2602.07774 null
2026-02-08 Compressed Sensing Methods for Memory Reduction in Monte Carlo Simulations Ethan Lame et.al. 2602.07771 null
2026-02-08 Structure Preserving Approximation of Semiconcave Functions Karl Kunisch et.al. 2602.07770 null
2026-02-08 BFTS: Thompson Sampling with Bayesian Additive Regression Trees Ruizhe Deng et.al. 2602.07767 null
2026-02-08 Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Tanmay Ambadkar et.al. 2602.07764 null
2026-02-08 The Development of a Preclinical Alpha Irradiation Platform with Versatile Control of Dose, Dose Rate, and Spatiotemporal Irradiation Patterns Harsh Arya et.al. 2602.07760 null
2026-02-07 The Laplacian Keyboard: Beyond the Linear Span Siddarth Chandrasekar et.al. 2602.07730 null
2026-02-06 MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images Ankan Deria et.al. 2602.06965 null
2026-02-06 InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning Yuchen Yan et.al. 2602.06960 null
2026-02-06 Optimal Derivative Feedback Control for an Active Magnetic Levitation System: An Experimental Study on Data-Driven Approaches Saber Omidi et.al. 2602.06944 null
2026-02-06 Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics Zuyuan Zhang et.al. 2602.06939 null
2026-02-06 When RL Meets Adaptive Speculative Training: A Unified Training-Serving System Junxiong Wang et.al. 2602.06932 null
2026-02-06 On micromodes in Bayesian posterior distributions and their implications for MCMC Sanket Agrawal et.al. 2602.06931 null
2026-02-06 Continuous-time reinforcement learning: ellipticity enables model-free value function approximation Wenlong Mou et.al. 2602.06930 null
2026-02-06 Towards a Fully Automated Pipeline for Short-Term Forecasting of In Situ Coronal Mass Ejection Magnetic Field Structure Hannah T. Rüdisser et.al. 2602.06926 null
2026-02-06 A first realization of reinforcement learning-based closed-loop EEG-TMS Dania Humaidan et.al. 2602.06907 null
2026-02-06 SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks Mingqian Feng et.al. 2602.06854 null
2026-02-06 AEGPO: Adaptive Entropy-Guided Policy Optimization for Diffusion Models Yuming Li et.al. 2602.06825 null
2026-02-06 On the Design of an Optimal Multi-Tone Jammer Against the Wiener Interpolation Filter Corentin Fonteneau et.al. 2602.06816 null
2026-02-06 Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling Kate Sanders et.al. 2602.06795 null
2026-02-06 UnifSrv: AP Selection for Achieving Uniformly Good Performance of CF-MIMO in Realistic Urban Networks Yunlu Xiao et.al. 2602.06780 null
2026-02-06 Soft Forward-Backward Representations for Zero-shot Reinforcement Learning with General Utilities Marco Bagatella et.al. 2602.06769 null
2026-02-06 R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging Yanlin Lai et.al. 2602.06763 null
2026-02-06 Semantically Labelled Automata for Multi-Task Reinforcement Learning with LTL Instructions Alessandro Abate et.al. 2602.06746 null
2026-02-06 F-GRPO: Don’t Let Your Policy Learn the Obvious and Forget the Rare Daniil Plyusov et.al. 2602.06717 null
2026-02-06 Evaluating and Enhancing the Vulnerability Reasoning Capabilities of Large Language Models Li Lu et.al. 2602.06687 null
2026-02-06 compar:IA: The French Government’s LLM arena to collect French-language human prompts and preference data Lucie Termignon et.al. 2602.06669 null
2026-02-05 InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions Sirui Xu et.al. 2602.06035 null
2026-02-05 V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval Dongyang Chen et.al. 2602.06034 null
2026-02-05 Can vision language models learn intuitive physics from interaction? Luca M. Schulze Buschoff et.al. 2602.06033 null
2026-02-05 Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory Haozhen Zhang et.al. 2602.06025 null
2026-02-05 On Computation and Reinforcement Learning Raj Ghugare et.al. 2602.05999 null
2026-02-05 VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation Jie Deng et.al. 2602.05998 null
2026-02-05 Diamond Maps: Efficient Reward Alignment via Stochastic Flow Maps Peter Holderrieth et.al. 2602.05993 null
2026-02-05 Time lags and their association with the Boundary Layer structure in a Z source GX 349+2 Abhishek M. V. R. et.al. 2602.05989 null
2026-02-05 Clifford Kolmogorov-Arnold Networks Matthias Wolff et.al. 2602.05977 null
2026-02-05 Learning to Share: Selective Memory for Efficient Parallel Agentic Systems Joseph Fioresi et.al. 2602.05965 null
2026-02-05 A POWHEG generator for di-jet production in polarized proton-proton collisions Ignacio Borsa et.al. 2602.05949 null
2026-02-05 $f$ -GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment Rajdeep Haldar et.al. 2602.05946 null
2026-02-05 Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training Zhenghao Xu et.al. 2602.05933 null
2026-02-05 Quantum Reinforcement Learning with Transformers for the Capacitated Vehicle Routing Problem Eva Andrés et.al. 2602.05920 null
2026-02-05 Reducing the Computational Cost Scaling of Tensor Network Algorithms via Field-Programmable Gate Array Parallelism Songtai Lv et.al. 2602.05900 null
2026-02-05 Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models Shuo Nie et.al. 2602.05897 null
2026-02-05 Residual Reinforcement Learning for Waste-Container Lifting Using Large-Scale Cranes with Underactuated Tools Qi Li et.al. 2602.05895 null
2026-02-05 DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training Dingwei Zhu et.al. 2602.05890 null
2026-02-05 Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations Wei Liu et.al. 2602.05885 null
2026-02-05 UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents Han Xiao et.al. 2602.05832 null
2026-02-04 Reinforced Attention Learning Bangzheng Li et.al. 2602.04884 null
2026-02-04 Rethinking the Trust Region in LLM Reinforcement Learning Penghui Qi et.al. 2602.04879 null
2026-02-04 CRoSS: A Continual Robotic Simulation Suite for Scalable Reinforcement Learning with High Task Diversity and Realistic Physics Simulation Yannick Denker et.al. 2602.04868 null
2026-02-04 PDF-HR: Pose Distance Fields for Humanoid Robots Yi Gu et.al. 2602.04851 null
2026-02-04 Safe Urban Traffic Control via Uncertainty-Aware Conformal Prediction and World-Model Reinforcement Learning Joydeep Chandra et.al. 2602.04821 null
2026-02-04 Site and bond percolation in linearly distorted triangular and square lattices Bishnu Bhowmik et.al. 2602.04818 null
2026-02-04 Beyond Rewards in Reinforcement Learning for Cyber Defence Elizabeth Bates et.al. 2602.04809 null
2026-02-04 Joint Sleep Mode Activation and Load Balancing with Dynamic Cell Load: A Combinatorial Bandit Approach Wajahat Bashir Gilkar et.al. 2602.04808 null
2026-02-04 Evolving Afferent Architectures: Biologically-inspired Models for Damage-Avoidance Learning Wolfgang Maass et.al. 2602.04807 null
2026-02-04 Skin Tokens: A Learned Compact Representation for Unified Autoregressive Rigging Jia-peng Zhang et.al. 2602.04805 null
2026-02-04 Benchmark Study of CEvNS Nuclear Recoil Observables for B, Mg, Ti, and Zr Targets Using Geant4 Yusuf Havvat et.al. 2602.04771 null
2026-02-04 Control Lyapunov Functions for Optimality in Sontag-Type Control Joscha F. Bongard et.al. 2602.04756 null
2026-02-04 When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond? Xinyu Zhou et.al. 2602.04755 null
2026-02-04 Multiple Imputation Methods under Extreme Values Enzo Porto Brasil et.al. 2602.04751 null
2026-02-04 Rationality Measurement and Theory for Reinforcement Learning Agents Kejiang Qian et.al. 2602.04737 null
2026-02-04 ERNIE 5.0 Technical Report Haifeng Wang et.al. 2602.04705 null
2026-02-04 Multi-Source Retrieval and Reasoning for Legal Sentencing Prediction Junjie Chen et.al. 2602.04690 null
2026-02-04 Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design Jaemoo Choi et.al. 2602.04663 null
2026-02-04 SAFE: Stable Alignment Finetuning with Entropy-Aware Predictive Control for RLHF Dipan Maity et.al. 2602.04651 null
2026-02-04 Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models Binghai Wang et.al. 2602.04649 null
2026-02-03 Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL Erfan Miahi et.al. 2602.03839 null
2026-02-03 SymPlex: A Structure-Aware Transformer for Symbolic PDE Solving Yesom Park et.al. 2602.03816 null
2026-02-03 Vacancy defects in square-triangle tilings and their implications for quasicrystals formed by square-shoulder particles Alptuğ Ulugöl et.al. 2602.03813 null
2026-02-03 Bridging Online and Offline RL: Contextual Bandit Learning for Multi-Turn Code Generation Ziru Chen et.al. 2602.03806 null
2026-02-03 Digital-Twin Empowered Deep Reinforcement Learning For Site-Specific Radio Resource Management in NextG Wireless Aerial Corridor Pulok Tarafder et.al. 2602.03801 null
2026-02-03 Efficient Estimation of Kernel Surrogate Models for Task Attribution Zhenshuo Zhang et.al. 2602.03783 null
2026-02-03 Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity Aneri Muni et.al. 2602.03778 null
2026-02-03 Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL Ian Wu et.al. 2602.03773 null
2026-02-03 RegionReasoner: Region-Grounded Multi-Round Visual Reasoning Wenfang Sun et.al. 2602.03733 null
2026-02-03 Efficient Variance-reduced Estimation from Generative EHR Models: The SCOPE and REACH Estimators Luke Solo et.al. 2602.03730 null
2026-02-03 Quantum Speedups for Derivative Pricing Beyond Black-Scholes Dylan Herman et.al. 2602.03725 null
2026-02-03 Training Multi-Turn Search Agent via Contrastive Dynamic Branch Sampling Yubao Zhao et.al. 2602.03719 null
2026-02-03 Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation Jiashuo Sun et.al. 2602.03689 null
2026-02-03 TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System Wenzhe Fan et.al. 2602.03688 null
2026-02-03 Resolving Quantum Criticality in the Honeycomb Hubbard Model Fo-Hong Wang et.al. 2602.03656 null
2026-02-03 Search-R2: Enhancing Search-Integrated Reasoning via Actor-Refiner Collaboration Bowei He et.al. 2602.03647 null
2026-02-03 Reinforcement Fine-Tuning for History-Aware Dense Retriever in RAG Yicheng Zhang et.al. 2602.03645 null
2026-02-03 TRE: Encouraging Exploration in the Trust Region Chao Huang et.al. 2602.03635 null
2026-02-03 A Method for Thermal Radiation Transport Using Backward Characteristic Tracing J. C. Dolence et.al. 2602.03621 null
2026-02-03 Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation Changze Lv et.al. 2602.03619 null
2026-02-02 Reward-free Alignment for Conflicting Objectives Peter Chen et.al. 2602.02495 null
2026-02-02 RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System Yinjie Wang et.al. 2602.02488 null
2026-02-02 Expanding the Capabilities of Reinforcement Learning via Text Feedback Yuda Song et.al. 2602.02482 null
2026-02-02 Flow Policy Gradients for Robot Control Brent Yi et.al. 2602.02481 null
2026-02-02 Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability Xiao Liang et.al. 2602.02477 null
2026-02-02 HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos Yinhuai Wang et.al. 2602.02473 null
2026-02-02 TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments Zhiyu Huang et.al. 2602.02459 null
2026-02-02 Conflict-Aware Client Selection for Multi-Server Federated Learning Mingwei Hong et.al. 2602.02458 null
2026-02-02 World-Gymnast: Training Robots with Reinforcement Learning in a World Model Ansh Kumar Sharma et.al. 2602.02454 null
2026-02-02 Multi-Agent Monte Carlo Tree Search for Makespan-Efficient Object Rearrangement in Cluttered Spaces Hanwen Ren et.al. 2602.02411 null
2026-02-02 Didactic to Constructive: Turning Expert Solutions into Learnable Reasoning Ethan Mendes et.al. 2602.02405 null
2026-02-02 PRISM: Performer RS-IMLE for Single-pass Multisensory Imitation Learning Amisha Bhaskar et.al. 2602.02396 null
2026-02-02 David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning Samuel Nellessen et.al. 2602.02395 null
2026-02-02 SLIME: Stabilized Likelihood Implicit Margin Enforcement for Preference Optimization Maksim Afanasyev et.al. 2602.02383 null
2026-02-02 Unified Personalized Reward Model for Vision Generation Yibin Wang et.al. 2602.02380 null
2026-02-02 Proof-RM: A Scalable and Generalizable Reward Model for Math Proof Haotong Yang et.al. 2602.02377 null
2026-02-02 SWE-Universe: Scale Real-World Verifiable Environments to Millions Mouxiang Chen et.al. 2602.02361 null
2026-02-02 Surface Interactions in Photon Monte Carlo Simulations J. R. Peterson et.al. 2602.02321 null
2026-02-02 Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient Changming Li et.al. 2602.02313 null
2026-02-02 Position: Explaining Behavioral Shifts in Large Language Models Requires a Comparative Approach Martino Ciaperoni et.al. 2602.02304 null
2026-01-30 IRL-DAL: Safe and Adaptive Trajectory Planning for Autonomous Driving via Energy-Guided Diffusion Models Seyed Ahmad Hosseini Miangoleh et.al. 2601.23266 null
2026-01-30 Particle-Guided Diffusion Models for Partial Differential Equations Andrew Millard et.al. 2601.23262 null
2026-01-30 How well do generative models solve inverse problems? A benchmark study Patrick Krüger et.al. 2601.23238 null
2026-01-30 Agile Reinforcement Learning through Separable Neural Architecture Rajib Mostakim et.al. 2601.23225 null
2026-01-30 Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning Xiangyu Zeng et.al. 2601.23224 null
2026-01-30 Med-Scout: Curing MLLMs’ Geometric Blindness in Medical Perception via Geometry-Aware RL Post-Training Anglin Liu et.al. 2601.23220 null
2026-01-30 Unsupervised Hierarchical Skill Discovery Damion Harvey et.al. 2601.23156 null
2026-01-30 On Safer Reinforcement Learning Policies for Sedation and Analgesia in Intensive Care Joel Romero-Hernandez et.al. 2601.23154 null
2026-01-30 THINKSAFE: Self-Generated Safety Alignment for Reasoning Models Seanie Lee et.al. 2601.23143 null
2026-01-30 Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients Cheng Ge et.al. 2601.23135 null
2026-01-30 TopoLS: Lattice Surgery Compilation via Topological Program Transformations Junyu Zhou et.al. 2601.23109 null
2026-01-30 Temporally Coherent Imitation Learning via Latent Action Flow Matching for Robotic Manipulation Wu Songwei et.al. 2601.23087 null
2026-01-30 RN-D: Discretized Categorical Actors with Regularized Networks for On-Policy Reinforcement Learning Yuexin Bian et.al. 2601.23075 null
2026-01-30 From Absolute to Relative: Rethinking Reward Shaping in Group-Based Reinforcement Learning Wenzhe Niu et.al. 2601.23058 null
2026-01-30 Theory of Little-Parks oscillations by vortices in two-dimensional superconductors Ying-Ming Xie et.al. 2601.23050 null
2026-01-30 Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning Siyu Gong et.al. 2601.23032 null
2026-01-30 Mem-T: Densifying Rewards for Long-Horizon Memory Agents Yanwei Yue et.al. 2601.23014 null
2026-01-30 Automatic Constraint Policy Optimization based on Continuous Constraint Interpolation Framework for Offline Reinforcement Learning Xinchen Han et.al. 2601.23010 null
2026-01-30 Unlocking the Power of Orbital-Free Density Functional Theory to Explore the Electronic Structure Under Extreme Conditions Cheng Ma et.al. 2601.23002 null
2026-01-30 Computationally efficient segmentation for non-stationary time series with oscillatory patterns Nicolas Bianco et.al. 2601.22999 null
2026-01-29 Exploring Reasoning Reward Model for Agents Kaixuan Fan et.al. 2601.22154 null
2026-01-29 DynaWeb: Model-Based Reinforcement Learning of Web Agents Hang Ding et.al. 2601.22149 null
2026-01-29 Alpha Discovery via Grammar-Guided Learning and Search Han Yang et.al. 2601.22119 null
2026-01-29 Defining Operational Conditions for Safety-Critical AI-Based Systems from Data Johann Christensen et.al. 2601.22118 null
2026-01-29 Boosting CVaR Policy Optimization with Quantile Gradients Yudong Luo et.al. 2601.22100 null
2026-01-29 A Gradient-Based Capacity Accreditation Framework in Resource Adequacy: Formulation, Computation, and Practical Implications Qian Zhang et.al. 2601.22087 null
2026-01-29 Learning to Dial-a-Ride: A Deep Graph Reinforcement Learning Approach to the Electric Dial-a-Ride Problem Sten Elling Tingstad Jacobsen et.al. 2601.22052 null
2026-01-29 SIA: Symbolic Interpretability for Anticipatory Deep Reinforcement Learning in Network Control MohammadErfan Jabbari et.al. 2601.22044 null
2026-01-29 SymbXRL: Symbolic Explainable Deep Reinforcement Learning for Mobile Networks Abhishek Duttagupta et.al. 2601.22024 null
2026-01-29 Post-Disaster Resource Redistribution and Cooperation Evolution Based on Two-Layer Network Evolutionary Games Yu Chen et.al. 2601.22021 null
2026-01-29 Efficient Stochastic Optimisation via Sequential Monte Carlo James Cuin et.al. 2601.22003 null
2026-01-29 Geometry of Drifting MDPs with Path-Integral Stability Certificates Zuyuan Zhang et.al. 2601.21991 null
2026-01-29 Elign: Equivariant Diffusion Model Alignment from Foundational Machine Learning Force Fields Yunyang Li et.al. 2601.21985 null
2026-01-29 Investigating Batch Inference in a Sequential Monte Carlo Framework for Neural Networks Andrew Millard et.al. 2601.21983 null
2026-01-29 Investigation into using stochastic embedding representations for evaluating the trustworthiness of the Fréchet Inception Distance Ciaran Bench et.al. 2601.21979 null
2026-01-29 Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic Shuo Liu et.al. 2601.21972 null
2026-01-29 MoE-ACT: Improving Surgical Imitation Learning Policies through Supervised Mixture-of-Experts Lorenzo Mazza et.al. 2601.21971 null
2026-01-29 Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding Yifan Zhu et.al. 2601.21969 null
2026-01-29 OVD: On-policy Verbal Distillation Jing Xiong et.al. 2601.21968 null
2026-01-29 From Tokens to Blocks: A Block-Diffusion Perspective on Molecular Generation Qianwei Yang et.al. 2601.21964 null
2026-01-28 End-to-end example-based sim-to-real RL policy transfer based on neural stylisation with application to robotic cutting Jamie Hathaway et.al. 2601.20846 null
2026-01-28 Reward Models Inherit Value Biases from Pretraining Brian Christian et.al. 2601.20838 null
2026-01-28 MemCtrl: Using MLLMs as Active Memory Controllers on Embodied Agents Vishnu Sashank Dorbala et.al. 2601.20831 null
2026-01-28 Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning Minwu Kim et.al. 2601.20829 null
2026-01-28 Reinforcement Learning via Self-Distillation Jonas Hübotter et.al. 2601.20802 null
2026-01-28 SERA: Soft-Verified Efficient Repository Agents Ethan Shen et.al. 2601.20789 null
2026-01-28 Neural Quantum States in Mixed Precision Massimo Solinas et.al. 2601.20782 null
2026-01-28 Less is More: Clustered Cross-Covariance Control for Offline RL Nan Qiao et.al. 2601.20765 null
2026-01-28 Exploring Re-inforcement Learning via Human Feedback under User Heterogeneity Sarvesh Shashidhar et.al. 2601.20760 null
2026-01-28 GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Zhiheng Jiang et.al. 2601.20753 null
2026-01-28 Comparing causal estimands from sequential nested versus single point target trials: A simulation study Catherine Wiener et.al. 2601.20725 null
2026-01-28 From biting to engulfment: curvature-actin coupling controls phagocytosis of soft, deformable targets Shubhadeep Sadhukhan et.al. 2601.20719 null
2026-01-28 Adapting the Behavior of Reinforcement Learning Agents to Changing Action Spaces and Reward Functions Raul de la Rosa et.al. 2601.20714 null
2026-01-28 A scalable flow-based approach to mitigate topological freezing Claudio Bonanno et.al. 2601.20708 null
2026-01-28 One Step Is Enough: Dispersive MeanFlow Policy Optimization Guowei Zou et.al. 2601.20701 null
2026-01-28 Tensor renormalization group study of cold and dense QCD in the strong coupling limit Yuto Sugimoto et.al. 2601.20690 null
2026-01-28 Grover’s Search-Inspired Quantum Reinforcement Learning for Massive MIMO User Scheduling Ruining Fan et.al. 2601.20688 null
2026-01-28 Positive-Unlabeled Reinforcement Learning Distillation for On-Premise Small Models Zhiqiang Kou et.al. 2601.20687 null
2026-01-28 GPO: Growing Policy Optimization for Legged Robot Locomotion and Whole-Body Control Shuhao Liao et.al. 2601.20668 null
2026-01-28 Deep Learning based Three-stage Solution for ISAC Beamforming Optimization Qian Gao et.al. 2601.20667 null
2026-01-27 Self-Distillation Enables Continual Learning Idan Shenfeld et.al. 2601.19897 null
2026-01-27 A Latent Space Framework for Modeling Transient Engine Emissions Using Joint Embedding Predictive Architectures Ganesh Sundaram et.al. 2601.19822 null
2026-01-27 Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed Goals Octavio Pappalardo et.al. 2601.19810 null
2026-01-27 Reimagining Peer Review Process Through Multi-Agent Mechanism Design Ahmad Farooq et.al. 2601.19778 null
2026-01-27 Reimagining Social Robots as Recommender Systems: Foundations, Framework, and Applications Jin Huang et.al. 2601.19761 null
2026-01-27 Run Dependent Monte Carlo at Belle II Giovanni Gaudino et.al. 2601.19728 null
2026-01-27 Zeroth-order parallel sampling Francesco Pozza et.al. 2601.19722 null
2026-01-27 Improving Policy Exploitation in Online Reinforcement Learning with Instant Retrospect Action Gong Gao et.al. 2601.19720 null
2026-01-27 Scalable Exploration for High-Dimensional Continuous Control via Value-Guided Flow Yunyue Wei et.al. 2601.19707 null
2026-01-27 AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion Tianyue Jiang et.al. 2601.19697 null
2026-01-27 Video-KTR: Reinforcing Video Reasoning via Key Token Attribution Ziyue Wang et.al. 2601.19686 null
2026-01-27 Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning Tongxi Wang et.al. 2601.19624 null
2026-01-27 R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning Zhizheng Jiang et.al. 2601.19620 null
2026-01-27 On the rarity of rocket-driven Penrose extraction in Kerr spacetime An T. Le et.al. 2601.19616 null
2026-01-27 Safe Exploration via Policy Priors Manuel Wendl et.al. 2601.19612 null
2026-01-27 LLM-Enhanced Reinforcement Learning for Long-Term User Satisfaction in Interactive Recommendation Chongjun Xia et.al. 2601.19585 null
2026-01-27 A Fast, Closed-Form Bandwidth Selector for the Beta Kernel Density Estimator Johan Hallberg Szabadváry et.al. 2601.19553 null
2026-01-27 Generalizable Equivariant Diffusion Models for Non-Abelian Lattice Gauge Theory Gert Aarts et.al. 2601.19552 null
2026-01-27 Radiative return at NLOPS accuracy Ettore Budassi et.al. 2601.19530 null
2026-01-27 Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration Zhengjian Yao et.al. 2601.19506 null
2026-01-26 Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes Amrith Setlur et.al. 2601.18795 null
2026-01-26 Multi-Objective Reinforcement Learning for Efficient Tactical Decision Making for Trucks in Highway Traffic Deepthi Pathare et.al. 2601.18783 null
2026-01-26 OptiGAN for Crystal Arrays: Physics-Informed Generative Modeling of Optical Photon Transport in PET Detector Arrays Stephan Naunheim et.al. 2601.18780 null
2026-01-26 POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration Yuxiao Qu et.al. 2601.18779 null
2026-01-26 Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability Shobhita Sundaram et.al. 2601.18778 null
2026-01-26 Dep-Search: Learning Dependency-Aware Reasoning Traces with Persistent Memory Yanming Liu et.al. 2601.18771 null
2026-01-26 Trust, Don’t Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback Seyed Amir Hosseini et.al. 2601.18751 null
2026-01-26 Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models Siyan Zhao et.al. 2601.18734 null
2026-01-26 One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment Hongru Cai et.al. 2601.18731 null
2026-01-26 Reflect: Transparent Principle-Guided Reasoning for Constitutional Alignment at Scale Henry Bell et.al. 2601.18730 null
2026-01-26 Trustworthy Evaluation of Robotic Manipulation: A New Benchmark and AutoEval Methods Mengyuan Liu et.al. 2601.18723 null
2026-01-26 Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs Zhichao Yang et.al. 2601.18706 null
2026-01-26 ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule Yilie Huang et.al. 2601.18681 null
2026-01-26 New parameters for star cluster dynamics: observational results Barbara Lanzoni et.al. 2601.18679 null
2026-01-26 Quasi Monte Carlo methods enable extremely low-dimensional deep generative models Miles Martinez et.al. 2601.18676 null
2026-01-26 McSAS3: improved Monte Carlo small-angle scattering analysis software for dilute and dense scatterers Brian Richard Pauw et.al. 2601.18659 null
2026-01-26 AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning Mingyang Song et.al. 2601.18631 null
2026-01-26 Rank-1 Approximation of Inverse Fisher for Natural Policy Gradients in Deep Reinforcement Learning Yingxiao Huo et.al. 2601.18626 null
2026-01-26 Learning long term climate-resilient transport adaptation pathways under direct and indirect flood impacts using reinforcement learning Miguel Costa et.al. 2601.18586 null
2026-01-26 From Classification to Ranking: Enhancing LLM Reasoning Capabilities for MBTI Personality Detection Yuan Cao et.al. 2601.18582 null
2026-01-23 Autonomous Optical Alignment of Satellite-Based Entanglement Sources using Reinforcement Learning Andrzej Gajewski et.al. 2601.16968 null
2026-01-23 The Trajectory Alignment Coefficient in Two Acts: From Reward Tuning to Reward Learning Calarina Muslimani et.al. 2601.16906 null
2026-01-23 Boosting Deep Reinforcement Learning with Semantic Knowledge for Robotic Manipulators Lucía Güitta-López et.al. 2601.16866 null
2026-01-23 Distributional Instruments: Identification and Estimation with Quantile Least Squares Rowan Cherodian et.al. 2601.16865 null
2026-01-23 Reasoning Promotes Robustness in Theory of Mind Tasks Ian B. de Haan et.al. 2601.16853 null
2026-01-23 Flux-ratio anomalies in cusp quasars reveal dark matter beyond CDM Siyuan Hou et.al. 2601.16818 null
2026-01-23 Length spectrum rigidity and flexibility of spheres of revolution with one equator Alberto Abbondandolo et.al. 2601.16804 null
2026-01-23 Improved Kelbg Potentials for $Z>1$ and Application to Carbon Plasmas Heather D. Whitley et.al. 2601.16794 null
2026-01-23 Finite Population Inference for Factorial Designs and Panel Experiments with Imperfect Compliance Pedro Picchetti et.al. 2601.16749 null
2026-01-23 LongCat-Flash-Thinking-2601 Technical Report Meituan LongCat Team et.al. 2601.16725 null
2026-01-23 Faster parallel MCMC: Metropolis adjustment is best served warm Jakob Robnik et.al. 2601.16696 null
2026-01-23 Adaptive Reinforcement and Model Predictive Control Switching for Safe Human-Robot Cooperative Navigation Ning Liu et.al. 2601.16686 null
2026-01-23 Sim-to-Real Transfer via a Style-Identified Cycle Consistent Generative Adversarial Network: Zero-Shot Deployment on Robotic Manipulators through Visual Domain Adaptation Lucía Güitta-López et.al. 2601.16677 null
2026-01-23 Inference from high-frequency data: A subsampling approach Kim Christensen et.al. 2601.16668 null
2026-01-23 A Cognitive Framework for Autonomous Agents: Toward Human-Inspired Design Francesco Guidi et.al. 2601.16648 null
2026-01-23 Is the diurnal pattern sufficient to explain intraday variation in volatility? A nonparametric assessment Kim Christensen et.al. 2601.16613 null
2026-01-23 Variational approximate penalized credible regions for Bayesian grouped regression Weichang Yu et.al. 2601.16585 null
2026-01-23 Necessary Optimality Conditions for Integrated Learning and Optimization Problem in Contextual Optimization Yuan Tao et.al. 2601.16581 null
2026-01-23 Zero-Shot MARL Benchmark in the Cyber-Physical Mobility Lab Julius Beerwerth et.al. 2601.16578 null
2026-01-23 Spiking Neural Networks for Communication Systems: Encoding Schemes, Learning Algorithms, and Equalization~Techniques Eike-Manuel Edelmann et.al. 2601.16550 null
2026-01-22 CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback Wenhang Ge et.al. 2601.16214 null
2026-01-22 LLM-in-Sandbox Elicits General Agentic Intelligence Daixuan Cheng et.al. 2601.16206 null
2026-01-22 Cyclic sunspot activity during the first millennium CE as reconstructed from radiocarbon Ilya Usoskin et.al. 2601.16203 null
2026-01-22 Constraining dark energy models using Jackknife and Bootstrap resampling Roshna K et.al. 2601.16197 null
2026-01-22 Learning to Discover at Test Time Mert Yuksekgonul et.al. 2601.16175 null
2026-01-22 Structured Hints for Sample-Efficient Lean Theorem Proving Zachary Burton et.al. 2601.16172 null
2026-01-22 Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning Moo Jin Kim et.al. 2601.16163 null
2026-01-22 Efficiently Learning Robust Torque-based Locomotion Through Reinforcement with Model-Based Supervision Yashuai Yan et.al. 2601.16109 null
2026-01-22 SAMTok: Representing Any Mask with Two Words Yikang Zhou et.al. 2601.16093 null
2026-01-22 A forward-only scheme for online learning of proposal distributions in particle filters Sylvain Procope-Mamert et.al. 2601.16089 null
2026-01-22 Improve the autonomy of the SE2(3) group based Extended Kalman Filter for Integrated Navigation: Application Jiarui Cui et.al. 2601.16078 null
2026-01-22 Dynamic Tactile Sensing System and Soft Actor Critic Reinforcement Learning for Inclusion Characterization John Bannan et.al. 2601.16061 null
2026-01-22 A Fast Monte Carlo Newton-Raphson Algorithm to Estimate Generalized Linear Mixed Models with Dense Covariance Samuel I. Watson et.al. 2601.16022 null
2026-01-22 Keyframe-Based Feed-Forward Visual Odometry Weichen Dai et.al. 2601.16020 null
2026-01-22 PUMA: Perception-driven Unified Foothold Prior for Mobility Augmented Quadruped Parkour Liang Wang et.al. 2601.15995 null
2026-01-22 Decoupling Return-to-Go for Efficient Decision Transformer Yongyi Wang et.al. 2601.15953 null
2026-01-22 Stochastically forced compressible Navier-Stokes equations with slip boundary conditions of friction type Reo Tsuboya et.al. 2601.15768 null
2026-01-22 Three’s a crowd: Identification challenges in the triple difference model with spillover effects Silvia De Nicolò et.al. 2601.15764 null
2026-01-22 Off-Policy Actor-Critic with Sigmoid-Bounded Entropy for Real-World Robot Learning Xiefeng Wu et.al. 2601.15761 null
2026-01-22 PhysProver: Advancing Automatic Theorem Proving for Physics Hanning Zhang et.al. 2601.15737 null
2026-01-21 Bhabha scattering at future colliders with BHLUMI/BHWIDE Wiesław Płaczek et.al. 2601.15265 null
2026-01-21 Assessing Orbital Optimization in Variational and Diffusion Monte Carlo Cody A. Melton et.al. 2601.15169 null
2026-01-21 DeGAS: Gradient-Based Optimization of Probabilistic Programs without Sampling Francesca Randone et.al. 2601.15167 null
2026-01-21 The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Models Zanlin Ni et.al. 2601.15165 null
2026-01-21 Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning Yuval Kansal et.al. 2601.15160 null
2026-01-21 Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data Yuval Ran-Milo et.al. 2601.15158 null
2026-01-21 CLEANER: Self-Purified Trajectories Boost Agentic Reinforcement Learning Tianshi Xu et.al. 2601.15141 null
2026-01-21 Efficient prior sensitivity analysis for Bayesian model comparison Zixiao Hu et.al. 2601.15132 null
2026-01-21 Vehicle Routing with Finite Time Horizon using Deep Reinforcement Learning with Improved Network Embedding Ayan Maity et.al. 2601.15131 null
2026-01-21 Memory Retention Is Not Enough to Master Memory Tasks in Reinforcement Learning Oleg Shchendrigin et.al. 2601.15086 null
2026-01-21 A combined dose and microdosimetric modeling framework incorporating volume effects correlates with tissue sparing in proton minibeam radiotherapy Giulio Bordieri et.al. 2601.15073 null
2026-01-21 A Curriculum-Based Deep Reinforcement Learning Framework for the Electric Vehicle Routing Problem Mertcan Daysalilar et.al. 2601.15038 null
2026-01-21 Physical Layer Security in Massive MIMO: Challenges and Open Research Directions Against Passive Eavesdroppers Nipun Agarwal et.al. 2601.15024 null
2026-01-21 Plug-and-Play Benchmarking of Reinforcement Learning Algorithms for Large-Scale Flow Control Jannis Becktepe et.al. 2601.15015 null
2026-01-21 Alternative Shapes of Modulation Schemes Detailed Exposition and Simulation Methodology Nipun Agarwal et.al. 2601.15004 null
2026-01-21 Ab initio path-integral Monte Carlo results for the one-particle spectral function of the warm dense electron gas Paul Hamann et.al. 2601.14992 null
2026-01-21 Dielectric formalism of the 2D uniform electron gas at finite temperatures Fotios Kalkavouras et.al. 2601.14989 null
2026-01-21 Improving Regret Approximation for Unsupervised Dynamic Environment Generation Harry Mead et.al. 2601.14957 null
2026-01-21 Numerical study of multiple solar flare induced modulation of Very Low Frequency (VLF) diurnal profile Sourav Palit et.al. 2601.14948 null
2026-01-21 Language-Coupled Reinforcement Learning for Multilingual Retrieval-Augmented Generation Rui Qi et.al. 2601.14896 null
2026-01-20 Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow Haocheng Xi et.al. 2601.14243 null
2026-01-20 Spatiotemporal Wildfire Prediction and Reinforcement Learning for Helitack Suppression Shaurya Mathur et.al. 2601.14238 null
2026-01-20 Q-learning with Adjoint Matching Qiyang Li et.al. 2601.14234 null
2026-01-20 KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning Egor Cherepanov et.al. 2601.14232 null
2026-01-20 Attention-Based Offline Reinforcement Learning and Clustering for Interpretable Sepsis Treatment Punit Kumar et.al. 2601.14228 null
2026-01-20 InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning Matthew Y. R. Yang et.al. 2601.14209 null
2026-01-20 Differentiated Pickup Point Offering for Emission Reduction in Last-Mile Delivery Albina Galiullina et.al. 2601.14196 null
2026-01-20 Toward Efficient Agents: Memory, Tool learning, and Planning Xiaofang Yang et.al. 2601.14192 null
2026-01-20 Symmetry Breaking and Phase Transitions in Random Non-Commutative Geometries and Related Random-Matrix Ensembles Mauro D’Arcangelo et.al. 2601.14141 null
2026-01-20 CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems Tong Xie et.al. 2601.14140 null
2026-01-20 Log-optimality with small liability stream Michail Anthropelos et.al. 2601.14139 null
2026-01-20 Diffusion-Guided Backdoor Attacks in Real-World Reinforcement Learning Tairan Huang et.al. 2601.14104 null
2026-01-20 Optimizing Energy and Data Collection in UAV-aided IoT Networks using Attention-based Multi-Objective Reinforcement Learning Babacar Toure et.al. 2601.14092 null
2026-01-20 Onset of stripe order in classical fluids: Lessons from lattice-gas mixtures Gabriele Costa et.al. 2601.14082 null
2026-01-20 Tail-Aware Density Forecasting of Locally Explosive Time Series: A Neural Network Approach Elena Dumitrescu et.al. 2601.14049 null
2026-01-20 RM-Distiller: Exploiting Generative LLM for Reward Model Distillation Hongli Zhou et.al. 2601.14032 null
2026-01-20 BACH-V: Bridging Abstract and Concrete Human-Values in Large Language Models Junyu Zhang et.al. 2601.14007 null
2026-01-20 The Transparency Paradox in Explainable AI: A Theory of Autonomy Depletion Through Cognitive Load Ancuta Margondai et.al. 2601.13973 null
2026-01-20 RL-BioAug: Label-Efficient Reinforcement Learning for Self-Supervised EEG Representation Learning Cheol-Hui Lee et.al. 2601.13964 null
2026-01-20 Glance-or-Gaze: Incentivizing LMMs to Adaptively Focus Search via Reinforcement Learning Hongbo Bai et.al. 2601.13942 null
2026-01-16 Do explanations generalize across large reasoning models? Koyena Pal et.al. 2601.11517 null
2026-01-16 Generative Scenario Rollouts for End-to-End Autonomous Driving Rajeev Yasarla et.al. 2601.11475 null
2026-01-16 Globally Optimal Contour Deformations with Neural Networks Stephen Jones et.al. 2601.11448 null
2026-01-16 When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models Raphaël Razafindralambo et.al. 2601.11444 null
2026-01-16 The Great March 100: 100 Detail-oriented Tasks for Evaluating Embodied AI Agents Ziyu Wang et.al. 2601.11421 null
2026-01-16 New Adaptive Mechanism for Large Neighborhood Search using Dual Actor-Critic Shaohua Yu et.al. 2601.11414 null
2026-01-16 Factored Value Functions for Graph-Based Multi-Agent Reinforcement Learning Ahmed Rashwan et.al. 2601.11401 null
2026-01-16 The Mini Wheelbot Dataset: High-Fidelity Data for Robot Learning Henrik Hose et.al. 2601.11394 null
2026-01-16 Reward Modeling for Scientific Writing Evaluation Furkan Şahinuç et.al. 2601.11374 null
2026-01-16 Offline Reinforcement-Learning-Based Power Control for Application-Agnostic Energy Efficiency Akhilesh Raj et.al. 2601.11352 null
2026-01-16 Optimal Abatement Schedules for Excess Carbon Emissions Towards a Net-Zero Target Hansjoerg Albrecher et.al. 2601.11348 null
2026-01-16 Controlled Interacting Branching Diffusion Processes: A Viscosity Approach Antonio Ocello et.al. 2601.11294 null
2026-01-16 Oriented Triplet $p$ -Wave Pairing from Fermi surface Anisotropy and Nonlocal Attraction Shuning Tan et.al. 2601.11267 null
2026-01-16 Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation Pingzhi Tang et.al. 2601.11258 null
2026-01-16 Model-free policy gradient for discrete-time mean-field control Matthieu Meunier et.al. 2601.11217 null
2026-01-16 Policy-Based Deep Reinforcement Learning Hyperheuristics for Job-Shop Scheduling Problems Sofiene Lassoued et.al. 2601.11189 null
2026-01-16 TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech Girish A. Koushik et.al. 2601.11178 null
2026-01-16 Ring isomorphisms in norm between Banach algebras of continuous complex-valued functions T. Miura et.al. 2601.11165 null
2026-01-16 Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems Zixu Wang et.al. 2601.11147 null
2026-01-16 Deep GraphRAG: A Balanced Approach to Hierarchical Retrieval and Adaptive Integration Yuejie Li et.al. 2601.11144 null
2026-01-15 MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite Matching Changle Qu et.al. 2601.10712 null
2026-01-15 Observation Timelines for the Potential Lunar Impact of Asteroid 2024 YR4 Yifan He et.al. 2601.10666 null
2026-01-15 Dynamics of Late time cosmology in $f(Q,L_{m})$ Gravity with Constraints from DESI DR2 BAO Data Rajdeep Mazumdar et.al. 2601.10627 null
2026-01-15 Institutional AI: A Governance Framework for Distributional AGI Safety Federico Pierucci et.al. 2601.10599 null
2026-01-15 Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay Hao Wang et.al. 2601.10589 null
2026-01-15 Comparison of viscosity solutions for a class of non-linear PDEs on the space of finite nonnegative measures Ibrahim Ekren et.al. 2601.10586 null
2026-01-15 Combinatorial Optimization Augmented Machine Learning Maximilian Schiffer et.al. 2601.10583 null
2026-01-15 Flat-band Ferromagnetism of SU $(N)$ Hubbard Model on the Kagome Lattices Hao Jin et.al. 2601.10549 null
2026-01-15 PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models Chengbing Wang et.al. 2601.10532 null
2026-01-15 Scalable Algorithms for Approximate DNF Model Counting Paul Burkhardt et.al. 2601.10511 null
2026-01-15 Projected Microbatch Accumulation yields reference-free proximal policy updates for reinforcement learning Nilin Abrahamsen et.al. 2601.10498 null
2026-01-15 Urban Socio-Semantic Segmentation with Vision-Language Reasoning Yu Wang et.al. 2601.10477 null
2026-01-15 DeFlow: Decoupling Manifold Modeling and Value Maximization for Offline Policy Extraction Zhancun Mu et.al. 2601.10471 null
2026-01-15 Active interrogation of underground piezoelectric fabrics using high energy muon beams propagating across seismogenic faults L. Serafini et.al. 2601.10430 null
2026-01-15 On the reconstruction of kinematic distributions computed with Monte Carlo methods using orthogonal basis functions Kirill Melnikov et.al. 2601.10420 null
2026-01-15 Reinforcement Learning with Multi-Step Lookahead Information Via Adaptive Batching Nadav Merlis et.al. 2601.10418 null
2026-01-15 CS-GBA: A Critical Sample-based Gradient-guided Backdoor Attack for Offline Reinforcement Learning Yuanjie Zhao et.al. 2601.10407 null
2026-01-15 Discrete Feynman-Kac Correctors Mohsin Hasan et.al. 2601.10403 null
2026-01-15 A comparison of simulation tools for Muon-Induced X-ray Emission (MIXE) in thin films: a study case with lithium batteries Maxime Lamotte et.al. 2601.10401 null
2026-01-15 Algebraic Farkas Lemma and Strong Duality for Perturbed Conic Linear Programming P. D. Khanh et.al. 2601.10390 null
2026-01-14 Revisiting Jahn–Teller Transitions in Correlated Oxides with Monte Carlo Modeling Liam A. V. Nagle-Cocco et.al. 2601.09705 null
2026-01-14 Bayesian Semi-Blind Deconvolution at Scale Guillermina Senn et.al. 2601.09677 null
2026-01-14 STEP3-VL-10B Technical Report Ailin Huang et.al. 2601.09668 null
2026-01-14 Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning Zhiyuan Hu et.al. 2601.09667 null
2026-01-14 DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing Qian Cao et.al. 2601.09609 null
2026-01-14 Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets Jeremiah Coholich et.al. 2601.09605 null
2026-01-14 Higgs Decays at NLO in the SMEFT Luigi Bellafronte et.al. 2601.09599 null
2026-01-14 Dialogue Telemetry: Turn-Level Instrumentation for Autonomous Information Gathering Dimitris Panagopoulos et.al. 2601.09570 null
2026-01-14 Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations Wei-Jin Huang et.al. 2601.09518 null
2026-01-14 Terminally constrained flow-based generative models from an optimal control perspective Weiguo Gao et.al. 2601.09474 null
2026-01-14 Data Scaling for Navigation in Unknown Environments Lauri Suomela et.al. 2601.09444 null
2026-01-14 Draw it like Euclid: Teaching transformer models to generate CAD profiles using ruler and compass construction steps Siyi Li et.al. 2601.09428 null
2026-01-14 Semi-Contention-Free Access in IoT NOMA Networks: A Reinforcement Learning Framework Abhishek Kumar et.al. 2601.09422 null
2026-01-14 Semi-physical Gamma-Process Degradation Modeling and Performance-Driven Opportunistic Maintenance Optimization for LED Lighting Systems Haohao Shi et.al. 2601.09380 null
2026-01-14 GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR Jiaying Zhang et.al. 2601.09361 null
2026-01-14 Monte-Carlo Tree Search with Neural Network Guidance for Lane-Free Autonomous Driving Ioannis Peridis et.al. 2601.09353 null
2026-01-14 Optimal control of McKean-Vlasov systems under partial observation and hidden Markov switching Marco Fuhrman et.al. 2601.09311 null
2026-01-14 Policy-Based Reinforcement Learning with Action Masking for Dynamic Job Shop Scheduling under Uncertainty: Handling Random Arrivals and Machine Failures Sofiene Lassoued et.al. 2601.09293 null
2026-01-14 An $O(\log N)$ Monte Carlo method for periodic Coulomb systems Xuanzhao Gao et.al. 2601.09288 null
2026-01-14 Enhancing Spatial Reasoning in Large Language Models for Metal-Organic Frameworks Structure Prediction Mianzhi Pan et.al. 2601.09285 null
2026-01-13 Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge Yao Tang et.al. 2601.08808 null
2026-01-13 Identifying Latent Intentions via Inverse Reinforcement Learning in Repeated Linear Public Good Games Carina I. Hausladen et.al. 2601.08803 null
2026-01-13 Enhancing classical simulation with noisy quantum devices Ruiqi Zhang et.al. 2601.08772 null
2026-01-13 Rewarding the Rare: Uniqueness-Aware RL for Creative Problem Solving in LLMs Zhiyuan Hu et.al. 2601.08763 null
2026-01-13 TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback Prithwish Jana et.al. 2601.08734 null
2026-01-13 Learning from Demonstrations via Capability-Aware Goal Sampling Yuanlin Duan et.al. 2601.08731 null
2026-01-13 Model-Agnostic Solutions for Deep Reinforcement Learning in Non-Ergodic Contexts Bert Verbruggen et.al. 2601.08726 null
2026-01-13 Kernel Learning for Regression via Quantum Annealing Based Spectral Sampling Yasushi Hasegawa et.al. 2601.08724 null
2026-01-13 QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models Zhaolu Kang et.al. 2601.08689 null
2026-01-13 PersonaDual: Balancing Personalization and Objectivity via Adaptive Reasoning Xiaoyou Liu et.al. 2601.08679 null
2026-01-13 VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory Shaoan Wang et.al. 2601.08665 null
2026-01-13 From Classical to Quantum Reinforcement Learning and Its Applications in Quantum Control: A Beginner’s Tutorial Abhijit Sen et.al. 2601.08662 null
2026-01-13 Prism: Towards Lowering User Cognitive Load in LLMs via Complex Intent Understanding Zenghua Liao et.al. 2601.08653 null
2026-01-13 Provably Safe Reinforcement Learning using Entropy Regularizer Abhijit Mazumdar et.al. 2601.08646 null
2026-01-13 Systemic Risk Surveillance Timo Dimitriadis et.al. 2601.08598 null
2026-01-13 Sparsifying transform priors in Gaussian graphical models Marcus Gehrmann et.al. 2601.08596 null
2026-01-13 Linear Canonical-Ensemble Quantum Monte Carlo: From Dilute Fermi Gas to Flat-Band Ferromagnetism Tu Hong et.al. 2601.08552 null
2026-01-13 Your Group-Relative Advantage Is Biased Fengkai Yang et.al. 2601.08521 null
2026-01-13 A New Duality-Free Framework for Convex Optimisation with Superlinear Convergence and Effective Warm-Starting Michael Cummins et.al. 2601.08494 null
2026-01-13 AUV Trajectory Learning for Underwater Acoustic Energy Transfer and Age Minimization Mohamed Afouene Melki et.al. 2601.08491 null
2026-01-12 Computing quantum magic of state vectors Piotr Sierant et.al. 2601.07824 null
2026-01-12 Video Generation Models in Robotics – Applications, Research Challenges, Future Directions Zhiting Mei et.al. 2601.07823 null
2026-01-12 Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation Huanyu Li et.al. 2601.07821 null
2026-01-12 Data-driven control of hydraulic impact hammers under strict operational and control constraints Francisco Leiva et.al. 2601.07813 null
2026-01-12 Beyond Single-Shot: Multi-step Tool Retrieval via Query Planning Wei Fang et.al. 2601.07782 null
2026-01-12 Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding Yanxiang Huang et.al. 2601.07761 null
2026-01-12 Hiking in the Wild: A Scalable Perceptive Parkour Framework for Humanoids Shaoting Zhu et.al. 2601.07718 null
2026-01-12 Smooth Operator: Smooth Verifiable Reward Activates Spatial Reasoning Ability of Vision-Language Model Siwen Jiao et.al. 2601.07695 null
2026-01-12 Inversion of Sea Ice Spectral Albedo to Estimate Under-Ice Transmittance Christophe Perron et.al. 2601.07672 null
2026-01-12 To report or not to report: Optimal claim reporting in a bonus-malus system Lea Enzi et.al. 2601.07655 null
2026-01-12 Reinforcement Learning for Micro-Level Claims Reserving Benjamin Avanzi et.al. 2601.07637 null
2026-01-12 Impact of Nuclear Reaction Rates on Calcium Production in Population III Stars: A Global Analysis Qing Wang et.al. 2601.07623 null
2026-01-12 Clipped Affine Policy: Low-Complexity Near-Optimal Online Power Control for Energy Harvesting Communications over Fading Channels Hao Wu et.al. 2601.07622 null
2026-01-12 Quasi-optimal quantum Markov chain spectral gap estimation Adam Connolly et.al. 2601.07601 null
2026-01-12 GRPO with State Mutations: Improving LLM-Based Hardware Test Plan Generation Dimple Vijay Kochar et.al. 2601.07593 null
2026-01-12 Large Language Models for Physics Instrument Design Sara Zoccheddu et.al. 2601.07580 null
2026-01-12 Stagewise Reinforcement Learning and the Geometry of the Regret Landscape Chris Elliott et.al. 2601.07524 null
2026-01-12 Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions Yongqi Li et.al. 2601.07516 null
2026-01-12 Graph Inference Towards ICD Coding Xiaoxiao Deng et.al. 2601.07496 null
2026-01-12 Online Markov Decision Processes with Terminal Law Constraints Bianca Marin Moreno et.al. 2601.07492 null
2026-01-09 Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards Jiajie Zhang et.al. 2601.06021 null
2026-01-09 Constraining Hamiltonians from chiral effective field theory with neutron-star data Cassandra L. Armstrong et.al. 2601.05999 null
2026-01-09 A Framework for Optimizing Human-Machine Interaction in Classification Systems Goran Muric et.al. 2601.05974 null
2026-01-09 Universal Dilation of Linear Itô SDEs: Quantum Trajectories and Lindblad Simulation of Second Moments Hsuan-Cheng Wu et.al. 2601.05928 null
2026-01-09 First-principles study of hydrogen diffusion in polycrystalline Nickel Bhanuj Jain et.al. 2601.05917 null
2026-01-09 TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents Dawei Wang et.al. 2601.05899 null
2026-01-09 StackPlanner: A Centralized Hierarchical Multi-Agent System with Task-Experience Memory Management Ruizhe Zhang et.al. 2601.05890 null
2026-01-09 Estimating optimal interpretable individualized treatment regimes from a classification perspective using adaptive LASSO Yunshu Zhang et.al. 2601.05875 null
2026-01-09 IIB-LPO: Latent Policy Optimization via Iterative Information Bottleneck Huilin Deng et.al. 2601.05870 null
2026-01-09 Sequential Bayesian Optimal Experimental Design in Infinite Dimensions via Policy Gradient Reinforcement Learning Kaichen Shen et.al. 2601.05868 null
2026-01-09 Neural Methods for Multiple Systems Estimation Models Joseph Marsh et.al. 2601.05859 null
2026-01-09 Intelligent Singularity Avoidance in UR10 Robotic Arm Path Planning Using Hybrid Fuzzy Logic and Reinforcement Learning Sheng-Kai Chen et.al. 2601.05836 null
2026-01-09 EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis Xiaoshuai Song et.al. 2601.05808 null
2026-01-09 From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy Assimilation Zezhou Wang et.al. 2601.05787 null
2026-01-09 Inclusion of Inter-crystal Scattering in PET: Analytical Models and Dedicated Reconstruction Jorge Roser et.al. 2601.05717 null
2026-01-09 SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More Muye Huang et.al. 2601.05688 null
2026-01-09 CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action Space Bingyi Liu et.al. 2601.05675 null
2026-01-09 The Hadronization Impact on $J/ψ$ Energy Correlators: A Pythia8 Study from Partonic to Hadronic Observables Jin-peng Zhang et.al. 2601.05658 null
2026-01-09 EvoQRE: Modeling Bounded Rationality in Safety-Critical Traffic Simulation via Evolutionary Quantal Response Equilibrium Phu-Hoa Pham et.al. 2601.05653 null
2026-01-09 GIFT: Games as Informal Training for Generalizable LLMs Nuoyan Lyu et.al. 2601.05633 null
2026-01-08 RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes Yuan-Kang Lee et.al. 2601.05249 null
2026-01-08 GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Shih-Yang Liu et.al. 2601.05242 null
2026-01-08 On the Value Function of Convex Bolza Problems Governed by Stochastic Difference Equations Sebastián Álvarez et.al. 2601.05207 null
2026-01-08 EARL: Energy-Aware Optimization of Liquid State Machines for Pervasive AI Zain Iqbal et.al. 2601.05205 null
2026-01-08 SimuAgent: An LLM-Based Simulink Modeling Assistant Enhanced with Reinforcement Learning Yanchang Liang et.al. 2601.05187 null
2026-01-08 Inside Out: Evolving User-Centric Core Memory Trees for Long-Term Personalized Dialogue Systems Jihao Zhao et.al. 2601.05171 null
2026-01-08 Safe Continual Reinforcement Learning Methods for Nonstationary Environments. Towards a Survey of the State of the Art Timofey Tomashevskiy et.al. 2601.05152 null
2026-01-08 Revealing the Truth: Calculating True Values in Causal Inference Simulation Studies via Gaussian Quadrature Alex Ocampo et.al. 2601.05128 null
2026-01-08 Unitary fault-tolerant encoding of Pauli states in surface codes Luis Colmenarez et.al. 2601.05113 null
2026-01-08 Token-Level LLM Collaboration via FusionRoute Nuoya Xiong et.al. 2601.05106 null
2026-01-08 Online Bayesian Learning of Agent Behavior in Differential Games Francesco Bianchin et.al. 2601.05087 null
2026-01-08 Reinforced Efficient Reasoning via Semantically Diverse Exploration Ziqi Zhao et.al. 2601.05053 null
2026-01-08 Hán Dān Xué Bù (Mimicry) or Qīng Chū Yú Lán (Mastery)? A Cognitive Perspective on Reasoning Distillation in Large Language Models Yueqing Hu et.al. 2601.05019 null
2026-01-08 From Idea to Co-Creation: A Planner-Actor-Critic Framework for Agent Augmented 3D Modeling Jin Gao et.al. 2601.05016 null
2026-01-08 On the Hidden Objective Biases of Group-based Reinforcement Learning Aleksandar Fontana et.al. 2601.05002 null
2026-01-08 AlgBench: To What Extent Do Large Reasoning Models Understand Algorithms? Henan Sun et.al. 2601.04996 null
2026-01-08 A DQN-based model for intelligent network selection in heterogeneous wireless systems Fayssal Bendaoud et.al. 2601.04978 null
2026-01-08 ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning Minda Hu et.al. 2601.04973 null
2026-01-08 Text as a Universal Interface for Transferable Personalization Yuting Liu et.al. 2601.04963 null
2026-01-08 Safe Reinforcement Learning Beyond Baseline Control: A Hierarchical Framework for Space Triangle Tethered Formation System Xinyi Tao et.al. 2601.04957 null
2026-01-07 A Comprehensive Computational Framework for Materials Design, Ab Initio Modeling, and Molecular Docking Md Rakibul Karim Akanda et.al. 2601.04186 null
2026-01-07 Hierarchical GNN-Based Multi-Agent Learning for Dynamic Queue-Jump Lane and Emergency Vehicle Corridor Formation Haoran Su et.al. 2601.04177 null
2026-01-07 Agentic Rubrics as Contextual Verifiers for SWE Agents Mohit Raghavendra et.al. 2601.04171 null
2026-01-07 Diffusion-DRF: Differentiable Reward Flow for Video Diffusion Fine-Tuning Yifan Wang et.al. 2601.04153 null
2026-01-07 On the Distributed Estimation for Scalar-on-Function Regression Models Peilun He et.al. 2601.04138 null
2026-01-07 InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training Ziyun Zhang et.al. 2601.04126 null
2026-01-07 GeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement Learning Wenshuai Li et.al. 2601.04118 null
2026-01-07 Universality in driven systems with a multiply-degenerate umbilic point Johannes Schmidt et.al. 2601.04116 null
2026-01-07 Cells on Autopilot: Adaptive Cell (Re)Selection via Reinforcement Learning Marvin Illian et.al. 2601.04083 null
2026-01-07 Stable Language Guidance for Vision-Language-Action Models Zhihao Zhan et.al. 2601.04052 null
2026-01-07 Quantum computing for multidimensional option pricing: End-to-end pipeline Julien Hok et.al. 2601.04049 null
2026-01-07 Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model Yuan Wang et.al. 2601.04033 null
2026-01-07 An SU(2n)-valued nonlinear Fourier transform Michel Alexis et.al. 2601.03987 null
2026-01-07 On-Device Deep Reinforcement Learning for Decentralized Task Offloading Performance trade-offs in the training process Gorka Nieto et.al. 2601.03976 null
2026-01-07 Anti-Length Shift: Dynamic Outlier Truncation for Training Efficient Reasoning Models Wei Wu et.al. 2601.03969 null
2026-01-07 CoINS: Counterfactual Interactive Navigation via Skill-Aware VLM Kangjie Zhou et.al. 2601.03956 null
2026-01-07 Quantum Monte Carlo Simulations for predicting electron-positron pair production via the linear Breit-Wheeler process Lucas I. Iñigo Gamiz et.al. 2601.03953 null
2026-01-07 Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification Rui Sun et.al. 2601.03948 null
2026-01-07 Adaptive-Boundary-Clipping GRPO: Ensuring Bounded Ratios for Stable and Generalizable Training Chi Liu et.al. 2601.03895 null
2026-01-07 IndexTTS 2.5 Technical Report Yunpei Li et.al. 2601.03888 null
2026-01-06 STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning Juntong Ni et.al. 2601.03248 null
2026-01-06 Critic-Guided Reinforcement Unlearning in Text-to-Image Diffusion Mykola Vysotskyi et.al. 2601.03213 null
2026-01-06 UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward Yile Liu et.al. 2601.03205 null
2026-01-06 MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory Shengtao Zhang et.al. 2601.03192 null
2026-01-06 Breaking the Dimensional Barrier: Dynamic Portfolio Choice with Parameter Uncertainty via Pontryagin Projection Jeonggyu Huh et.al. 2601.03175 null
2026-01-06 WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning Yu Xinmiao et.al. 2601.03164 null
2026-01-06 Unified Thinker: A General Reasoning Modular Core for Image Generation Sashuai Zhou et.al. 2601.03127 null
2026-01-06 A Bayesian Statistical Study of Bianchi Type-I Universe in $f(R,T^ψ)$ Modified Gravity Mohit Thakre et.al. 2601.03116 null
2026-01-06 One Sample to Rule Them All: Extreme Data Efficiency in RL Scaling Yiyuan Li et.al. 2601.03111 null
2026-01-06 Post-Decision State-Based Online Learning for Delay-Energy-Aware Flow Allocation in Wireless Systems Mahesh Ganesh Bhat et.al. 2601.03108 null
2026-01-06 Epicyclic motion and accretion disk around a charged black hole in Einstein-ModMax theory with a quintessence field Hamza Rehman et.al. 2601.03088 null
2026-01-06 IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation Yankai Jiang et.al. 2601.03054 null
2026-01-06 SOP: A Scalable Online Post-Training System for Vision-Language-Action Models Mingjie Pan et.al. 2601.03044 null
2026-01-06 Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning Sindhuja Chaduvula et.al. 2601.03027 null
2026-01-06 Dementia-R1: Reinforced Pretraining and Reasoning from Unstructured Clinical Notes for Real-World Dementia Prognosis Choonghan Kim et.al. 2601.03018 null
2026-01-06 Adaptive Control of Unknown Linear Switched Systems via Policy Gradient Methods Felix Laurent et.al. 2601.03016 null
2026-01-06 In-Context Reinforcement Learning through Bayesian Fusion of Context and Value Prior Anaïs Berkes et.al. 2601.03015 null
2026-01-06 P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Kwangwook Seo et.al. 2601.02986 null
2026-01-06 Interpretable All-Type Audio Deepfake Detection with Audio LLMs via Frequency-Time Reinforcement Learning Yuankun Xie et.al. 2601.02983 null
2026-01-06 Correct, Concise and Complete: Multi-stage Training For Adaptive Reasoning Nathanaël Carraz Rakotonirina et.al. 2601.02972 null
2026-01-05 Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes Jing Tan et.al. 2601.02356 null
2026-01-05 The Sequential Monte Carlo goes NUTS: Boosting Gravitational-Wave Inference Gabriele Demasi et.al. 2601.02336 null
2026-01-05 A comprehensive search for high-velocity X-ray sources: New compact object binary candidates in the Gaia era Yue Zhao et.al. 2601.02287 null
2026-01-05 VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive Generation Shikun Sun et.al. 2601.02256 null
2026-01-05 Enabling Deep Reinforcement Learning Research for Energy Saving in Open RAN Matteo Bordin et.al. 2601.02240 null
2026-01-05 Extended real number arithmetics via Dedekind cuts Andreas H Hamel et.al. 2601.02229 null
2026-01-05 NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation Huichao Zhang et.al. 2601.02204 null
2026-01-05 CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents Keyu Wang et.al. 2601.02201 null
2026-01-05 ACDZero: Graph-Embedding-Based Tree Search for Mastering Automated Cyber Defense Yu Li et.al. 2601.02196 null
2026-01-05 Accurate Helium-Benzene Potential: from CCSD(T) to Gaussian Process Regression Shahzad Akram et.al. 2601.02166 null
2026-01-05 Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting Muxi Diao et.al. 2601.02151 null
2026-01-05 Simulation of Radiation Chemistry by a One-Shot Hybrid Continuum / Monte Carlo Method Charlie Fynn Perkins et.al. 2601.02132 null
2026-01-05 MDAgent2: Large Language Model for Code Generation and Knowledge Q&A in Molecular Dynamics Zhuofan Shi et.al. 2601.02075 null
2026-01-05 Reinforcement Learning Based Computationally Efficient Conditional Choice Simulation Estimation of Dynamic Discrete Choice Models Ahmed Khwaja et.al. 2601.02069 null
2026-01-05 Higher-Order Action Regularization in Deep Reinforcement Learning: From Continuous Control to Building Energy Management Faizan Ahmed et.al. 2601.02061 null
2026-01-05 GDRO: Group-level Reward Post-training Suitable for Diffusion Models Yiyang Wang et.al. 2601.02036 null
2026-01-05 AgentVNE: LLM-Augmented Graph Reinforcement Learning for Affinity-Aware Multi-Agent Placement in Edge Agentic AI Runze Zheng et.al. 2601.02021 null
2026-01-05 Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation Weijian Ma et.al. 2601.01984 null
2026-01-05 Distorted Distributional Policy Evaluation for Offline Reinforcement Learning Ryo Iwaki et.al. 2601.01917 null
2026-01-05 Evaluating Feature Dependent Noise in Preference-based Reinforcement Learning Yuxuan Li et.al. 2601.01904 null
2026-01-02 Exponentially Accelerated Sampling of Pauli Strings for Nonstabilizerness Zhenyu Xiao et.al. 2601.00761 null
2026-01-02 Training-Free Certified Bounds for Quantum Regression: A Scalable Framework Demerson N. Gonçalves et.al. 2601.00745 null
2026-01-02 Materials Informatics: Emergence To Autonomous Discovery In The Age Of AI Turab Lookman et.al. 2601.00742 null
2026-01-02 Stochastic Actor-Critic: Mitigating Overestimation via Temporal Aleatoric Uncertainty Uğurcan Özalp et.al. 2601.00737 null
2026-01-02 Precision Autotuning for Linear Solvers via Contextual Bandit-Based RL Erin Carson et.al. 2601.00728 null
2026-01-02 ARISE: Adaptive Reinforcement Integrated with Swarm Exploration Rajiv Chaitanya M et.al. 2601.00693 null
2026-01-02 IRPO: Scaling the Bradley-Terry Model via Reinforcement Learning Haonan Song et.al. 2601.00677 null
2026-01-02 RoboReward: General-Purpose Vision-Language Reward Models for Robotics Tony Lee et.al. 2601.00675 null
2026-01-02 Spectral Shapes of Pair Annihilation Line Emission in Magnetar Giant Flares Tomoki Wada et.al. 2601.00666 null
2026-01-02 Integrating Multi-Armed Bandit, Active Learning, and Distributed Computing for Scalable Optimization Foo Hui-Mean et.al. 2601.00615 null
2026-01-02 Vision-based Goal-Reaching Control for Mobile Robots Using a Hierarchical Learning Framework Mehdi Heydari Shahna et.al. 2601.00610 null
2026-01-02 Traffic-Aware Optimal Taxi Placement Using Graph Neural Network-Based Reinforcement Learning Sonia Khetarpaul et.al. 2601.00607 null
2026-01-02 Coordination-driven magic numbers in protonated argon clusters Saajid Chowdhury et.al. 2601.00591 null
2026-01-02 Elaboration on the kinetic approach of Derbenev and Kondratenko to spin-polarized beams in electron storage rings Klaus Heinemann et.al. 2601.00586 null
2026-01-02 Parametrized Sharing for Multi-Agent Hybrid DRL for Multiple Multi-Functional RISs-Aided Downlink NOMA Networks Chi-Te Kuo et.al. 2601.00538 null
2026-01-02 Fair Policy Learning under Bipartite Network Interference: Learning Fair and Cost-Effective Environmental Policies Raphael C. Kim et.al. 2601.00531 null
2026-01-01 MIMO-AFDM Outperforms MIMO-OFDM in the Face of Hardware Impairments Zeping Sui et.al. 2601.00502 null
2026-01-01 CPPO: Contrastive Perception for Vision Language Policy Optimization Ahmad Rezaei et.al. 2601.00501 null
2026-01-01 Discovering pulsars in compact binaries with a hidden Markov model Joseph O’Leary et.al. 2601.00500 null
2026-01-01 Fisher-Information-Driven Adaptive Acquisition for Photon-Efficient FLIM: A Dual-Implementation Framework for TCSPC and Programmable Time-Gating J. Sumaya-Martinez et.al. 2601.00490 null
2025-12-31 Coordinated Humanoid Manipulation with Choice Policies Haozhi Qi et.al. 2512.25072 null
2025-12-31 Scaling Open-Ended Reasoning to Predict the Future Nikhil Chandak et.al. 2512.25070 null
2025-12-31 Many Minds from One Model: Bayesian Transformers for Population Intelligence Diji Yang et.al. 2512.25063 null
2025-12-31 ResponseRank: Data-Efficient Reward Modeling through Preference Strength Learning Timo Kaufmann et.al. 2512.25023 null
2025-12-31 MSACL: Multi-Step Actor-Critic Learning with Lyapunov Certificates for Exponentially Stabilizing Control Yongwei Zhang et.al. 2512.24955 null
2025-12-31 Multi-particle quantum systems within the Worldline Monte Carlo formalism Ivan Ahumada et.al. 2512.24942 null
2025-12-31 Iterative Deployment Improves Planning Skills in LLMs Augusto B. Corrêa et.al. 2512.24940 null
2025-12-31 A Pontryagin Maximum Principle on the Belief Space for Continuous-Time Optimal Control with Discrete Observations Christian Bayer et.al. 2512.24916 null
2025-12-31 Adaptive Resource Orchestration for Distributed Quantum Computing Systems Kuan-Cheng Chen et.al. 2512.24902 null
2025-12-31 Adaptive Clutter Suppression via Convex Optimization Yifan He et.al. 2512.24889 null
2025-12-31 Symmetric mass generation as a multicritical point with enhanced symmetry Sandip Maiti et.al. 2512.24836 null
2025-12-31 Explaining Why Things Go Where They Go: Interpretable Constructs of Human Organizational Preferences Emmanuel Fashae et.al. 2512.24829 null
2025-12-31 Upscaling from ab initio atomistic simulations to electrode scale: The case of manganese hexacyanoferrate, a cathode material for Na-ion batteries Yuan-Chi Yang et.al. 2512.24816 null
2025-12-31 Nonlinear Noise2Noise for Efficient Monte Carlo Denoiser Training Andrew Tinits et.al. 2512.24794 null
2025-12-31 Throughput Optimization in UAV-Mounted RIS under Jittering and Imperfect CSI via DRL Anas K. Saeed et.al. 2512.24773 null
2025-12-31 Sparse Offline Reinforcement Learning with Corruption Robustness Nam Phuong Tran et.al. 2512.24768 null
2025-12-31 Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow Karthik Dharmarajan et.al. 2512.24766 null
2025-12-31 Quasi-Maximum Likelihood Estimation for a Genuinely Unbalanced Dynamic Network Panel Data Model Zhijian Wang et.al. 2512.24748 null
2025-12-31 Control of Microrobots with Reinforcement Learning under On-Device Compute Constraints Yichen Liu et.al. 2512.24740 null
2025-12-31 Evolving, Not Training: Zero-Shot Reasoning Segmentation via Evolutionary Prompting Kai Ye et.al. 2512.24702 null
2025-12-29 Training AI Co-Scientists Using Rubric Rewards Shashwat Goel et.al. 2512.23707 null
2025-12-29 Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation Huajie Tan et.al. 2512.23703 null
2025-12-29 Bellman Calibration for V-Learning in Offline Reinforcement Learning Lars van der Laan et.al. 2512.23694 null
2025-12-29 Rethinking conditioning in polarimetry: a new framework beyond $\ell^2$ -based metrics Yuxi Cai et.al. 2512.23689 null
2025-12-29 Scrutinizing the KNT model with vacuum stability conditions Tim Huesmann et.al. 2512.23662 null
2025-12-29 Le Cam Distortion: A Decision-Theoretic Framework for Robust Transfer Learning Deniz Akdemir et.al. 2512.23617 null
2025-12-29 Analysis of kinetic-diffusion Monte Carlo simulation and source term estimation scheme in nuclear fusion applications Zhirui Tang et.al. 2512.23580 null
2025-12-29 ProGuard: Towards Proactive Multimodal Safeguard Shaohan Yu et.al. 2512.23573 null
2025-12-29 Considering parallel tempering and comparing post-treatment procedures in Bayesian Profile Regression Models for a survival outcome and correlated exposures Fendler Julie et.al. 2512.23571 null
2025-12-29 ThinkGen: Generalized Thinking for Visual Generation Siyu Jiao et.al. 2512.23568 null
2025-12-29 A NEAT Approach to Evolving Neural-Network-based Optimization of Chiral Photonic Metasurfaces: Application of a Neuro-Evolution Pipeline Davide Filippozzi et.al. 2512.23558 null
2025-12-29 PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis Shengyi Hua et.al. 2512.23545 null
2025-12-29 Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning Zuoyou Jiang et.al. 2512.23515 null
2025-12-29 Hierarchical Decision Mamba Meets Agentic AI: A Novel Approach for RAN Slicing in 6G Md Arafat Habib et.al. 2512.23502 null
2025-12-29 Joint Link Adaptation and Device Scheduling Approach for URLLC Industrial IoT Network: A DRL-based Method with Bayesian Optimization Wei Gao et.al. 2512.23493 null
2025-12-29 Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation Toqeer Ali Syed et.al. 2512.23480 null
2025-12-29 Simulation of tau decays, ambiguities and anomalous couplings Zbigniew Was et.al. 2512.23475 null
2025-12-29 HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation Yuxin Wen et.al. 2512.23464 null
2025-12-29 Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance Zhuo Li et.al. 2512.23461 null
2025-12-29 Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following Kongcheng Zhang et.al. 2512.23457 null
2025-12-26 Charge-Informed Quantum Error Correction Vlad Temkin et.al. 2512.22119 null
2025-12-26 Index-Tracking Portfolio Construction and Rebalancing under Bayesian Sparse Modelling and Uncertainty Quantification Dimitrios Roxanas et.al. 2512.22109 null
2025-12-26 Hybrid Deep Reinforcement Learning for Joint Resource Allocation in Multi-Active RIS-Aided Uplink Communications Mohamed Shalma et.al. 2512.22107 null
2025-12-26 Exact inference via quasi-conjugacy in two-parameter Poisson-Dirichlet hidden Markov models Marco Dalla Pria et.al. 2512.22098 null
2025-12-26 Heterogeneous fragmentation of empty sites promotes cooperation in phenotypically diverse populations with tag-mediated interactions Hui Zhang et.al. 2512.22077 null
2025-12-26 Rotationally invariant dynamical lattice regulators for Euclidean quantum field theories Tsogtgerel Gantumur et.al. 2512.22072 null
2025-12-26 MAI-UI Technical Report: Real-World Centric Foundation GUI Agents Hanzhang Zhou et.al. 2512.22047 null
2025-12-26 Meta-Learning-Based Handover Management in NextG O-RAN Michail Kalntis et.al. 2512.22022 null
2025-12-26 Spacetime Spins: Statistical mechanics for error correction with stabilizer circuits Cory T. Aitchison et.al. 2512.21991 null
2025-12-26 Latency-Optimal Cache-aided Multicast Streaming via Forward-Backward Reinforcement Learning Mohsen Amidzadeh et.al. 2512.21954 null
2025-12-26 Multi-reference Trial State for Lattice Quantum Monte Carlo Simulations Teng Wang et.al. 2512.21942 null
2025-12-26 SWE-RM: Execution-free Feedback For Software Engineering Agents KaShun Shum et.al. 2512.21919 null
2025-12-26 A Comedy of Estimators: On KL Regularization in RL Training of LLMs Vedant Shah et.al. 2512.21852 null
2025-12-26 A Cohomological Framework for Topological Phases from Momentum-Space Crystallographic Groups T. R. Liu et.al. 2512.21844 null
2025-12-26 Contextual Biasing for LLM-Based ASR with Hotword Retrieval and Reinforcement Learning YuXiang Kong et.al. 2512.21828 null
2025-12-26 Q-A3C2: Quantum Reinforcement Learning with Time-Series Dynamic Clustering for Adaptive ETF Stock Selection Yen-Ku Liu et.al. 2512.21819 null
2025-12-25 Direct Deep Neural-network Extraction of Generalized Parton Distributions Dima Watkins et.al. 2512.21761 null
2025-12-25 On Critical Temperature and Finite Size Scaling of Continuous Spin $2d$ Ising Model Swapna Mahapatra et.al. 2512.21748 null
2025-12-25 Concentration-Dependent Tungsten Effects on Short-Range Order and Deformation Behavior in Ni-W alloys Shaozun Liu et.al. 2512.21745 null
2025-12-25 Multiconnectivity for SAGIN: Current Trends, Challenges, AI-driven Solutions, and Opportunities Abd Ullah Khan et.al. 2512.21717 null
2025-12-24 Autonomous Uncertainty Quantification for Computational Point-of-care Sensors Artem Goncharov et.al. 2512.21335 null
2025-12-24 Minijets and Broken Stationarity in a Blazar : Novel Insights into the Origin of $γ$ -ray Variability in CTA 102 Agniva Roychowdhury et.al. 2512.21240 null
2025-12-24 RoboCade: Gamifying Robot Data Collection Suvir Mirchandani et.al. 2512.21235 null
2025-12-24 MiST: Understanding the Role of Mid-Stage Scientific Training in Developing Chemical Reasoning Models Andres M Bran et.al. 2512.21231 null
2025-12-24 Production of charmed particles in proton-proton and light nucleus-nucleus interactions in Geant4 FTF model A. Galoyan et.al. 2512.21184 null
2025-12-24 Equivariant Multiscale Learned Invertible Reconstruction for Cone Beam CT: From Simulated to Real Data Nikita Moriakov et.al. 2512.21180 null
2025-12-24 Equilibrium investment under dynamic preference uncertainty Luca De Gennaro Aquino et.al. 2512.21149 null
2025-12-24 Thermodynamic sampling of materials using neutral-atom quantum computers Bruno Camino et.al. 2512.21142 null
2025-12-24 Global End-Effector Pose Control of an Underactuated Aerial Manipulator via Reinforcement Learning Shlok Deshmukh et.al. 2512.21085 null
2025-12-24 Dyna-Style Reinforcement Learning Modeling and Control of Non-linear Dynamics Karim Abdelsalam et.al. 2512.21081 null
2025-12-24 LSTM-Based Modeling and Reinforcement Learning Control of a Magnetically Actuated Catheter Arya Rashidinejad Meibodi et.al. 2512.21063 null
2025-12-24 Policy-Conditioned Policies for Multi-Agent Task Solving Yue Lin et.al. 2512.21024 null
2025-12-24 LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics Jiashuo Liu et.al. 2512.21010 null
2025-12-24 LLM-Empowered Agentic AI for QoE-Aware Network Slicing Management in Industrial IoT Xudong Wang et.al. 2512.20997 null
2025-12-24 IMPACTX: An X-ray Spectral Model for Polar Dust and Clumpy Torus Kanta Fujiwara et.al. 2512.20993 null
2025-12-24 Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions Jingyang You et.al. 2512.20974 null
2025-12-24 ReACT-Drug: Reaction-Template Guided Reinforcement Learning for de novo Drug Design R Yadunandan et.al. 2512.20958 null
2025-12-24 One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents Zhaoxi Zhang et.al. 2512.20957 null
2025-12-24 Guardrailed Elasticity Pricing: A Churn-Aware Forecasting Playbook for Subscription Strategy Deepit Sapru et.al. 2512.20932 null
2025-12-24 Model-free stochastic linear quadratic control for discrete-time systems with multiplicative and additive noises via semidefinite programming Jing Guo et.al. 2512.20911 null
2025-12-23 LongVideoAgent: Multi-Agent Reasoning with Long Videos Runtao Liu et.al. 2512.20618 null
2025-12-23 Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning Seijin Kobayashi et.al. 2512.20605 null
2025-12-23 Leveraging High-Fidelity Digital Models and Reinforcement Learning for Mission Engineering: A Case Study of Aerial Firefighting Under Perfect Information İbrahim Oğuz Çetinkaya et.al. 2512.20589 null
2025-12-23 Certified Lower Bounds and Efficient Estimation of Minimum Accuracy in Quantum Kernel Methods Demerson N. Gonçalves et.al. 2512.20588 null
2025-12-23 Performative Policy Gradient: Optimality in Performative Reinforcement Learning Debabrota Basu et.al. 2512.20576 null
2025-12-23 LEAD: Minimizing Learner-Expert Asymmetry in End-to-End Driving Long Nguyen et.al. 2512.20563 null
2025-12-23 Comparative study of plasmons in half-filled graphene via Quantum Monte Carlo and Random Phase Approximation Maksim Ulybyshev et.al. 2512.20559 null
2025-12-23 Recurrent Off-Policy Deep Reinforcement Learning Doesn’t Have to be Slow Tyler Clark et.al. 2512.20513 null
2025-12-23 Quantum vs thermal fluctuations in phase transitions of two-dimensional superconductors Andrea Ponticelli et.al. 2512.20476 null
2025-12-23 Characterization of the BIFROST spectrometer through virtual experiments Kristine M. L. Krighaar et.al. 2512.20463 null
2025-12-23 Topic-informed dynamic mixture model for occupational heterogeneity in health risk behaviors Lorenzo Schiavon et.al. 2512.20408 null
2025-12-23 Resilient Packet Forwarding: A Reinforcement Learning Approach to Routing in Gaussian Interconnected Networks with Clustered Faults Mohammad Walid Charrwi et.al. 2512.20394 null
2025-12-23 Identifying Appropriately-Sized Services with Deep Reinforcement Learning Syeda Tasnim Fabiha et.al. 2512.20381 null
2025-12-23 RIS-Empowered OTFS Modulation With Faster-than-Nyquist Signaling in High-Mobility Wireless Communications Chaorong Zhang et.al. 2512.20332 null
2025-12-23 TableGPT-R1: Advancing Tabular Reasoning Through Reinforcement Learning Saisai Yang et.al. 2512.20312 null
2025-12-23 Graph-Symbolic Policy Enforcement and Control (G-SPEC): A Neuro-Symbolic Framework for Safe Agentic AI in 5G Autonomous Networks Divya Vijay et.al. 2512.20275 null
2025-12-23 Koopman for stochastic dynamics: error bounds for kernel extended dynamic mode decomposition Maximiliano Hertel et.al. 2512.20247 null
2025-12-23 Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning Kausthubh Manda et.al. 2512.20220 null
2025-12-23 Joint Design of Embedded Index Coding and Beamforming for MIMO-based Distributed Computing via Multi-Agent Reinforcement Learning Heekang Song et.al. 2512.20201 null
2025-12-23 Decay of $f(R)$ quintessence into dark matter: mitigating the Hubble tension? Giovanni Montani et.al. 2512.20193 null
2025-12-22 Scalably Enhancing the Clinical Validity of a Task Benchmark with Physician Oversight Junze Ye et.al. 2512.19691 null
2025-12-22 Partition Function Estimation Using Analog Quantum Processors Thinh Le et.al. 2512.19685 null
2025-12-22 VA- $π$ : Variational Policy Alignment for Pixel-Aware Autoregressive Generation Xinyao Liao et.al. 2512.19680 null
2025-12-22 Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies Yuqiao Tan et.al. 2512.19673 null
2025-12-22 CORE: Compensable Reward as a Catalyst for Improving Offline RL in Wireless Networks Lipeng Zu et.al. 2512.19671 null
2025-12-22 Generative diffusion models for agricultural AI: plant image generation, indoor-to-outdoor translation, and expert preference alignment Da Tan et.al. 2512.19632 null
2025-12-22 Learning Generalizable Hand-Object Tracking from Synthetic Demonstrations Yinhuai Wang et.al. 2512.19583 null
2025-12-22 LeLaR: The First In-Orbit Demonstration of an AI-Based Satellite Attitude Controller Kirill Djebko et.al. 2512.19576 null
2025-12-22 Variational Autoregressive Networks Applied to $φ^4$ Field Theory Systems Moxian Qian et.al. 2512.19575 null
2025-12-22 CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Yongxin Wang et.al. 2512.19554 null
2025-12-22 LacaDM: A Latent Causal Diffusion Model for Multiobjective Reinforcement Learning Xueming Yan et.al. 2512.19516 null
2025-12-22 A Gauss-Newton-Induced Structure-Exploiting Algorithm for Differentiable Optimal Control Yuankun Chen et.al. 2512.19447 null
2025-12-22 CodeSimpleQA: Scaling Factuality in Code Large Language Models Jian Yang et.al. 2512.19424 null
2025-12-22 EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration Runze Li et.al. 2512.19396 null
2025-12-22 Superconductivity in Electron Liquids: Precision Many-Body Treatment of Coulomb Interaction Xiansheng Cai et.al. 2512.19382 null
2025-12-22 Learning General Policies with Policy Gradient Methods Simon Ståhlberg et.al. 2512.19366 null
2025-12-22 From Points to Coalitions: Hierarchical Contrastive Shapley Values for Prioritizing Data Samples Canran Xiao et.al. 2512.19363 null
2025-12-22 Interpretable Hybrid Deep Q-Learning Framework for IoT-Based Food Spoilage Prediction with Synthetic Data Generation and Hardware Validation Isshaan Singh et.al. 2512.19361 null
2025-12-22 On the inclusion of the pion form factor in $e^+e^- \to π^+π^-$ beyond leading order Francesco P. Ucci et.al. 2512.19359 null
2025-12-22 First-Order Representation Languages for Goal-Conditioned RL Simon Ståhlberg et.al. 2512.19355 null
2025-12-19 Plane Strong Connectivity Augmentation Stéphane Bessy et.al. 2512.17904 null
2025-12-19 Feasibility to probe the dynamical scotogenic model at the LHC Gustavo Ardila-Tafurth et.al. 2512.17903 null
2025-12-19 Distributionally Robust Imitation Learning: Layered Control Architecture for Certifiable Autonomy Aditya Gahlawat et.al. 2512.17899 null
2025-12-19 Delayed Acceptance Slice Sampling Kevin Bitterlich et.al. 2512.17868 null
2025-12-19 AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning Ran Gong et.al. 2512.17853 null
2025-12-19 Planning as Descent: Goal-Conditioned Latent Trajectory Synthesis in Learned Energy Landscapes Carlos Vélez García et.al. 2512.17846 null
2025-12-19 NeuRehab: A Reinforcement Learning and Spiking Neural Network-Based Rehab Automation Framework Phani Pavan Kambhampati et.al. 2512.17841 null
2025-12-19 Quenching statistics in Si and Ge SPADs using particle Monte Carlo simulation Philippe Dollfus et.al. 2512.17818 null
2025-12-19 Quantum Monte Carlo studies of U(1) lattice gauge models of Kondo breakdown Gaopei Pan et.al. 2512.17801 null
2025-12-19 Recursive state estimation via approximate modal paths Filip Tronarp et.al. 2512.17737 null
2025-12-19 Revisiting the Broken Symmetry Phase of Solid Hydrogen: A Neural Network Variational Monte Carlo Study Shengdu Chai et.al. 2512.17703 null
2025-12-19 Generative Multi-Objective Bayesian Optimization with Scalable Batch Evaluations for Sample-Efficient De Novo Molecular Design Madhav R. Muthyala et.al. 2512.17659 null
2025-12-19 Estimating Spatially Resolved Radiation Fields Using Neural Networks Felix Lehner et.al. 2512.17654 null
2025-12-19 About Time: Model-free Reinforcement Learning with Timed Reward Machines Anirban Majumdar et.al. 2512.17637 null
2025-12-19 Trust-Region Adaptive Policy Optimization Mingyu Su et.al. 2512.17636 null
2025-12-19 SCOPE: Sequential Causal Optimization of Process Interventions Jakob De Moor et.al. 2512.17629 null
2025-12-19 Surrogate-Accelerated Bayesian Inversion for Exoplanet Interior Characterization Tijn De Wringer et.al. 2512.17626 null
2025-12-19 Learning Safe Autonomous Driving Policies Using Predictive Safety Representations Mahesh Keswani et.al. 2512.17586 null
2025-12-19 Kinematics-Aware Diffusion Policy with Consistent 3D Observation and Action Space for Whole-Arm Robotic Manipulation Kangchen Lv et.al. 2512.17568 null
2025-12-19 HydroGym: A Reinforcement Learning Platform for Fluid Dynamics Christian Lagemann et.al. 2512.17534 null
2025-12-18 Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification Qihao Liu et.al. 2512.16921 null
2025-12-18 AdaTooler-V: Adaptive Tool-Use for Images and Videos Chaoyang Wang et.al. 2512.16918 null
2025-12-18 Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning Qihao Liu et.al. 2512.16917 null
2025-12-18 Exploration v.s. Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward Peter Chen et.al. 2512.16912 null
2025-12-18 Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning Andrew Wagenmaker et.al. 2512.16911 null
2025-12-18 MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning Yuanchen Ju et.al. 2512.16909 null
2025-12-18 Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image Yushi Hu et.al. 2512.16899 null
2025-12-18 AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning Tzu-Han Lin et.al. 2512.16883 null
2025-12-18 A survey of the orienteering problem: model evolution, algorithmic advances, and future directions Songhao Shen et.al. 2512.16865 null
2025-12-18 RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing Tianyuan Qu et.al. 2512.16864 null
2025-12-18 ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning Zihan Zhou et.al. 2512.16861 null
2025-12-18 Meta-RL Induces Exploration in Language Agents Yulun Jiang et.al. 2512.16848 null
2025-12-18 Coordinated Anti-Jamming Resilience in Swarm Networks via Multi-Agent Reinforcement Learning Bahman Abolhassani et.al. 2512.16813 null
2025-12-18 Efficient Monte-Carlo sampling of metastable systems using non-local collective variable updates Christoph Schönle et.al. 2512.16812 null
2025-12-18 Strain-Controlled Magnetic Phase Transitions through Anisotropic Exchange Interactions: A Combined DFT and Monte Carlo Study Sudip Mandal et.al. 2512.16765 null
2025-12-18 Pattern recognition in complex systems via vector-field representations of spatio-temporal data Ingrid Amaranta Membrillo Solis et.al. 2512.16763 null
2025-12-18 Correlation between the first-reaction time and the acquired boundary local time Yilin Ye et.al. 2512.16747 null
2025-12-18 Discovering and Learning Probabilistic Models of Black-Box AI Capabilities Daniel Bramblett et.al. 2512.16733 null
2025-12-18 Olaf: Bringing an Animated Character to Life in the Physical World David Müller et.al. 2512.16705 null
2025-12-18 QMCkl: A Kernel Library for Quantum Monte Carlo Applications Emiel Slootman et.al. 2512.16677 null
2025-12-17 Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning Zhenwen Liang et.al. 2512.15687 null
2025-12-17 Revisiting the Phase Diagram of Hard Sphere Dumbbells with Nested Sampling: Known Phases and New Packing Variants Omar-Farouk Adesida et.al. 2512.15665 null
2025-12-17 Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning Jiaqi Xu et.al. 2512.15662 null
2025-12-17 Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction Mathieu Blondel et.al. 2512.15605 null
2025-12-17 Deep Reinforcement Learning for EH-Enabled Cognitive-IoT Under Jamming Attacks Nadia Abdolkhani et.al. 2512.15558 null
2025-12-17 OMCL: Open-vocabulary Monte Carlo Localization Evgenii Kruzhkov et.al. 2512.15557 null
2025-12-17 Autonomous Pressure Control in MuVacAS via Deep Reinforcement Learning and Deep Learning Surrogate Models Guillermo Rodriguez-Llorente et.al. 2512.15521 null
2025-12-17 On regularity of compressions and diagonals of operator functions Vladimir Müller et.al. 2512.15467 null
2025-12-17 Double Horizon Model-Based Policy Optimization Akihiro Kubo et.al. 2512.15439 null
2025-12-17 Statistical repulsion on hyperons in two-color dense QCD Masato Nagatsuka et.al. 2512.15434 null
2025-12-17 FM-EAC: Feature Model-based Enhanced Actor-Critic for Multi-Task Control in Dynamic Environments Quanxi Zhou et.al. 2512.15430 null
2025-12-17 Can AI Generate more Comprehensive Test Scenarios? Review on Automated Driving Systems Test Scenario Generation Methods Ji Zhou et.al. 2512.15422 null
2025-12-17 EUBRL: Epistemic Uncertainty Directed Bayesian Reinforcement Learning Jianfei Ma et.al. 2512.15405 null
2025-12-17 Deep Learning-Driven Quantitative Spectroscopic Photoacoustic Imaging for Segmentation and Oxygen Saturation Estimation Ruibo Shang et.al. 2512.15394 null
2025-12-17 Environmental Policy and Firm Performance in Europe: A Difference-in-Differences Approach with Spillovers Andrea Ciaccio et.al. 2512.15377 null
2025-12-17 Amplitude-amplified coherence detection and estimation Rhea Alexander et.al. 2512.15352 null
2025-12-17 Explicit Solution to a government debt reduction problem: a stochastic control approach Claudia Ceci et.al. 2512.15296 null
2025-12-17 Graph Contextual Reinforcement Learning for Efficient Directed Controller Synthesis Toshihide Ubukata et.al. 2512.15295 null
2025-12-17 Learning-Based Phase Shift Optimization of Liquid Crystal RIS in Dynamic mmWave Networks Le Hao et.al. 2512.15279 null
2025-12-17 Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM Reasoning Yiliu Sun et.al. 2512.15274 null
2025-12-16 TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs Jun Zhang et.al. 2512.14698 null
2025-12-16 CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives Zihan Wang et.al. 2512.14696 null
2025-12-16 Model-Based Reinforcement Learning in Discrete-Action Non-Markovian Reward Decision Processes Alessandro Trapasso et.al. 2512.14617 null
2025-12-16 Estimating Program Participation with Partial Validation Augustine Denteh et.al. 2512.14616 null
2025-12-16 Fair sampling of ground-state configurations using hybrid quantum-classical MCMC algorithms Yuichiro Nakano et.al. 2512.14552 null
2025-12-16 RecGPT-V2 Technical Report Chao Yi et.al. 2512.14503 null
2025-12-16 PushGen: Push Notifications Generation with LLM Shifu Bie et.al. 2512.14490 null
2025-12-16 Hybrid Cognitive IoT with Cooperative Caching and SWIPT-EH: A Hierarchical Reinforcement Learning Framework Nadia Abdolkhani et.al. 2512.14488 null
2025-12-16 Context-Picker: Dynamic context selection using multi-stage reinforcement learning Siyuan Zhu et.al. 2512.14465 null
2025-12-16 Leggett’s bound and superfluidity in strongly interacting bosons Lorenzo Pizzino et.al. 2512.14346 null
2025-12-16 A data-physics hybrid generative model for patient-specific post-stroke motor rehabilitation using wearable sensor data Yanning Dai et.al. 2512.14329 null
2025-12-16 Multi-Agent Medical Decision Consensus Matrix System: An Intelligent Collaborative Framework for Oncology MDT Consultations Xudong Han et.al. 2512.14321 null
2025-12-16 A Threshold-Triggered Deep Q-Network-Based Framework for Self-Healing in Autonomic Software-Defined IIoT-Edge Networks Agrippina Mwangi et.al. 2512.14297 null
2025-12-16 GLM-TTS Technical Report Jiayan Cui et.al. 2512.14291 null
2025-12-16 Understanding and Improving Hyperbolic Deep Reinforcement Learning Timo Klein et.al. 2512.14202 null
2025-12-16 Incentivizing Tool-augmented Thinking with Images for Medical Image Analysis Yankai Jiang et.al. 2512.14157 null
2025-12-16 Vertically resolved minimal-set k-distribution for thermal infrared absorption in the Venus atmosphere Boris Fomin et.al. 2512.14120 null
2025-12-16 A First-Order Logic-Based Alternative to Reward Models in RLHF Chunjin Jian et.al. 2512.14100 null
2025-12-16 RADAR: Accelerating Large Language Model Inference With RL-Based Dynamic Draft Trees Junjie Ma et.al. 2512.14069 null
2025-12-16 Context Representation via Action-Free Transformer encoder-decoder for Meta Reinforcement Learning Amir M. Soufi Enayati et.al. 2512.14057 null
2025-12-15 AgentIAD: Tool-Augmented Single-Agent for Industrial Anomaly Detection Junwen Miao et.al. 2512.13671 null
2025-12-15 A Scientific Reasoning Model for Organic Synthesis Procedure Generation Guoqing Liu et.al. 2512.13668 null
2025-12-15 Advancing Machine Learning Optimization of Chiral Photonic Metasurface: Comparative Study of Neural Network and Genetic Algorithm Approaches Davide Filippozzi et.al. 2512.13656 null
2025-12-15 MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning Haoyu Fu et.al. 2512.13636 null
2025-12-15 SCR2-ST: Combine Single Cell with Spatial Transcriptomics for Efficient Active Sampling via Reinforcement Learning Junchao Zhu et.al. 2512.13635 null
2025-12-15 Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models Boxin Wang et.al. 2512.13607 null
2025-12-15 Image Diffusion Preview with Consistency Solver Fu-Yun Wang et.al. 2512.13592 null
2025-12-15 MMhops-R1: Multimodal Multi-hop Reasoning Tao Zhang et.al. 2512.13573 null
2025-12-15 Memory in the Age of AI Agents Yuyang Hu et.al. 2512.13564 null
2025-12-15 Disability insurance with collective health claims: A mean-field approach Christian Furrer et.al. 2512.13562 null
2025-12-15 Magnetic order and novel quantum criticality in the strongly interacting quasicrystals Cong Zhang et.al. 2512.13546 null
2025-12-15 How Low Can You Go? The Data-Light SE Challenge Kishan Kumar Ganguly et.al. 2512.13524 null
2025-12-15 Reinforcement Learning based 6-DoF Maneuvers for Microgravity Intravehicular Docking: A Simulation Study with Int-Ball2 in ISS-JEM Aman Arora et.al. 2512.13514 null
2025-12-15 MedCEG: Reinforcing Verifiable Medical Reasoning with Critical Evidence Graph Linjie Mu et.al. 2512.13510 null
2025-12-15 Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model Siyan Chen et.al. 2512.13507 null
2025-12-15 Differentiable Evolutionary Reinforcement Learning Sitao Cheng et.al. 2512.13399 null
2025-12-15 QoS-Aware State-Augmented Learnable Framework for 5G NR-U/Wi-Fi Coexistence: Impact of Parameter Selection and Enhanced Collision Resolution Mohammad Reza Fasihi et.al. 2512.13393 null
2025-12-15 Universal Dexterous Functional Grasping via Demonstration-Editing Reinforcement Learning Chuan Mao et.al. 2512.13380 null
2025-12-15 Fast Policy Learning for 6-DOF Position Control of Underwater Vehicles Sümer Tunçay et.al. 2512.13359 null
2025-12-15 Control of a Twin Rotor using Twin Delayed Deep Deterministic Policy Gradient (TD3) Zeyad Gamal et.al. 2512.13356 null
2025-12-12 AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis Junjie Ye et.al. 2512.11797 null
2025-12-12 Agile Flight Emerges from Multi-Agent Competitive Racing Vineet Pasumarti et.al. 2512.11781 null
2025-12-12 SUMFORU: An LLM-Based Review Summarization Framework for Personalized Purchase Decision Support Yuming Feng et.al. 2512.11755 null
2025-12-12 Spatially Varying Gene Regulatory Networks via Bayesian Nonparametric Covariate-Dependent Directed Cyclic Graphical Models Trisha Dawn et.al. 2512.11732 null
2025-12-12 Transfer Learning (Il)liquidity Andrea Conti et.al. 2512.11731 null
2025-12-12 Order statistics for multijet events G. Chachamis et.al. 2512.11697 null
2025-12-12 2 $k_F$ instability and chiral spin density wave at the 1/9 magnetization plateau in the kagome antiferromagnets Tanja Đurić et.al. 2512.11670 null
2025-12-12 Fast and Explicit: Slice-to-Volume Reconstruction via 3D Gaussian Primitives with Analytic Point Spread Function Modeling Maik Dannecker et.al. 2512.11624 null
2025-12-12 UniBYD: A Unified Framework for Learning Robotic Manipulation Across Embodiments Beyond Imitation of Human Demonstrations Tingyu Yuan et.al. 2512.11609 null
2025-12-12 Detecting changes in the mean of spatial random fields on a regular grid Sheila T. Görz et.al. 2512.11599 null
2025-12-12 DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry Zhenyang Cai et.al. 2512.11558 null
2025-12-12 Rethinking Expert Trajectory Utilization in LLM Post-training Bowen Ding et.al. 2512.11470 null
2025-12-12 Three methods, one problem: Classical and AI approaches to no-three-in-line Pranav Ramanathan et.al. 2512.11469 null
2025-12-12 On Łojasiewicz Ideals and Flatness for Zero Sets with Infinite Tangential Geometry Abdelhafed El Khadiri et.al. 2512.11432 null
2025-12-12 Conditional Copula models using loss-based Bayesian Additive Regression Trees Tathagata Basu et.al. 2512.11427 null
2025-12-12 Towards Trustworthy Multi-Turn LLM Agents via Behavioral Guidance Gonca Gürsun et.al. 2512.11421 null
2025-12-12 Intergalactic magnetic field lower limits up to the redshift $z\approx3$ Ievgen Vovk et.al. 2512.11397 null
2025-12-12 Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization Yifan Niu et.al. 2512.11391 null
2025-12-12 A Molecular Gas Dynamics Study of Hypersonic Boundary Layer Second Mack Mode Instabilities Mert Senkardesler et.al. 2512.11390 null
2025-12-12 Computing quantum entanglement with machine learning Andrea Bulgarelli et.al. 2512.11389 null
2025-12-11 Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation Yiwen Tang et.al. 2512.10949 null
2025-12-11 Curriculum-Based Reinforcement Learning for Autonomous UAV Navigation in Unknown Curved Tubular Conduit Zamirddine Mari et.al. 2512.10934 null
2025-12-11 Digital Twin Supervised Reinforcement Learning Framework for Autonomous Underwater Navigation Zamirddine Mari et.al. 2512.10925 null
2025-12-11 Reinforcement Learning in Financial Decision Making: A Systematic Review of Performance, Challenges, and Implementation Strategies Mohammad Rezoanul Hoque et.al. 2512.10913 null
2025-12-11 Iterative Compositional Data Generation for Robot Control Anh-Quan Pham et.al. 2512.10891 null
2025-12-11 Impact of geometry on 1D molecular-kinetics simulations of acoustic-gravity wave propagation into the exosphere Jose A. Perez Chavez et.al. 2512.10887 null
2025-12-11 Bayesian Symbolic Regression via Posterior Sampling Geoffrey F. Bomarito et.al. 2512.10849 null
2025-12-11 Learning Controllable and Diverse Player Behaviors in Multi-Agent Environments Atahan Cilan et.al. 2512.10835 null
2025-12-11 V-OCBF: Learning Safety Filters from Offline Data via Value-Guided Offline Control Barrier Functions Mumuksh Tayal et.al. 2512.10822 null
2025-12-11 Complexity and multi-functional variants of the Quantum-to-Quantum Bernoulli Factories Francesco Hoch et.al. 2512.10810 null
2025-12-11 OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification Zijian Wu et.al. 2512.10756 null
2025-12-11 Learning to Split: A Reinforcement-Learning-Guided Splitting Heuristic for Neural Network Verification Maya Swisa et.al. 2512.10747 null
2025-12-11 Long-horizon Reasoning Agent for Olympiad-Level Mathematical Problem Solving Songyang Gao et.al. 2512.10739 null
2025-12-11 Disperon QED Yizhou Fang et.al. 2512.10709 null
2025-12-11 How to Brake? Ethical Emergency Braking with Deep Reinforcement Learning Jianbo Wang et.al. 2512.10698 null
2025-12-11 Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning Benjamin Gundersen et.al. 2512.10691 null
2025-12-11 A Cryogenic Muon Tagging System Based on Kinetic Inductance Detectors for Superconducting Quantum Processors Ambra Mariani et.al. 2512.10679 null
2025-12-11 AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence Bo Yang et.al. 2512.10624 null
2025-12-11 Dynamics of multidimensional Simple Clock Auctions Jad Zeroual et.al. 2512.10614 null
2025-12-11 Multi-Objective Reward and Preference Optimization: Theory and Algorithms Akhil Agnihotri et.al. 2512.10601 null
2025-12-10 STACHE: Local Black-Box Explanations for Reinforcement Learning Policies Andrew Elashkin et.al. 2512.09909 null
2025-12-10 FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning Khurram Khalil et.al. 2512.09872 null
2025-12-10 Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation Yuyang Li et.al. 2512.09851 null
2025-12-10 ChronusOmni: Improving Time Awareness of Omni Large Language Models Yijing Chen et.al. 2512.09841 null
2025-12-10 RIFT: A Scalable Methodology for LLM Accelerator Fault Assessment using Reinforcement Learning Khurram Khalil et.al. 2512.09829 null
2025-12-10 Prefrontal scaling of reward prediction error readout gates reinforcement-derived adaptive behavior in primates Tian Sang et.al. 2512.09761 null
2025-12-10 MOA: Multi-Objective Alignment for Role-Playing Agents Chonghua Liao et.al. 2512.09756 null
2025-12-10 Dimensional crossover and finite-range effects in a quasi-two-dimensional gas of fermionic dimers Giovanni Midei et.al. 2512.09753 null
2025-12-10 Tachyonic dark energy- Constraints from current observations Ramanpreet Singh et.al. 2512.09744 null
2025-12-10 Gaussian Process Aggregation for Root-Parallel Monte Carlo Tree Search with Continuous Actions Junlin Xiao et.al. 2512.09727 null
2025-12-10 Bayesian Model Selection with an Application to Cosmology Nikoloz Gigiberia et.al. 2512.09724 null
2025-12-10 Flexible Reconfigurable Intelligent Surface-Aided Covert Communications in UAV Networks Chong Huang et.al. 2512.09714 null
2025-12-10 Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning Kaichen He et.al. 2512.09706 null
2025-12-10 Dynamic one-time delivery of critical data by small and sparse UAV swarms: a model problem for MARL scaling studies Mika Persson et.al. 2512.09682 null
2025-12-10 d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models Leyi Pan et.al. 2512.09675 null
2025-12-10 A Monotone–Operator Proof of Existence and Uniqueness for a Simple Stationary Mean Field Game Hikmatullo Ismatov et.al. 2512.09671 null
2025-12-10 Persistent Cycle Representatives and Generalized Landscapes for Codimension 1 Persistent Homology Fabian Lenzen et.al. 2512.09668 null
2025-12-10 SynthPix: A lightspeed PIV images generator Antonio Terpin et.al. 2512.09664 null
2025-12-10 Theoretical Characterization of the Magnetic Properties of Vanadium-doped Ti2C MXenes Carlos Patiño et.al. 2512.09657 null
2025-12-10 Graph-Based Bayesian Optimization for Quantum Circuit Architecture Search with Uncertainty Calibrated Surrogates Prashant Kumar Choudhary et.al. 2512.09586 null
2025-12-09 No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers Damiano Marsili et.al. 2512.08889 null
2025-12-09 IPPO Learns the Game, Not the Team: A Study on Generalization in Heterogeneous Agent Teams Ryan LeRoy et.al. 2512.08877 null
2025-12-09 Toward Quantitative Modeling of Cybersecurity Risks Due to AI Misuse Steve Barrett et.al. 2512.08864 null
2025-12-09 Reinforcement Learning From State and Temporal Differences Lex Weaver et.al. 2512.08855 null
2025-12-09 Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages David Samuel et.al. 2512.08777 null
2025-12-09 Triangular $J_1$-$J_2$ Heisenberg Antiferromagnet in a Magnetic Field Thomas Bader et.al. 2512.08768 null
2025-12-09 Optimal navigation in two-dimensional regular and turbulent flows Vladimir Parfenyev et.al. 2512.08766 null
2025-12-09 Learning and Editing Universal Graph Prompt Tuning via Reinforcement Learning Jinfeng Xu et.al. 2512.08763 null
2025-12-09 Adaptive cut reveals multiscale complexity in networks Louis Boucherie et.al. 2512.08741 null
2025-12-09 Gradient-Informed Monte Carlo Fine-Tuning of Diffusion Models for Low-Thrust Trajectory Design Jannik Graebner et.al. 2512.08705 null
2025-12-09 Direct transfer of optimized controllers to similar systems using dimensionless MPC Josip Kir Hromatko et.al. 2512.08667 null
2025-12-09 Sim2Swim: Zero-Shot Velocity Control for Agile AUV Maneuvering in 3 Minutes Lauritz Rismark Fosso et.al. 2512.08656 null
2025-12-09 CogMCTS: A Novel Cognitive-Guided Monte Carlo Tree Search Framework for Iterative Heuristic Evolution with Large Language Models Hui Wang et.al. 2512.08609 null
2025-12-09 Heuristics for Combinatorial Optimization via Value-based Reinforcement Learning: A Unified Framework and Analysis Orit Davidovich et.al. 2512.08601 null
2025-12-09 Study of a small-scale gamma-ray detection system employing Compton scattering with a monolithic CeBr3 crystal and segmented photodetector array Veronika Asova et.al. 2512.08593 null
2025-12-09 Mind to Hand: Purposeful Robotic Control via Embodied Reasoning Peijun Tang et.al. 2512.08580 null
2025-12-09 Thinking with Images via Self-Calling Agent Wenxi Yang et.al. 2512.08511 null
2025-12-09 Developing Distance-Aware Uncertainty Quantification Methods in Physics-Guided Neural Networks for Reliable Bearing Health Prediction Waleed Razzaq et.al. 2512.08499 null
2025-12-09 Optimal Perturbation Budget Allocation for Data Poisoning in Offline Reinforcement Learning Junnan Qiu et.al. 2512.08485 null
2025-12-09 Using reinforcement learning to probe the role of feedback in skill acquisition Antonio Terpin et.al. 2512.08463 null
2025-12-08 An Adaptive Multi-Layered Honeynet Architecture for Threat Behavior Analysis via Deep Learning Lukas Johannes Möller et.al. 2512.07827 null
2025-12-08 Trapped Fermions Through Kolmogorov-Arnold Wavefunctions Paulo F. Bedaque et.al. 2512.07800 null
2025-12-08 On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models Charlie Zhang et.al. 2512.07783 null
2025-12-08 RL-MTJail: Reinforcement Learning for Automated Black-Box Multi-Turn Jailbreaking of Large Language Models Xiqiao Xiong et.al. 2512.07761 null
2025-12-08 DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving Jialv Zou et.al. 2512.07745 null
2025-12-08 SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery Meng Cao et.al. 2512.07733 null
2025-12-08 Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE Anxiang Zeng et.al. 2512.07710 null
2025-12-08 Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic Tasks Aileen Liao et.al. 2512.07697 null
2025-12-08 Prospects for measuring electroweak production of $Zγγ$ and 2 jets at the LHC Ran Ding et.al. 2512.07689 null
2025-12-08 Observability of eccentricity in a population of merging compact binaries Mukesh Kumar Singh et.al. 2512.07688 null
2025-12-08 The Agent Capability Problem: Predicting Solvability Through Information-Theoretic Bounds Shahar Lutati et.al. 2512.07631 null
2025-12-08 Comparative Analysis and Parametric Tuning of PPO, GRPO, and DAPO for LLM Reasoning Enhancement Yongsheng Lian et.al. 2512.07611 null
2025-12-08 Critical Density-Wave Vestigial Phases of Commensurate Pair Density Wave Chu-Tian Gao et.al. 2512.07591 null
2025-12-08 Understanding Individual Decision-Making in Multi-Agent Reinforcement Learning: A Dynamical Systems Approach James Rudd-Jones et.al. 2512.07588 null
2025-12-08 LongCat-Image Technical Report Meituan LongCat Team et.al. 2512.07584 null
2025-12-08 ReLaX: Reasoning with Latent Exploration for Large Reasoning Models Shimin Zhang et.al. 2512.07558 null
2025-12-08 Model-Based Reinforcement Learning Under Confounding Nishanth Venkatesh et.al. 2512.07528 null
2025-12-08 How Do LLMs Fail In Agentic Scenarios? A Qualitative Analysis of Success and Failure Scenarios of Various LLMs in Agentic Simulations JV Roig et.al. 2512.07497 null
2025-12-08 Enhancing Agentic RL with Progressive Reward Shaping and Value-based Sampling Policy Optimization Zhuoran Zhuang et.al. 2512.07478 null
2025-12-08 Recovery of the optimal control value function in reproducing kernel Hilbert spaces from verification conditions Tobias Ehring et.al. 2512.07477 null
2025-12-05 EditThinker: Unlocking Iterative Reasoning for Any Image Editor Hongyu Li et.al. 2512.05965 null
2025-12-05 Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping Diversity Germán Kruszewski et.al. 2512.05962 null
2025-12-05 MaxShapley: Towards Incentive-compatible Generative Search with Fair Context Attribution Sara Patel et.al. 2512.05958 null
2025-12-05 Correspondence-Oriented Imitation Learning: Flexible Visuomotor Control with 3D Conditioning Yunhao Cao et.al. 2512.05953 null
2025-12-05 Variational Quantum Rainbow Deep Q-Network for Optimizing Resource Allocation Problem Truong Thanh Hung Nguyen et.al. 2512.05946 null
2025-12-05 A poset representation for stable contracts in a two-sided market generated by integer choice functions Alexander V. Karzanov et.al. 2512.05942 null
2025-12-05 A Residual Variance Matching Recursive Least Squares Filter for Real-time UAV Terrain Following Xiaobo Wu et.al. 2512.05918 null
2025-12-05 Functional dual-slope frequency-domain near-infrared spectroscopy data interpreted with two- and three-layer models Jodee Frias et.al. 2512.05877 null
2025-12-05 Invariant Price of Anarchy: a Metric for Welfarist Traffic Control Ilia Shilov et.al. 2512.05843 null
2025-12-05 Toward Efficient and Robust Behavior Models for Multi-Agent Driving Simulation Fabian Konstantinidis et.al. 2512.05812 null
2025-12-05 Twisting the Hagedorn temperature in planar $\mathcal{N}=4$ super Yang-Mills Simon Ekhammar et.al. 2512.05810 null
2025-12-05 Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling Saurav Jha et.al. 2512.05809 null
2025-12-05 Real-time Remote Tracking and Autonomous Planning for Whale Rendezvous using Robots Sushmita Bhattacharya et.al. 2512.05808 null
2025-12-05 Three-loop jet function for boosted top quarks Alberto M. Clavero et.al. 2512.05795 null
2025-12-05 Machine Learning-Informed 3+1 Sterile Neutrino Global Fits using Posterior Density Estimation of Electron Disappearance Data Joshua Villarreal et.al. 2512.05784 null
2025-12-05 Towards agent-based-model informed neural networks Nino Antulov-Fantulin et.al. 2512.05764 null
2025-12-05 USV: Unified Sparsification for Accelerating Video Diffusion Models Xinjian Wu et.al. 2512.05754 null
2025-12-05 A Fast Anti-Jamming Cognitive Radar Deployment Algorithm Based on Reinforcement Learning Wencheng Cai et.al. 2512.05753 null
2025-12-05 Stochastic Reconfiguration with Warm-Started SVD Dexuan Zhou et.al. 2512.05749 null
2025-12-05 Capturing Classic Authorial Style in Long-Form Story Generation with GRPO Fine-Tuning Jinlong Liu et.al. 2512.05747 null
2025-12-05 A High-Order Immersed Boundary Method for Fluid-Structure Interaction Problems Yingjie Xia et.al. 2512.05733 null
2025-12-05 Taylor Approximation Variance Reduction for Approximation Errors in PDE-constrained Bayesian Inverse Problems Ruanui Nicholson et.al. 2512.05723 null
2025-12-05 $α$ -Potential Games for Decentralized Control of Connected and Automated Vehicles Xuan Di et.al. 2512.05712 null
2025-12-05 Bayesian Active Inference for Intelligent UAV Anti-Jamming and Adaptive Trajectory Planning Ali Krayani et.al. 2512.05711 null
2025-12-05 LA-RL: Language Action-guided Reinforcement Learning with Safety Guarantees for Autonomous Highway Driving Yiming Shu et.al. 2512.05686 null
2025-12-05 MedTutor-R1: Socratic Personalized Medical Teaching with Multi-Agent Simulation Zhitao He et.al. 2512.05671 null
2025-12-05 Efficient sequential Bayesian inference for state-space epidemic models using ensemble data assimilation Dhorasso Temfack et.al. 2512.05650 null
2025-12-05 Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning Zhenpeng Su et.al. 2512.05591 null
2025-12-05 RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs Jonathan Geuter et.al. 2512.05542 null
2025-12-04 Value Gradient Guidance for Flow Matching Alignment Zhen Liu et.al. 2512.05116 null
2025-12-04 ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning Shengyuan Ding et.al. 2512.05111 null
2025-12-04 STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models Feng Xu et.al. 2512.05107 null
2025-12-04 Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning Purbesh Mitra et.al. 2512.05105 null
2025-12-04 Structured Document Translation via Format Reinforcement Learning Haiyue Song et.al. 2512.05100 null
2025-12-04 SA-IQA: Redefining Image Quality Assessment for Spatial Aesthetics with Multi-Dimensional Rewards Yuan Gao et.al. 2512.05098 null
2025-12-04 From Generated Human Videos to Physically Plausible Robot Trajectories James Ni et.al. 2512.05094 null
2025-12-04 The Geometry of Intelligence: Deterministic Functional Topology as a Foundation for Real-World Perception Eduardo Di Santi et.al. 2512.05089 null
2025-12-04 Efficient Decoders for Sensing Subspace Code Siva Aditya Gooty et.al. 2512.05028 null
2025-12-04 Model-Free Assessment of Simulator Fidelity via Quantile Curves Garud Iyengar et.al. 2512.05024 null
2025-12-04 Lévy sources in UrQMD in Ar+Sc collisions at SPS energies Barnabas Porfy et.al. 2512.05019 null
2025-12-04 Geophysical intensity problems: the axisymmetric case Ralf Kaiser et.al. 2512.05010 null
2025-12-04 A meta-GGA perspective on the altermagnetism of RuO2 Markus Meinert et.al. 2512.04995 null
2025-12-04 Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction Nex-AGI Team et.al. 2512.04987 null
2025-12-04 Inflationary relics from an Ultra-Slow-Roll plateau Albert Escrivà et.al. 2512.04986 null
2025-12-04 Towards a unified framework for guided diffusion models Yuchen Jiao et.al. 2512.04985 null
2025-12-04 Internal superfluid response and torque evolution in the giant glitch of PSR J1718-3718 Peng Liu et.al. 2512.04972 null
2025-12-04 Hybrid-Diffusion Models: Combining Open-loop Routines with Visuomotor Diffusion Policies Jonne Van Haastregt et.al. 2512.04960 null
2025-12-04 Realizable Abstractions: Near-Optimal Hierarchical Reinforcement Learning Roberto Cipollone et.al. 2512.04958 null
2025-12-04 CARL: Critical Action Focused Reinforcement Learning for Multi-Step Agent Leyang Shen et.al. 2512.04949 null
2025-12-03 PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design Jiazhe Wei et.al. 2512.04082 null
2025-12-03 Semi-Markov Decision Process Framework for Age of Incorrect Information Minimization Ismail Cosandal et.al. 2512.04077 null
2025-12-03 SkillFactory: Self-Distillation For Learning Cognitive Behaviors Zayne Sprague et.al. 2512.04072 null
2025-12-03 SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL Siyi Chen et.al. 2512.04069 null
2025-12-03 Learning Steerable Clarification Policies with Collaborative Self-play Jonathan Berant et.al. 2512.04068 null
2025-12-03 Sign-Resolved Statistics and the Origin of Bias in Quantum Monte Carlo Ryan Larson et.al. 2512.04056 null
2025-12-03 MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking Yizhou Zhao et.al. 2512.04044 null
2025-12-03 Predicting parameters of a model cuprate superconductor using machine learning V. A. Ulitko et.al. 2512.04024 null
2025-12-03 Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics Connall Garrod et.al. 2512.04006 null
2025-12-03 A Strict Comparison Principle for Integro-Differential Hamilton-Jacobi-Bellman Equations on Domains with Boundary Serena Della Corte et.al. 2512.04005 null
2025-12-03 Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation Hang Xu et.al. 2512.03996 null
2025-12-03 Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning Franki Nguimatsia Tiofack et.al. 2512.03973 null
2025-12-03 Training for Identity, Inference for Controllability: A Unified Approach to Tuning-Free Face Personalization Lianyu Pang et.al. 2512.03964 null
2025-12-03 TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Tao Wu et.al. 2512.03963 null
2025-12-03 Primary gravitational waves at high frequencies I: Origin of suppression in the power spectrum Alipriyo Hoory et.al. 2512.03959 null
2025-12-03 Performance and efficiency of a transformer-based quark/gluon jet tagger in the ATLAS experiment ATLAS Collaboration et.al. 2512.03949 null
2025-12-03 Discontinuous Strongly Quasiconvex Functions Nguyen Thi Van Hang et.al. 2512.03934 null
2025-12-03 Integrating High Performance In-Memory Data Streaming and In-Situ Visualization in Hybrid MPI+OpenMP PIC MC Simulations Towards Exascale Jeremy J. Williams et.al. 2512.03914 null
2025-12-03 Hierarchical Vision Language Action Model Using Success and Failure Demonstrations Jeongeun Park et.al. 2512.03913 null
2025-12-03 Autonomous Reinforcement Learning Robot Control with Intel’s Loihi 2 Neuromorphic Hardware Kenneth Stewart et.al. 2512.03911 null
2025-12-02 OneThinker: All-in-one Reasoning Model for Image and Video Kaituo Feng et.al. 2512.03043 null
2025-12-02 Entanglement evolution from entangled multipodal states Konstantinos Chalas et.al. 2512.03032 null
2025-12-02 SMP: Reusable Score-Matching Motion Priors for Physics-Based Character Control Yuxuan Mu et.al. 2512.03028 null
2025-12-02 LORE: A Large Generative Model for Search Relevance Chenji Lu et.al. 2512.03025 null
2025-12-02 New insights into hydrogen-assisted intergranular cracking in nickel S. Quan et.al. 2512.02979 null
2025-12-02 Spinons and Spin-Charge Separation at the Deconfined Quantum Critical Point Sibin Yang et.al. 2512.02962 null
2025-12-02 Exceptional Point Dynamics in Photonic Time Crystals for Enhanced Optical Sensing Saurabh Mani Tripathi et.al. 2512.02945 null
2025-12-02 The Convex Matching Distance in Multiparameter Persistence Patrizio Frosini et.al. 2512.02944 null
2025-12-02 Eisenstein cohomology and congruences for the ratios of Rankin–Selberg $L$ -functions P. Narayanan et.al. 2512.02927 null
2025-12-02 Asymptotics for additive functionals of particle systems via Stein’s method Arturo Jaramillo et.al. 2512.02922 null
2025-12-02 Congruences for the ratios of Rankin–Selberg $L$ -functions P. Narayanan et.al. 2512.02919 null
2025-12-02 MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm Wei Chen et.al. 2512.02895 null
2025-12-02 OptPO: Optimal Rollout Allocation for Test-time Policy Optimization Youkang Wang et.al. 2512.02882 null
2025-12-02 Insights into Extragalactic Background Light constraints with MAGIC archival data R. Grau et.al. 2512.02880 null
2025-12-02 Taming Camera-Controlled Video Generation with Verifiable Geometry Reward Zhaoqing Wang et.al. 2512.02870 null
2025-12-02 Assessing the performance of correlation-based multi-fidelity neural emulators Cristian J. Villatoro et.al. 2512.02868 null
2025-12-02 ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning Yifan Li et.al. 2512.02835 null
2025-12-02 Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach Siyuan Yang et.al. 2512.02834 null
2025-12-02 A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models Kunning Li et.al. 2512.02816 null
2025-12-02 Phase-Adaptive LLM Framework with Multi-Stage Validation for Construction Robot Task Allocation: A Systematic Benchmark Against Traditional Optimization Algorithms Shyam prasad reddy Kaitha et.al. 2512.02810 null
2025-12-01 A Diffusion Model Framework for Maximum Entropy Reinforcement Learning Sebastian Sanokowski et.al. 2512.02019 null
2025-12-01 Learning Dexterous Manipulation Skills from Imperfect Simulations Elvis Hsieh et.al. 2512.02011 null
2025-12-01 Learning Sim-to-Real Humanoid Locomotion in 15 Minutes Younggyo Seo et.al. 2512.01996 null
2025-12-01 RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies Guillermo Garcia-Cobo et.al. 2512.01993 null
2025-12-01 Artemis: Structured Visual Reasoning for Perception Policy Learning Wei Tang et.al. 2512.01988 null
2025-12-01 Forecasting in Offline Reinforcement Learning for Non-stationary Environments Suzan Ece Ada et.al. 2512.01987 null
2025-12-01 From Atomic to Composite: Reinforcement Learning Enables Generalization in Complementary Reasoning Sitao Cheng et.al. 2512.01970 null
2025-12-01 Learned-Rule-Augmented Large Language Model Evaluators Jie Meng et.al. 2512.01958 null
2025-12-01 Spontaneous Symmetry Breaking in Two-dimensional Long-range Heisenberg Model Dingyun Yao et.al. 2512.01956 null
2025-12-01 GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment Haoyang He et.al. 2512.01952 null
2025-12-01 Agentic Policy Optimization via Instruction-Policy Co-Evolution Han Zhou et.al. 2512.01945 null
2025-12-01 Prescribed energy solutions of concave-convex type problems involving sign-changing or vanishing weights Kanishka Perera et.al. 2512.01931 null
2025-12-01 Jacobi Forms of Affine Weight in Higher Cogenus and Nearly Holomorphic Functions Jan Feldmann et.al. 2512.01926 null
2025-12-01 Rectifying LLM Thought from Lens of Optimization Junnan Liu et.al. 2512.01925 null
2025-12-01 Tight Bounds for Feedback Vertex Set Parameterized by Clique-width Narek Bojikian et.al. 2512.01900 null
2025-12-01 New Spiking Architecture for Multi-Modal Decision-Making in Autonomous Vehicles Aref Ghoreishee et.al. 2512.01882 null
2025-12-01 Graph Distance as Surprise: Free Energy Minimization in Knowledge Graph Reasoning Gaganpreet Jhajj et.al. 2512.01878 null
2025-12-01 Topological Order in Deep State Ahmed Abouelkomsan et.al. 2512.01863 null
2025-12-01 Beyond SFT: Reinforcement Learning for Safer Large Reasoning Models with Better Reasoning Ability Jinghan Jia et.al. 2512.01848 null
2025-12-01 OpenREAD: Reinforced Open-Ended Reasoing for End-to-End Autonomous Driving with LLM-as-Critic Songyan Zhang et.al. 2512.01830 null
2025-11-28 Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models Muhammad Maaz et.al. 2511.23478 null
2025-11-28 Video-CoM: Interactive Video Reasoning via Chain of Manipulations Hanoona Rasheed et.al. 2511.23477 null
2025-11-28 Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction Bao Shu et.al. 2511.23476 null
2025-11-28 ThetaEvolve: Test-time Learning on Open Problems Yiping Wang et.al. 2511.23473 null
2025-11-28 The $L$-test: Increasing the Linear Model $F$ -test’s Power Under Sparsity Without Sacrificing Validity Danielle Paulson et.al. 2511.23466 null
2025-11-28 SmallWorlds: Assessing Dynamics Understanding of World Models in Isolated Environments Xinyi Li et.al. 2511.23465 null
2025-11-28 Arbitrary control of the temporal waveform of photons during spontaneous emission Carl Thomas et.al. 2511.23462 null
2025-11-28 Convergence rates of self-repellent random walks, their local time and Event Chain Monte Carlo Andreas Eberle et.al. 2511.23453 null
2025-11-28 ASTRO: Adaptive Stitching via Dynamics-Guided Trajectory Rollouts Hang Yu et.al. 2511.23442 null
2025-11-28 A Heuristic for Matrix Product State Simulation of Out-of-Equilibrium Dynamics of Two-Dimensional Transverse-Field Ising Models Salvatore Mandrà et.al. 2511.23438 null
2025-11-28 Towards Continuous Intelligence Growth: Self-Training, Continual Learning, and Dual-Scale Memory in SuperIntelliAgent Jianzhe Lin et.al. 2511.23436 null
2025-11-28 Variation of Microphysical Parameters in Reverse-shock Scenario Nissim Fraija et.al. 2511.23426 null
2025-11-28 SU( $N_c$) confinement and color $N_c$ -ality V. Tomas Mari Surkau et.al. 2511.23413 null
2025-11-28 From CAD to POMDP: Probabilistic Planning for Robotic Disassembly of End-of-Life Products Jan Baumgärtner et.al. 2511.23407 null
2025-11-28 Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization Jian Li et.al. 2511.23391 null
2025-11-28 Quantum Matrix Spherical Functions Stein Meereboer et.al. 2511.23367 null
2025-11-28 Homomorphism Testing with Resilience to Online Manipulations Esty Kelman et.al. 2511.23363 null
2025-11-28 Synchrotron Self-Compton Model of TeV Afterglows in Gamma-Ray Bursts Edilberto Aguilar-Ruiz et.al. 2511.23349 null
2025-11-28 Efficient Estimation of Sum-Parameters for Multi-Component Complex Exponential Signals with Theoretical Cramer-Rao Bound Analysis Huiguang Zhang et.al. 2511.23318 null
2025-11-28 Emergent Coordination and Phase Structure in Independent Multi-Agent Reinforcement Learning Azusa Yamaguchi et.al. 2511.23315 null
2025-11-26 ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration Hongjin Su et.al. 2511.21689 null
2025-11-26 Multivalued backward stochastic differential equations with jumps and moving boundary Badr Elmansouri et.al. 2511.21679 null
2025-11-26 Escaping the Verifier: Learning to Reason via Demonstrations Locke Cai et.al. 2511.21667 null
2025-11-26 EvilGenie: A Reward Hacking Benchmark Jonathan Gabor et.al. 2511.21654 null
2025-11-26 Optimal Bit Detection in Thermal Noise Communication Systems Under Rician Fading Mohamed El Jbari et.al. 2511.21649 null
2025-11-26 Stochastic Optimal Control of Interacting Particle Systems in Hilbert Spaces and Applications Filippo de Feo et.al. 2511.21646 null
2025-11-26 Estimates for convolution operators on Hardy spaces associated with ball quasi-Banach function spaces Pablo Rocha et.al. 2511.21642 null
2025-11-26 Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative PPO Daniel R. Jiang et.al. 2511.21638 null
2025-11-26 Bang-Bang Evasion: Its Stochastic Optimality and a Terminal-Set-Based Implementation Liraz Mudrik et.al. 2511.21633 null
2025-11-26 Dynamics of generalized abcd Boussinesq solitary waves under a slowly variable bottom André de Laire et.al. 2511.21632 null
2025-11-26 A complete solution of the Erdős-Kleitman matching problem for $n\le 3s$ Andrey Kupavskii et.al. 2511.21628 null
2025-11-26 Two behavioural pseudometrics for continuous-time Markov processes Linan Chen et.al. 2511.21621 null
2025-11-26 Closed Form HJB Solution for Continuous-Time Optimal Control of a Non-Linear Input-Affine System Akash Vyas et.al. 2511.21593 null
2025-11-26 MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training Haotian Xue et.al. 2511.21592 null
2025-11-26 Approximate Bayesian Computation Made Easy: A Practical Guide to ABC-SMC for Dynamical Systems with \texttt{pymc} Mario Castro et.al. 2511.21587 null
2025-11-26 Learning When to Stop: Adaptive Latent Reasoning via Reinforcement Learning Alex Ning et.al. 2511.21581 null
2025-11-26 A Generalized Control Function Approach to Production Function Estimation Ulrich Doraszelski et.al. 2511.21578 null
2025-11-26 BAMAS: Structuring Budget-Aware Multi-Agent Systems Liming Yang et.al. 2511.21572 null
2025-11-26 Some aspects of robustness in modern Markov Chain Monte Carlo Sam Power et.al. 2511.21563 null
2025-11-26 Efficient bayesian spatially varying coefficients modeling for censored data using the vecchia approximation Yacine Mohamed Idir et.al. 2511.21553 null
2025-11-25 RubricRL: Simple Generalizable Rewards for Text-to-Image Generation Xuelu Feng et.al. 2511.20651 null
2025-11-25 Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization Tahira Kazimi et.al. 2511.20647 null
2025-11-25 Reinforcing Action Policies by Prophesying Jiahui Zhang et.al. 2511.20633 null
2025-11-25 Multivariable Wold-Type Decomposition and Analytic Models for a class of left-inverse commuting pairs Monojit Bhattacharjee et.al. 2511.20632 null
2025-11-25 MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models Chieh-Yun Chen et.al. 2511.20629 null
2025-11-25 Optimization of Sums of Bivariate Functions: An Introduction to Relaxation-Based Methods for the Case of Finite Domains Nils Müller et.al. 2511.20607 null
2025-11-25 Precision thermodynamics of the strongly interacting Fermi gas in two dimensions S. Ramachandran et.al. 2511.20599 null
2025-11-25 Inferring the Impacts of Baryonic Feedback from Kinetic Sunyaev-Zeldovich Cross-Correlations Alex Laguë et.al. 2511.20595 null
2025-11-25 Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning Charlotte Beylier et.al. 2511.20591 null
2025-11-25 Hurst exponent and planetary rings Horacio Salomone et.al. 2511.20583 null
2025-11-25 Multi-Resonant-Line Radiative Transfer: Lyman-Alpha Fine Structure and Deuterium Coupling Ethan Stace et.al. 2511.20580 null
2025-11-25 $MC^2$ Mixed Integer and Linear Programming Nick Polson et.al. 2511.20575 null
2025-11-25 Active learning with physics-informed neural networks for optimal sensor placement in deep tunneling through transversely isotropic elastic rocks Alec Tristani et.al. 2511.20574 null
2025-11-25 A Reason-then-Describe Instruction Interpreter for Controllable Video Generation Shengqiong Wu et.al. 2511.20563 null
2025-11-25 Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning Guanjie Chen et.al. 2511.20549 null
2025-11-25 Feature-Modulated UFNO for Improved Prediction of Multiphase Flow in Porous Media Alhasan Abdellatif et.al. 2511.20543 null
2025-11-25 SBP-FDEC: Summation-by-Parts Finite Difference Exterior Calculus Daniel Bach et.al. 2511.20529 null
2025-11-25 Physically Interpretable Interatomic Potentials via Symbolic Regression and Reinforcement Learning Bilvin Varughese et.al. 2511.20506 null
2025-11-25 DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs Yuanhao Li et.al. 2511.20468 null
2025-11-25 Single-hole spectral functions in 1D quantum magnets with different ground states Sibin Yang et.al. 2511.20447 null
2025-11-24 SLMFix: Leveraging Small Language Models for Error Fixing with Reinforcement Learning David Jiahao Fu et.al. 2511.19422 null
2025-11-24 Stochastic Adaptive Optimization with Unreliable Inputs: A Unified Framework for High-Probability Complexity Analysis Katya Scheinberg et.al. 2511.19411 null
2025-11-24 Learning Robust Social Strategies with Large Language Models Dereck Piche et.al. 2511.19405 null
2025-11-24 DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research Rulin Shao et.al. 2511.19399 null
2025-11-24 Identification, estimation and inference in Panel Vector Autoregressions using external instruments Raimondo Pala et.al. 2511.19372 null
2025-11-24 LLM-Driven Stationarity-Aware Expert Demonstrations for Multi-Agent Reinforcement Learning in Mobile Systems Tianyang Duan et.al. 2511.19368 null
2025-11-24 Growing with the Generator: Self-paced GRPO for Video Generation Rui Li et.al. 2511.19356 null
2025-11-24 Leveraging LLMs for reward function design in reinforcement learning control tasks Franklin Cardenoso et.al. 2511.19355 null
2025-11-24 Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning Qihan Huang et.al. 2511.19343 null
2025-11-24 Simulating dynamics of the two-dimensional transverse-field Ising model: a comparative study of large-scale classical numerics Joseph Vovrosh et.al. 2511.19340 null
2025-11-24 On the PB Sudakov: NNLL coefficient, CS kernel and intrinsic-kt Aleksandra Lelek et.al. 2511.19318 null
2025-11-24 PRInTS: Reward Modeling for Long-Horizon Information Seeking Jaewoo Lee et.al. 2511.19314 null
2025-11-24 Diagnosis of mixed-state topological phases in strongly correlated systems via disorder parameters Shao-Hang Shi et.al. 2511.19311 null
2025-11-24 Universal scaling limits at the spectral singularity of structured random matrices Markus Ebke et.al. 2511.19308 null
2025-11-24 AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning Jiayi Zhang et.al. 2511.19304 null
2025-11-24 Tilings of a bounded region of the plane by maximal one-dimensional tiles Eduardo J. Aguilar et.al. 2511.19288 null
2025-11-24 Numerical Approximation In Real Domain Of Special Function Of Product Of A Variable And Its Double Exponential Narinder Kumar Wadhawan et.al. 2511.19270 null
2025-11-24 Leveraging Spatiotemporal Graph Neural Networks for Multi-Store Sales Forecasting Manish Singh et.al. 2511.19267 null
2025-11-24 Interpreting GFlowNets for Drug Discovery: Extracting Actionable Insights for Medicinal Chemistry Amirtha Varshini A S et.al. 2511.19264 null
2025-11-24 MAESTRO: Multi-Agent Environment Shaping through Task and Reward Optimization Boyuan Wu et.al. 2511.19253 null
2025-11-21 Exploring fixed points and eigenstates of quantum systems with reinforcement learning María Laura Olivera-Atencio et.al. 2511.17491 null
2025-11-21 Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination Yolo Yunlong Tang et.al. 2511.17490 null
2025-11-21 Harnessing Data from Clustered LQR Systems: Personalized and Collaborative Policy Optimization Vinay Kanakeri et.al. 2511.17489 null
2025-11-21 Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards Zhen Wang et.al. 2511.17473 null
2025-11-21 InTAct: Interval-based Task Activation Consolidation for Continual Learning Patryk Krukowski et.al. 2511.17439 null
2025-11-21 Iterating marginalized Bayes maps for likelihood maximization with application to nonlinear panel models Jesse Wheeler et.al. 2511.17438 null
2025-11-21 Multi-Agent Pointer Transformer: Seq-to-Seq Reinforcement Learning for Multi-Vehicle Dynamic Pickup-Delivery Problems Zengyu Zou et.al. 2511.17435 null
2025-11-21 Bayesian Bridge Gaussian Process Regression Minshen Xu et.al. 2511.17415 null
2025-11-21 Beyond Multiple Choice: A Hybrid Framework for Unifying Robust Evaluation and Verifiable Reasoning Training Yesheng Liu et.al. 2511.17405 null
2025-11-21 MorphSeek: Fine-grained Latent Representation-Level Policy Optimization for Deformable Image Registration Runxun Zhang et.al. 2511.17392 null
2025-11-21 Human Imitated Bipedal Locomotion with Frequency Based Gait Generator Network Yusuf Baran Ates et.al. 2511.17387 null
2025-11-21 Abstract fractional linear transformations David Handelman et.al. 2511.17383 null
2025-11-21 Critical BKT dynamics in the archetypal 2D spin system Ba $_2$CuSi$_2$O$_6$Cl$_2$ K. M. Ranjith et.al. 2511.17381 null
2025-11-21 Agility Meets Stability: Versatile Humanoid Control with Heterogeneous Data Yixuan Pan et.al. 2511.17373 null
2025-11-21 R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability Runyu Lu et.al. 2511.17367 null
2025-11-21 SVRecon: Sparse Voxel Rasterization for Surface Reconstruction Seunghun Oh et.al. 2511.17364 null
2025-11-21 Convergence and stability of Q-learning in Hierarchical Reinforcement Learning Massimiliano Manenti et.al. 2511.17351 null
2025-11-21 ReBaPL: Repulsive Bayesian Prompt Learning Yassir Bendou et.al. 2511.17339 null
2025-11-21 Law-Strength Frontiers and a No-Free-Lunch Result for Law-Seeking Reinforcement Learning on Volatility Law Manifolds Jian’an Zhang et.al. 2511.17304 null
2025-11-21 MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning Wenrui Zhang et.al. 2511.17300 null
2025-11-20 EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards Omkat Thawakar et.al. 2511.16672 null
2025-11-20 Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation Ziyu Guo et.al. 2511.16671 null
2025-11-20 Learning to Think Fast and Slow for Visual Language Models Chenyu Lin et.al. 2511.16670 null
2025-11-20 Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO Junhao Cheng et.al. 2511.16669 null
2025-11-20 SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation Zhenyuan Qin et.al. 2511.16666 null
2025-11-20 Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter Qinghao Hu et.al. 2511.16665 null
2025-11-20 Dexterity from Smart Lenses: Multi-Fingered Robot Manipulation with In-the-Wild Human Demonstrations Irmak Guzey et.al. 2511.16661 null
2025-11-20 Stabilizing Policy Gradient Methods via Reward Profiling Shihab Ahmed et.al. 2511.16629 null
2025-11-20 MedBayes-Lite: Bayesian Uncertainty Quantification for Safe Clinical Decision Support Elias Hossain et.al. 2511.16625 null
2025-11-20 Deep Learning Framework for Enhanced Neutrino Reconstruction of Single-line Events in the ANTARES Telescope A. Albert et.al. 2511.16614 null
2025-11-20 KrkNLO matching and phenomenology for vector boson processes Pratixan Sarmah et.al. 2511.16605 null
2025-11-20 Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization Yi Zhang et.al. 2511.16602 null
2025-11-20 Green Resilience of Cyber-Physical Systems: Doctoral Dissertation Diaeddin Rimawi et.al. 2511.16593 null
2025-11-20 gfnx: Fast and Scalable Library for Generative Flow Networks in JAX Daniil Tiapkin et.al. 2511.16592 null
2025-11-20 Differential decay rate of $B^+ \to J/ψK^+$ with the LHCb Upgrade I experiment LHCb collaboration et.al. 2511.16564 null
2025-11-20 Reinforcement learning of quantum circuit architectures for molecular potential energy curves Maureen Krumtünger et.al. 2511.16559 null
2025-11-20 MiMo-Embodied: X-Embodied Foundation Model Technical Report Xiaoshuai Hao et.al. 2511.16518 null
2025-11-20 Polynomial-Time Algorithms for Computing the Nucleolus: An Assessment Holger I. Meinhardt et.al. 2511.16517 null
2025-11-20 Quasi-metric spaces on which real-valued continuous functions are uniformly continuous Om Dev Singh et.al. 2511.16503 null
2025-11-20 Quantum corrections in general relativity explored through a GUP-inspired maximal acceleration analysis Christian Corda et.al. 2511.16502 null
2025-11-19 GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization Yikun Wang et.al. 2511.15705 null
2025-11-19 The Impact of Quantization on Large Reasoning Model Reinforcement Learning Medha Kumar et.al. 2511.15694 null
2025-11-19 Semiparametric Estimation of Fractional Integration: An Evaluation of Local Whittle Methods Jason R. Blevins et.al. 2511.15689 null
2025-11-19 Quantum measurement tomography with mini-batch stochastic gradient descent Akshay Gaikwad et.al. 2511.15682 null
2025-11-19 VisPlay: Self-Evolving Vision-Language Models from Images Yicheng He et.al. 2511.15661 null
2025-11-19 Continual Reinforcement Learning for Cyber-Physical Systems: Lessons Learned and Open Challenges Kim N. Nolle et.al. 2511.15652 null
2025-11-19 A Green’s function approach to linearized Monge-Ampère equations in divergence form and application to singular Abreu type equations Chong Gu et.al. 2511.15621 null
2025-11-19 Magnetic electron-hole asymmetry in cuprates: a computational revisit Jiong Mei et.al. 2511.15608 null
2025-11-19 SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models Senyu Fei et.al. 2511.15605 null
2025-11-19 Learning from Mistakes: Loss-Aware Memory Enhanced Continual Learning for LiDAR Place Recognition Xufei Wang et.al. 2511.15597 null
2025-11-19 A sharp threshold for arithmetic effects on the tail probabilities of lacunary sums Christoph Aistleitner et.al. 2511.15595 null
2025-11-19 QTIS: A QAOA-Based Quantum Time Interval Scheduler José A. Tirado-Domínguez et.al. 2511.15590 null
2025-11-19 Meta-Black-Box Optimization with Bi-Space Landscape Analysis and Dual-Control Mechanism for SAEA Yukun Du et.al. 2511.15551 null
2025-11-19 A Physics Informed Machine Learning Framework for Optimal Sensor Placement and Parameter Estimation Georgios Venianakis et.al. 2511.15543 null
2025-11-19 Theoretical Closed-loop Stability Bounds for Dynamical System Coupled with Diffusion Policies Gabriel Lauzier et.al. 2511.15520 null
2025-11-19 Robust H-infinity control and worst-case search in constrained parametric space Ervan Kassarian et.al. 2511.15480 null
2025-11-19 Challenging the $ω_0ω_a$ CDM parametrization through rational expansions in view of DESI data release Youri Carloni et.al. 2511.15472 null
2025-11-19 Gini Score under Ties and Case Weights Alexej Brauer et.al. 2511.15446 null
2025-11-19 Measure finite topology on the ring of measurable functions Soumajit Dey et.al. 2511.15436 null
2025-11-19 Viscous Dark Energy and Mass-Varying Dark Matter in Lyra Manifold: Cosmological Dynamics and Observational Constraints Giridhari Deogharia et.al. 2511.15422 null
2025-11-18 UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning Rui Tian et.al. 2511.14760 null
2025-11-18 $π^{*}_{0.6}$ : a VLA That Learns From Experience Ali Amin et.al. 2511.14759 null
2025-11-18 Heterogeneous Multi-Agent Proximal Policy Optimization for Power Distribution System Restoration Parya Dolatyabi et.al. 2511.14730 null
2025-11-18 Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization Zonghao Chen et.al. 2511.14710 null
2025-11-18 Nonparametric Uniform Inference in Binary Classification and Policy Values Nan Liu et.al. 2511.14700 null
2025-11-18 SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction Biaojie Zeng et.al. 2511.14684 null
2025-11-18 Quadratic Term Correction on Heaps’ Law Oscar Fontanelli et.al. 2511.14683 null
2025-11-18 Estimation of Spatial and Temporal Autoregressive Effects using LASSO - An Example of Hourly Particulate Matter Concentrations Elkanah Nyabuto et.al. 2511.14666 null
2025-11-18 NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards Chia-Yu Hung et.al. 2511.14659 null
2025-11-18 Failure to Mix: Large language models struggle to answer according to desired probability distributions Ivy Yuqian Yang et.al. 2511.14630 null
2025-11-18 Fusing Biomechanical and Spatio-Temporal Features for Fall Prediction: Characterizing and Mitigating the Simulation-to-Reality Gap Md Fokhrul Islam et.al. 2511.14620 null
2025-11-18 Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning Ruoyu Qin et.al. 2511.14617 null
2025-11-18 XAttn-BMD: Multimodal Deep Learning with Cross-Attention for Femoral Neck Bone Mineral Density Estimation Yilin Zhang et.al. 2511.14604 null
2025-11-18 Anomalous spontaneous induction of magnetic and electric fields in dense quark matter E. J. Ferrer et.al. 2511.14602 null
2025-11-18 ReflexGrad: Three-Way Synergistic Architecture for Zero-Shot Generalization in LLM Agents Ankush Kadu et.al. 2511.14584 null
2025-11-18 Masked IRL: LLM-Guided Reward Disambiguation from Demonstrations and Language Minyoung Hwang et.al. 2511.14565 null
2025-11-18 Linear Combinations of Logarithms of $L$ -functions over Function Fields at Microscopic Shifts and Beyond Fatma Çiçek et.al. 2511.14563 null
2025-11-18 DeepBlip: Estimating Conditional Average Treatment Effects Over Time Haorui Ma et.al. 2511.14545 null
2025-11-18 Solving Navier-Stokes Equations Using Data-free Physics-Informed Neural Networks With Hard Boundary Conditions Ritik Pal et.al. 2511.14497 null
2025-11-18 Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning Mingyue Cheng et.al. 2511.14460 null
2025-11-17 TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models Harold Haodong Chen et.al. 2511.13704 null
2025-11-17 Ultra Low Overhead Syndrome Extraction for the Steane code Boldizsár Poór et.al. 2511.13700 null
2025-11-17 Resilient Distribution Network Planning against Dynamic Malicious Power Injection Attacks Hampei Sasahara et.al. 2511.13698 null
2025-11-17 The physical properties of post-mass-transfer binaries Rhys Seeburger et.al. 2511.13692 null
2025-11-17 Distribution Matching Distillation Meets Reinforcement Learning Dengyang Jiang et.al. 2511.13649 null
2025-11-17 Sense and Sensitivity - I. Uncertainty analysis of the gas-phase chemistry in AGB outflows M. Van de Sande et.al. 2511.13638 null
2025-11-17 Towards Multimodal Representation Learning in Paediatric Kidney Disease Ana Durica et.al. 2511.13637 null
2025-11-17 Exploring the production of Terbium-161 in the Brazilian Multipurpose Reactor Fernando Cozim Melges et.al. 2511.13634 null
2025-11-17 P1: Mastering Physics Olympiads with Reinforcement Learning Jiacheng Chen et.al. 2511.13612 null
2025-11-17 Thermodynamics of the Fermi-Hubbard Model through Stochastic Calculus and Girsanov Transformation Detlef Lehmann et.al. 2511.13581 null
2025-11-17 Electron Correlation by Exchange Mapping in Electronic Structure Calculations Jerry L. Whitten et.al. 2511.13570 null
2025-11-17 Infinite-Horizon Optimal Control of Jump-Diffusion Models for Pollution-Dependent Disasters Daria Sakhanda et.al. 2511.13568 null
2025-11-17 Artificial Intelligence-driven Intelligent Wearable Systems: A full-stack Integration from Material Design to Personalized Interaction Jingyi Zhao et.al. 2511.13565 null
2025-11-17 Efficient Simulation of Hawkes Processes using their Affine Volterra Structure Eduardo Abi Jaber et.al. 2511.13554 null
2025-11-17 TSE-Net: Semi-supervised Monocular Height Estimation from Single Remote Sensing Images Sining Chen et.al. 2511.13552 null
2025-11-17 SO(3) real algebra method for finite baryon-number density QCD Hideo Suganuma et.al. 2511.13551 null
2025-11-17 MDIntrinsicDimension: Dimensionality-Based Analysis of Collective Motions in Macromolecules from Molecular Dynamics Trajectories Irene Cazzaniga et.al. 2511.13550 null
2025-11-17 Full range of infinite point blow-up exponents for the critical generalized KdV equation Nailya Manatova et.al. 2511.13538 null
2025-11-17 Smoothed-Cubic Spin-Glass Model of Random Lasers Marcello Benedetti et.al. 2511.13508 null
2025-11-17 Sampling Density for Gabor Phase Retrieval Ting Chen et.al. 2511.13500 null
2025-11-14 The Maximal Variance of Unilaterally Truncated Gaussian and Chi Distributions Robert J. Petrella et.al. 2511.11566 null
2025-11-14 Low-energy enhancement in the magnetic dipole radiation of actinide nuclei C. Rodgers et.al. 2511.11565 null
2025-11-14 Two Useful Facts About Generating Functions Alex Kasman et.al. 2511.11559 null
2025-11-14 Drone Swarm Energy Management Michael Z. Zgurovsky et.al. 2511.11557 null
2025-11-14 Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping Dena Mujtaba et.al. 2511.11551 null
2025-11-14 Parameterized complexity of the f-Critical Set problem Thiago Marcilon et.al. 2511.11546 null
2025-11-14 STEM EBIC as a Quantitative Probe of Semiconductor Devices Sebastian Schneider et.al. 2511.11528 null
2025-11-14 W2S-AlignTree: Weak-to-Strong Inference-Time Alignment for Large Language Models via Monte Carlo Tree Search Zhenyu Ding et.al. 2511.11518 null
2025-11-14 Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation Mohamad Amin Mohamadi et.al. 2511.11500 null
2025-11-14 Learning and Testing Convex Functions Renato Ferreira Pinto et.al. 2511.11498 null
2025-11-14 A Recursive Theory of Variational State Estimation: The Dynamic Programming Approach Filip Tronarp et.al. 2511.11497 null
2025-11-14 Intrinsic Dimension Estimation for Radio Galaxy Zoo using Diffusion Models Joan Font-Quer Roset et.al. 2511.11490 null
2025-11-14 Risk-Aware Deep Reinforcement Learning for Dynamic Portfolio Optimization Emmanuel Lwele et.al. 2511.11481 null
2025-11-14 Context-aware Adaptive Visualizations for Critical Decision Making Angela Lopez-Cardona et.al. 2511.11476 null
2025-11-14 Estimating the Effects of Heatwaves on Health: A Causal Inference Framework Giulio Grossi et.al. 2511.11433 null
2025-11-14 Multi-Phase Spacecraft Trajectory Optimization via Transformer-Based Reinforcement Learning Amit Jain et.al. 2511.11402 null
2025-11-14 Variational Quantum Algorithms for Particle Track Reconstruction Vincenzo Lipardi et.al. 2511.11397 null
2025-11-14 Modeling the Multi-Wavelength Afterglow of Short Gamma-Ray Bursts with a Plateau Phase Chen Deng et.al. 2511.11396 null
2025-11-14 Robust and Efficient Communication in Multi-Agent Reinforcement Learning Zejiao Liu et.al. 2511.11393 null
2025-11-14 Optimal Dividend, Reinsurance and Capital Injection Strategies for Collaborating Business Lines: The Case of Excess-of-Loss Reinsurance Tim J. Boonen et.al. 2511.11383 null
2025-11-13 Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling Jiahao Wang et.al. 2511.10648 null
2025-11-13 Black-Box On-Policy Distillation of Large Language Models Tianzhu Ye et.al. 2511.10643 null
2025-11-13 Robot Crash Course: Learning Soft and Stylized Falling Pascal Strauch et.al. 2511.10635 null
2025-11-13 Instella: Fully Open Language Models with Stellar Performance Jiang Liu et.al. 2511.10628 null
2025-11-13 Global Solutions to Non-Convex Functional Constrained Problems with Hidden Convexity Ilyas Fatkhullin et.al. 2511.10626 null
2025-11-13 Algorithm Design and Stronger Guarantees for the Improving Multi-Armed Bandits Problem Avrim Blum et.al. 2511.10619 null
2025-11-13 The $L_p$ -error rate for randomized quasi-Monte Carlo self-normalized importance sampling of unbounded integrands Jiarui Du et.al. 2511.10599 null
2025-11-13 A Fast Earth-scattering Formalism for Light Dark Matter with Dark Photon Mediators Agustín Lantero-Barreda et.al. 2511.10589 null
2025-11-13 Towards Emotionally Intelligent and Responsible Reinforcement Learning Garapati Keerthana et.al. 2511.10573 null
2025-11-13 Non-Resonant Alpha-Induced Neutron-Emission: A Multi- Method Comparison Of Nuclear Reaction Rates Bhavay Luthra et.al. 2511.10536 null
2025-11-13 Parallel and GPU accelerated code for phase-field and reaction-diffusion simulations Steven A. Silber et.al. 2511.10508 null
2025-11-13 Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following Yun He et.al. 2511.10507 null
2025-11-13 Strategic Opponent Modeling with Graph Neural Networks, Deep Reinforcement Learning and Probabilistic Topic Modeling Georgios Chalkiadakis et.al. 2511.10501 null
2025-11-13 Spin Liquids on the Tetratrillium Lattice Matías G. Gonzalez et.al. 2511.10489 null
2025-11-13 Revealing the Connection Between the Filamentary Hierarchy and Star Cluster Formation in a Simulated NGC 628 Galaxy Tamara Koletic et.al. 2511.10486 null
2025-11-13 Reasoning About Intent for Ambiguous Requests Irina Saparina et.al. 2511.10453 null
2025-11-13 Continuum Dropout for Neural Differential Equations Jonghun Lee et.al. 2511.10446 null
2025-11-13 A Decomposition Approach to Solving Numerical Constraint Satisfaction Problems on Directed Acyclic Graphs Max Mowbray et.al. 2511.10426 null
2025-11-13 Generalizing Analogical Inference from Boolean to Continuous Domains Francisco Cunha et.al. 2511.10416 null
2025-11-13 Strangeness enhancement at its extremes: multiple (multi-)strange hadron production in pp collisions at $\mathbf{\sqrt{\textit{s}} = 5.02}$ TeV ALICE Collaboration et.al. 2511.10413 null
2025-11-10 Robot Learning from a Physical World Model Jiageng Mao et.al. 2511.07416 null
2025-11-10 Unified Humanoid Fall-Safety Policy from a Few Demonstrations Zhengjie Xu et.al. 2511.07407 null
2025-11-10 SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards Hunar Batra et.al. 2511.07403 null
2025-11-10 Policy Learning for Perturbance-wise Linear Quadratic Control Problem Haoran Zhang et.al. 2511.07388 null
2025-11-10 samsara: A Continuous-Time Markov Chain Monte Carlo Sampler for Trans-Dimensional Bayesian Analysis Gabriele Astorino et.al. 2511.07385 null
2025-11-10 Real-Time LiDAR Super-Resolution via Frequency-Aware Multi-Scale Fusion June Moh Goo et.al. 2511.07377 null
2025-11-10 Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training Dake Bu et.al. 2511.07372 null
2025-11-10 Consistency Is Not Always Correct: Towards Understanding the Role of Exploration in Post-Training Reasoning Dake Bu et.al. 2511.07368 null
2025-11-10 UAV-Assisted Resilience in 6G and Beyond Network Energy Saving: A Multi-Agent DRL Approach Dao Lan Vy Dinh et.al. 2511.07366 null
2025-11-10 Bayesian compartmental modelling of MRSA transmission within hospitals in Edmonton, Canada Ruoyu Li et.al. 2511.07353 null
2025-11-10 Smoothing Out Sticking Points: Sampling from Discrete-Continuous Mixtures with Dynamical Monte Carlo by Mapping Discrete Mass into a Latent Universe Andrew Chin et.al. 2511.07340 null
2025-11-10 Grounding Computer Use Agents on Human Demonstrations Aarash Feizi et.al. 2511.07332 null
2025-11-10 Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Training Artyom Sorokin et.al. 2511.07328 null
2025-11-10 IterResearch: Rethinking Long-Horizon Agents via Markovian State Reconstruction Guoxin Chen et.al. 2511.07327 null
2025-11-10 FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation Song Jin et.al. 2511.07322 null
2025-11-10 RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments Zhiyuan Zeng et.al. 2511.07317 null
2025-11-10 Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time Search Samuel Sokota et.al. 2511.07312 null
2025-11-10 Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization Sayambhu Sen et.al. 2511.07288 null
2025-11-10 Bridging the divide: axion searches and axino phenomenology at colliders Gabe Hoshino et.al. 2511.07224 null
2025-11-10 Unlocking the Regression Space Liudas Giraitis et.al. 2511.07183 null
2025-11-07 Visual Spatial Tuning Rui Yang et.al. 2511.05491 null
2025-11-07 TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning Junwen Pan et.al. 2511.05489 null
2025-11-07 Minority-Aware Satisfaction Estimation in Dialogue Systems via Preference-Adaptive Reinforcement Learning Yahui Fu et.al. 2511.05407 null
2025-11-07 Sample Complexity of Distributionally Robust Off-Dynamics Reinforcement Learning with Online Interaction Yiting He et.al. 2511.05396 null
2025-11-07 PreResQ-R1: Towards Fine-Grained Rank-and-Score Reinforcement Learning for Visual Quality Assessment via Preference-Response Disentangled Policy Optimization Zehui Feng et.al. 2511.05393 null
2025-11-07 TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework Chao Zhang et.al. 2511.05385 null
2025-11-07 EMPEROR I. Exoplanet MCMC parallel tempering for RV orbit retrieval Pablo A. Peña R. et.al. 2511.05331 null
2025-11-07 QUESTER: Query Specification for Generative Retrieval Arthur Satouf et.al. 2511.05301 null
2025-11-07 Reflective Personalization Optimization: A Post-hoc Rewriting Framework for Black-Box Large Language Models Teqi Hao et.al. 2511.05286 null
2025-11-07 DeepEyesV2: Toward Agentic Multimodal Model Jack Hong et.al. 2511.05271 null
2025-11-07 An End-to-End Deep Reinforcement Learning Approach for Solving the Traveling Salesman Problem with Drones Taihelong Zeng et.al. 2511.05265 null
2025-11-07 Adaptive Entanglement-Aware Routing for Satellite Quantum Networks under Orbital and Atmospheric Variability Dhrumil Bhatt et.al. 2511.05228 null
2025-11-07 Fast and Scalable Evaluation of Unbiased Atomic Forces in ab initio Variational Monte Carlo via the Lagrangian Technique Kousuke Nakano et.al. 2511.05222 null
2025-11-07 Emergence from Emergence: Financial Market Simulation via Learning with Heterogeneous Preferences Ryuko Hashimoto et.al. 2511.05207 null
2025-11-07 Follow-Me in Micro-Mobility with End-to-End Imitation Learning Sahar Salimpour et.al. 2511.05158 null
2025-11-07 Mass determination of the three long-period Neptune- and sub-Neptune-sized planets transiting TOI-282 A. Barone et.al. 2511.05147 null
2025-11-07 Exponential Spatiotemporal GARCH Model with Asymmetric Volatility Spillovers Ariane Nidelle Meli Chrisko et.al. 2511.05126 null
2025-11-07 Real-World Adverse Weather Image Restoration via Dual-Level Reinforcement Learning with High-Quality Cold Start Fuyang Liu et.al. 2511.05095 null
2025-11-07 FM4Com: Foundation Model for Scene-Adaptive Communication Strategy Optimization Zhaoyang Li et.al. 2511.05094 null
2025-11-07 Optical studies of scintillation detectors for precision beta-energy measurements S. Vanlangendonck et.al. 2511.05083 null
2025-11-06 GentleHumanoid: Learning Upper-body Compliance for Contact-rich Human and Object Interaction Qingzhou Lu et.al. 2511.04679 null
2025-11-06 On the Exoplanet Yield of Gaia Astrometry Caleb Lammers et.al. 2511.04673 null
2025-11-06 Forgetting is Everywhere Ben Sanati et.al. 2511.04666 null
2025-11-06 Environment Agnostic Goal-Conditioning, A Study of Reward-Free Autonomous Learning Hampus Åström et.al. 2511.04598 null
2025-11-06 Combining Harmonic Sampling with the Worm Algorithm to Improve the Efficiency of Path Integral Monte Carlo Sourav Karmakar et.al. 2511.04597 null
2025-11-06 Continuous matrix product operators for quantum fields Erickson Tjoa et.al. 2511.04545 null
2025-11-06 End-to-End Reinforcement Learning of Koopman Models for eNMPC of an Air Separation Unit Daniel Mayfrank et.al. 2511.04522 null
2025-11-06 Approaching the thermodynamic limit of a bounded one-component plasma D. I. Zhukhovitskii et.al. 2511.04516 null
2025-11-06 Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion Oskar Lappi et.al. 2511.04489 null
2025-11-06 V-Thinker: Interactive Thinking with Images Runqi Qiao et.al. 2511.04460 null
2025-11-06 Fitting Reinforcement Learning Model to Behavioral Data under Bandits Hao Zhu et.al. 2511.04454 null
2025-11-06 The Peril of Preference: Why GRPO fails on Ordinal Rewards Anisha Garg et.al. 2511.04439 null
2025-11-06 Temporal Action Selection for Action Chunking Yueyang Weng et.al. 2511.04421 null
2025-11-06 On the Estimation of Own Funds for Life Insurers: A Study of Direct, Indirect, and Control Variate Methods in a Risk-Neutral Pricing Framework Mark-Oliver Wolf et.al. 2511.04412 null
2025-11-06 Mixed-State Measurement-Induced Phase Transitions in Imaginary-Time Dynamics Yi-Ming Ding et.al. 2511.04402 null
2025-11-06 Artificial Precision Polarization Array: Sensitivity for the axion-like dark matter with clock satellites Hanyu Jiang et.al. 2511.04400 null
2025-11-06 GraSP-VLA: Graph-based Symbolic Action Representation for Long-Horizon Planning with VLA Policies Maëlic Neau et.al. 2511.04357 null
2025-11-06 Stochastic simulation of partial discharge inception Jannis Teunissen et.al. 2511.04356 null
2025-11-06 MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments Kuankuan Sima et.al. 2511.04320 null
2025-11-06 DeepPAAC: A New Deep Galerkin Method for Principal-Agent Problems Michael Ludkovski et.al. 2511.04309 null
2025-11-05 Outbidding and Outbluffing Elite Humans: Mastering Liar’s Poker via Self-Play and Reinforcement Learning Richard Dewey et.al. 2511.03724 null
2025-11-05 Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards Guanning Zeng et.al. 2511.03710 null
2025-11-05 AnaFlow: Agentic LLM-based Workflow for Reasoning-Driven Explainable and Sample-Efficient Analog Circuit Sizing Mohsen Ahmadzadeh et.al. 2511.03697 null
2025-11-05 Behavior-Adaptive Q-Learning: A Unifying Framework for Offline-to-Online RL Lipeng Zu et.al. 2511.03695 null
2025-11-05 Simulation-Based Validation of an Integrated 4D/5D Digital-Twin Framework for Predictive Construction Control Atena Khoshkonesh et.al. 2511.03684 null
2025-11-05 DQN Performance with Epsilon Greedy Policies and Prioritized Experience Replay Daniel Perkins et.al. 2511.03670 null
2025-11-05 Towards Formalizing Reinforcement Learning Theory Shangtong Zhang et.al. 2511.03618 null
2025-11-05 Going Beyond Expert Performance via Deep Implicit Imitation Reinforcement Learning Iason Chrysomallis et.al. 2511.03616 null
2025-11-05 Bayesian Topological Analysis of Functional Brain Networks Xukun Zhu et.al. 2511.03605 null
2025-11-05 Tensor-Efficient High-Dimensional Q-learning Junyi Wu et.al. 2511.03595 null
2025-11-05 PerfDojo: Automated ML Library Generation for Heterogeneous Architectures Andrei Ivanov et.al. 2511.03586 null
2025-11-05 Realization of repulsive polarons in the strongly correlated regime René Henke et.al. 2511.03569 null
2025-11-05 Imitation Learning in the Deep Learning Era: A Novel Taxonomy and Recent Advances Iason Chrysomallis et.al. 2511.03565 null
2025-11-05 Learning Without Critics? Revisiting GRPO in Classical Reinforcement Learning Environments Bryan L. M. de Oliveira et.al. 2511.03527 null
2025-11-05 Reinforcement Learning Using known Invariances Alexandru Cioba et.al. 2511.03473 null
2025-11-05 Unraveling Deconfined Quantum Criticality in Non-Hermitian Easy-Plane $J$-$Q$ Model Xuan Zou et.al. 2511.03456 null
2025-11-05 QMeCha: quantum Monte Carlo package for fermions in embedding environments Matteo Barborini et.al. 2511.03439 null
2025-11-05 The moment is here: a generalised class of estimators for fuzzy regression discontinuity designs Stuart Lane et.al. 2511.03424 null
2025-11-05 Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG Longpeng Qiu et.al. 2511.03410 null
2025-11-05 Adaptable Hindsight Experience Replay for Search-Based Learning Alexandros Vazaios et.al. 2511.03405 null
2025-11-04 Audience Amplified: Virtual Audiences in Asynchronously Performed AR Theater You-Jin Kim et.al. 2511.02807 null
2025-11-04 MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning Qianhao Yuan et.al. 2511.02805 null
2025-11-04 From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos Xun Wang et.al. 2511.02762 null
2025-11-04 Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning Bowen Jin et.al. 2511.02755 null
2025-11-04 Approximation by Certain Complex Nevai Operators : Theory and Applications Priyanka Majethiya et.al. 2511.02750 null
2025-11-04 From Densities to Potentials: Benchmarking Local Exchange-Correlation Approximations Visagan Ravindran et.al. 2511.02744 null
2025-11-04 Bayesian full waveform inversion with learned prior using deep convolutional autoencoder Shuhua Hu et.al. 2511.02737 null
2025-11-04 Observational tests of the conformal osculating Barthel-Kropina cosmological model Himanshu Chaudhary et.al. 2511.02729 null
2025-11-04 VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models Zhicheng Zhang et.al. 2511.02712 null
2025-11-04 The two-dimensional optical Su-Schrieffer-Heeger model: ground state and thermodynamic properties Jadson L. Portela e Silva et.al. 2511.02707 null
2025-11-04 Optimizing Kernel Discrepancies via Subset Selection Deyao Chen et.al. 2511.02706 null
2025-11-04 Policy Gradient Methods for Information-Theoretic Opacity in Markov Decision Processes Chongyang Shi et.al. 2511.02704 null
2025-11-04 Identification and Estimation of Continuous-Time Dynamic Discrete Choice Games Jason R. Blevins et.al. 2511.02701 null
2025-11-04 Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs Georgios Tzannetos et.al. 2511.02690 null
2025-11-04 RL-Aided Cognitive ISAC: Robust Detection and Sensing-Communication Trade-offs Adam Umra et.al. 2511.02672 null
2025-11-04 Natural-gas storage modelling by deep reinforcement learning Tiziano Balaconi et.al. 2511.02646 null
2025-11-04 Supernova Classification using the Recurrent Neural Network in the CSST Ultra-Deep Field Survey Minglin Wang et.al. 2511.02631 null
2025-11-04 Adaptive GR(1) Specification Repair for Liveness-Preserving Shielding in Reinforcement Learning Tiberiu-Andrei Georgescu et.al. 2511.02605 null
2025-11-04 CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency Ehsan Aghazadeh et.al. 2511.02603 null
2025-11-04 Directional-Clamp PPO Gilad Karpel et.al. 2511.02577 null
2025-10-31 Challenges in Credit Assignment for Multi-Agent Reinforcement Learning in Open Agent Systems Alireza Saleh Abadi et.al. 2510.27659 null
2025-10-31 Probing cosmic isotropy with Gamma-ray bursts: A dipole and quadrupole analysis of BATSE and Fermi GBM data Debosi Mondal et.al. 2510.27644 null
2025-10-31 A Comprehensive Stress Test of Truncated Hilbert Space Bases against Green’s function Monte Carlo in U(1) Lattice Gauge Theory Timo Jakobs et.al. 2510.27611 null
2025-10-31 Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning Yuhong Liu et.al. 2510.27606 null
2025-10-31 DiffstarPop: A generative physical model of galaxy star formation history Alex Alarcon et.al. 2510.27604 null
2025-10-31 MARAG-R1: Beyond Single Retriever via Reinforcement-Learned Multi-Tool Agentic Retrieval Qi Luo et.al. 2510.27569 null
2025-10-31 Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval Yulong Hui et.al. 2510.27566 null
2025-10-31 Holographic equation of state matched with hadron gas equation as a tool for the study of the quark-gluon plasma evolution A. V. Anufriev et.al. 2510.27541 null
2025-10-31 pDANSE: Particle-based Data-driven Nonlinear State Estimation from Nonlinear Measurements Anubhab Ghosh et.al. 2510.27503 null
2025-10-31 Study of Central Exclusive Production of $π^+π^-$, $K^+K^-$ and $p \bar{p}$ Pairs in Proton-Proton Collisions at $\sqrt{s} = 510$ GeV with the STAR Detector at RHIC Tomas Truhlar et.al. 2510.27482 null
2025-10-31 VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision Xuan Gong et.al. 2510.27462 null
2025-10-31 Modeling partially-ionized dense plasma using wavepacket molecular dynamics Daniel Plummer et.al. 2510.27446 null
2025-10-31 Learning Soft Robotic Dynamics with Active Exploration Hehui Zheng et.al. 2510.27428 null
2025-10-31 DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains Tian Liang et.al. 2510.27419 null
2025-10-31 Dialogue as Discovery: Navigating Human Intent Through Principled Inquiry Jianwen Sun et.al. 2510.27410 null
2025-10-31 Muon veto system for the CROSS double-beta decay search experiment A. S. Barabash et.al. 2510.27406 null
2025-10-31 Realistic pedestrian-driver interaction modelling using multi-agent RL with human perceptual-motor constraints Yueyang Wang et.al. 2510.27383 null
2025-10-31 Reasoning Models Sometimes Output Illegible Chains of Thought Arun Jose et.al. 2510.27338 null
2025-10-31 When AI Trading Agents Compete: Adverse Selection of Meta-Orders by Reinforcement Learning-Based Market Making Ali Raza Jafree et.al. 2510.27334 null
2025-10-31 Reinforcement Learning for Long-Horizon Unordered Tasks: From Boolean to Coupled Reward Machines Kristina Levina et.al. 2510.27329 null
2025-10-30 Defeating the Training-Inference Mismatch via FP16 Penghui Qi et.al. 2510.26788 null
2025-10-30 Automated event generation for S-wave quarkonium and leptonium production in NRQCD and NRQED Alice Colpani Serri et.al. 2510.26773 null
2025-10-30 The Oversight Game: Learning to Cooperatively Balance an AI Agent’s Safety and Autonomy William Overman et.al. 2510.26752 null
2025-10-30 Characterization of the H2M Monolithic CMOS Sensor Rafael Ballabriga et.al. 2510.26741 null
2025-10-30 A General Incentives-Based Framework for Fairness in Multi-agent Resource Allocation Ashwin Kumar et.al. 2510.26740 null
2025-10-30 Wavefront Curvature and Transverse Atomic Motion in Time-Resolved Atom Interferometry: Impact and Mitigation Noam Mouelle et.al. 2510.26739 null
2025-10-30 Emergence of charge- $4e$ superconductivity from 2D nematic superconductors Xuan Zou et.al. 2510.26720 null
2025-10-30 Stabilizing Rayleigh-Benard convection with reinforcement learning trained on a reduced-order model Qiwei Chen et.al. 2510.26705 null
2025-10-30 Kimi Linear: An Expressive, Efficient Attention Architecture Kimi Team et.al. 2510.26692 null
2025-10-30 Flinch: A Differentiable Framework for Field-Level Inference of Cosmological parameters from curved sky data Andrea Crespi et.al. 2510.26691 null
2025-10-30 Generative sampling with physics-informed kernels Friederike Ihssen et.al. 2510.26678 null
2025-10-30 Action-Driven Processes for Continuous-Time Control Ruimin He et.al. 2510.26672 null
2025-10-30 Hybrid Consistency Policy: Decoupling Multi-Modal Diversity and Real-Time Efficiency in Robotic Manipulation Qianyou Zhao et.al. 2510.26670 null
2025-10-30 The Era of Agentic Organization: Learning to Organize with Language Models Zewen Chi et.al. 2510.26658 null
2025-10-30 Hybrid DQN-TD3 Reinforcement Learning for Autonomous Navigation in Dynamic Environments Xiaoyi He et.al. 2510.26646 null
2025-10-30 Low-Altitude UAV-Carried Movable Antenna for Joint Wireless Power Transfer and Covert Communications Chuang Zhang et.al. 2510.26628 null
2025-10-30 Tests of exogeneity in duration models with censored data Gilles Crommen et.al. 2510.26613 null
2025-10-30 A DRL-Empowered Multi-Level Jamming Approach for Secure Semantic Communication Weixuan Chen et.al. 2510.26610 null
2025-10-30 Emu3.5: Native Multimodal Models are World Learners Yufeng Cui et.al. 2510.26583 null
2025-10-30 Two-Timescale Optimization Framework for IAB-Enabled Heterogeneous UAV Networks Jikang Deng et.al. 2510.26578 null
2025-10-30 PairUni: Pairwise Training for Unified Multimodal Language Models Jiani Zheng et.al. 2510.25682 null
2025-10-30 Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks Davide Romano et.al. 2510.25623 null
2025-10-29 MetaLore: Learning to Orchestrate Communication and Computation for Metaverse Synchronization Elif Ebru Ohri et.al. 2510.25705 null
2025-10-29 Scaling flow-based approaches for topology sampling in $\mathrm{SU}(3)$ gauge theory Claudio Bonanno et.al. 2510.25704 null
2025-10-29 3-Dimensional Adaptive Unstructured Tessellated Look-up Tables for the Approximation of Compton Form Factors Charles Hyde et.al. 2510.25699 null
2025-10-29 PyDPF: A Python Package for Differentiable Particle Filtering John-Joseph Brady et.al. 2510.25693 null
2025-10-29 Navigation in a Three-Dimensional Urban Flow using Deep Reinforcement Learning Federica Tonti et.al. 2510.25679 null
2025-10-29 ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents Tianyu Yang et.al. 2510.25668 null
2025-10-29 Universal Features of Chiral Symmetry Breaking in Large- $N$ QCD Claudio Bonanno et.al. 2510.25644 null
2025-10-29 Learning to Plan & Schedule with Reinforcement-Learned Bimanual Robot Skills Weikang Wan et.al. 2510.25634 null
2025-10-29 EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis Yusheng Liao et.al. 2510.25628 null
2025-10-29 Inference on Welfare and Value Functionals under Optimal Treatment Assignment Xiaohong Chen et.al. 2510.25607 null
2025-10-29 On the instability of local learning algorithms: Q-learning can fail in infinite state spaces Vittorio Puricelli et.al. 2510.25572 null
2025-10-29 Deep Reinforcement Learning-Based Cooperative Rate Splitting for Satellite-to-Underground Communication Networks Kaiqiang Lin et.al. 2510.25562 null
2025-10-29 Off-policy Reinforcement Learning with Model-based Exploration Augmentation Likun Wang et.al. 2510.25529 null
2025-10-29 Zero Reinforcement Learning Towards General Domains Yuyuan Zeng et.al. 2510.25528 null
2025-10-29 MTIR-SQL: Multi-turn Tool-Integrated Reasoning Reinforcement Learning for Text-to-SQL Zekun Xu et.al. 2510.25510 null
2025-10-29 Dynamic Beamforming and Power Allocation in ISAC via Deep Reinforcement Learning Duc Nguyen Dao et.al. 2510.25496 null
2025-10-29 Reinforcement Learning techniques for the flavor problem in particle physics A. Giarnetti et.al. 2510.25495 null
2025-10-29 Stochastic Control of Dividends with a Drawdown Penalty Kira Dudziak et.al. 2510.25494 null
2025-10-29 Prospects for a 95 GeV Higgs Boson at Future Higgs Factories with Transformer Networks Yabo Dong et.al. 2510.24662 null
2025-10-29 OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Ziyou Hu et.al. 2510.24636 null
2025-10-28 Cluster Dose Prediction in Carbon Ion Therapy: Using Transfer Learning from a Pretrained Dose Prediction U-Net Miriam Schwarze et.al. 2510.24703 null
2025-10-28 Greedy Sampling Is Provably Efficient for RLHF Di Wu et.al. 2510.24700 null
2025-10-28 How Flat is a Plateau? Evolution of Late-Time TDE Disks Yael Alush et.al. 2510.24696 null
2025-10-28 SPICE: Self-Play In Corpus Environments Improves Reasoning Bo Liu et.al. 2510.24684 null
2025-10-28 Fare: Failure Resilience in Learned Visual Navigation Control Zishuo Wang et.al. 2510.24680 null
2025-10-28 Learning to Drive Safely with Hybrid Options Bram De Cooman et.al. 2510.24674 null
2025-10-28 Evolving Diagnostic Agents in a Virtual Clinical Environment Pengcheng Qiu et.al. 2510.24654 null
2025-10-28 Advancing site-specific disease and pest management in precision agriculture: From reasoning-driven foundation models to adaptive, feedback-based learning Nitin Rai et.al. 2510.24650 null
2025-10-28 Fast Bayesian Multilevel Quasi-Monte Carlo Aleksei G. Sorokin et.al. 2510.24604 null
2025-10-28 Low-lying baryon resonances from lattice QCD Colin Morningstar et.al. 2510.24596 null
2025-10-28 Towards Quadrupedal Jumping and Walking for Dynamic Locomotion using Reinforcement Learning Jørgen Anker Olsen et.al. 2510.24584 null
2025-10-28 Dual-Mind World Models: A General Framework for Learning in Dynamic Wireless Networks Lingyi Wang et.al. 2510.24546 null
2025-10-28 Sample-efficient and Scalable Exploration in Continuous-Time RL Klemens Iten et.al. 2510.24482 null
2025-10-28 Adaptive Surrogate Gradients for Sequential Reinforcement Learning in Spiking Neural Networks Korneel Van den Berghe et.al. 2510.24461 null
2025-10-28 Pair Approximation Meets Reality: Diffusion of Innovation in Organizational Networks within the biased-independence q-Voter Model Angelika Abramiuk-Szurlej et.al. 2510.24447 null
2025-10-28 SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space Viktoriia Zinkovich et.al. 2510.24446 null
2025-10-28 Fill in the Blanks: Accelerating Q-Learning with a Handful of Demonstrations in Sparse Reward Settings Seyed Mahdi Basiri Azad et.al. 2510.24432 null
2025-10-28 MiniOneRec: An Open-Source Framework for Scaling Generative Recommendation Xiaoyu Kong et.al. 2510.24431 null
2025-10-28 Multi-Agent Evolve: LLM Self-Improve through Co-evolution Yixing Chen et.al. 2510.23595 null
2025-10-28 VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation Walid Bousselham et.al. 2510.23497 null
2025-10-28 SGFusion: Stochastic Geographic Gradient Fusion in Federated Learning Khoa Nguyen et.al. 2510.23455 null
2025-10-27 Think Twice: Branch-and-Rethink Reasoning Reward Model Yizhu Jiao et.al. 2510.23596 null
2025-10-27 Cosmic magnification on multi-catalogue Herschel submillimetre galaxies R. Fernandez-Fernandez et.al. 2510.23582 null
2025-10-27 Towards Stochastic (N-1)-Secure Redispatch Oleksii Molodchyk et.al. 2510.23551 null
2025-10-27 Variational Thermal State Preparation on Digital Quantum Processors Assisted by Matrix Product States Rui-Hao Li et.al. 2510.23546 null
2025-10-27 Approximately optimal distributed controls for high-dimensional stochastic systems with pairwise interaction through controls Elise Devey et.al. 2510.23537 null
2025-10-27 Sequential Multi-Agent Dynamic Algorithm Configuration Chen Lu et.al. 2510.23535 null
2025-10-27 Learning to Reason Efficiently with Discounted Reinforcement Learning Alex Ayoub et.al. 2510.23486 null
2025-10-27 MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding Xin Jin et.al. 2510.23479 null
2025-10-27 Video-Thinker: Sparking “Thinking with Videos” via Reinforcement Learning Shijian Wang et.al. 2510.23473 null
2025-10-27 Adaptive Multilevel Splitting: First Application to Rare-Event Derivative Pricing Riccardo Gozzo et.al. 2510.23461 null
2025-10-27 Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences Zhuoran Jin et.al. 2510.23451 null
2025-10-27 An Information-Theoretic Analysis of Out-of-Distribution Generalization in Meta-Learning with Applications to Meta-RL Xingtu Liu et.al. 2510.23448 null
2025-10-27 Causal Deep Q Network Elouanes Khelifi et.al. 2510.23424 null
2025-10-27 A Sequential Planning Framework for the Operational Reality of Interacting Air Traffic Flow Regulations and Traffic Flow Programs Thinh Hoang et.al. 2510.23402 null
2025-10-27 VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations Lu Dong et.al. 2510.23397 null
2025-10-27 The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation Farid Bagirov et.al. 2510.23393 null
2025-10-27 Ground-state phase diagram of S = 1/2 Heisenberg model on 2D square-hexagon-octagon lattice Yumeng Luo et.al. 2510.23376 null
2025-10-24 Mechanistic Interpretability for Neural TSP Solvers Reuben Narad et.al. 2510.21693 null
2025-10-24 Reduced Floating-Point Precision Implicit Monte Carlo Simon Butson et.al. 2510.21683 null
2025-10-24 Goal-based portfolio selection with fixed transaction costs Erhan Bayraktar et.al. 2510.21650 null
2025-10-24 Electroweak corrections to $gg\rightarrow γγ$ Gabriele Fiore et.al. 2510.21643 null
2025-10-24 Predicted observational effects of rapid rotation for Be stars Rina G. Rast et.al. 2510.21640 null
2025-10-24 DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection Tala Aljaafari et.al. 2510.21638 null
2025-10-24 DeepAgent: A General Reasoning Agent with Scalable Toolsets Xiaoxi Li et.al. 2510.21618 null
2025-10-24 Enhancing Tactile-based Reinforcement Learning for Robotic Control Elle Miller et.al. 2510.21609 null
2025-10-24 Multilevel Picard scheme for solving high-dimensional drift control problems with state constraints Yuan Zhong et.al. 2510.21607 null
2025-10-24 RETuning: Upgrading Inference-Time Scaling for Stock Movement Prediction with Large Language Models Xueyuan Lin et.al. 2510.21604 null
2025-10-24 Three-nucleon lepton-number-violating potentials in chiral EFT and their matrix elements in light nuclei Graham Chambers-Wall et.al. 2510.21564 null
2025-10-24 System-Theoretic Analysis of Dynamic Generalized Nash Equilibrium Problems – Turnpikes and Dissipativity Sophie Hall et.al. 2510.21556 null
2025-10-24 Cost Minimization for Space-Air-Ground Integrated Multi-Access Edge Computing Systems Weihong Qin et.al. 2510.21541 null
2025-10-24 A Unified Model for Multi-Task Drone Routing in Post-Disaster Road Assessment Huatian Gong et.al. 2510.21525 null
2025-10-24 Surrogate-based quantification of policy uncertainty in generative flow networks Ramón Nartallo-Kaluarachchi et.al. 2510.21523 null
2025-10-24 The population of Galactic young massive star clusters in the TeV range Rowan Batzofin et.al. 2510.21480 null
2025-10-24 MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward Optimization Chenglong Wang et.al. 2510.21473 null
2025-10-24 Constraints on ultra-heavy dark matter from the CDEX-10 experiment at the China Jinping Underground Laboratory Y. F. Wang et.al. 2510.21458 null
2025-10-24 Unified token representations for sequential decision models Zhuojing Tian et.al. 2510.21448 null
2025-10-24 Causality Meets Locality: Provably Generalizable and Scalable Policy Learning for Networked Systems Hao Liang et.al. 2510.21427 null
2025-10-24 Real-Time Gait Adaptation for Quadrupeds using Model Predictive Control and Reinforcement Learning Prakrut Kotecha et.al. 2510.20706 null
2025-10-23 KL-Regularized Reinforcement Learning is Designed to Mode Collapse Anthony GX-Chen et.al. 2510.20817 null
2025-10-23 GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation Guangqi Jiang et.al. 2510.20813 null
2025-10-23 A Microphysical Probe of Neutron Star Interiors: Constraining the Equation of State with Glitch Dynamics Zhonghao Tu et.al. 2510.20791 null
2025-10-23 Consumption-Investment Problem in Rank-Based Models David Itkin et.al. 2510.20763 null
2025-10-23 Reinforcement Learning and Consumption-Savings Behavior Brandon Kaplowitz et.al. 2510.20748 null
2025-10-23 No-Regret Thompson Sampling for Finite-Horizon Markov Decision Processes with Gaussian Processes Jasmine Bayrooti et.al. 2510.20725 null
2025-10-23 Measuring cosmic dipole with the GRB luminosity-time relation Jessica Santiago et.al. 2510.20705 null
2025-10-23 Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge Graphs Yanlin Song et.al. 2510.20691 null
2025-10-23 Downsizing Diffusion Models for Cardinality Estimation Xinhe Mu et.al. 2510.20681 null
2025-10-23 The Shape of Reasoning: Topological Analysis of Reasoning Traces in Large Language Models Xue Wen Tan et.al. 2510.20665 null
2025-10-23 Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence Jiahao Meng et.al. 2510.20579 null
2025-10-23 EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence Ding Zou et.al. 2510.20578 null
2025-10-23 Monte Carlo Sampling for Wave Functions Requiring (Anti)Symmetrization Koyena Bose et.al. 2510.20577 null
2025-10-23 AdaDoS: Adaptive DoS Attack via Deep Adversarial Reinforcement Learning in SDN Wei Shao et.al. 2510.20566 null
2025-10-23 GlobalRAG: Enhancing Global Reasoning in Multi-hop Question Answering via Reinforcement Learning Jinchang Luo et.al. 2510.20548 null
2025-10-23 A Unified Framework for Zero-Shot Reinforcement Learning Jacopo Di Ventura et.al. 2510.20542 null
2025-10-23 Detection of ultra-high-energy cosmic rays in the southern hemisphere with FAST: data acquisition and preliminary results Jakub Kmec et.al. 2510.20522 null
2025-10-23 Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence Kun Ouyang et.al. 2510.20470 null
2025-10-23 On Multiple Robustness of Proximal Dynamic Treatment Regimes Yuanshan Gao et.al. 2510.20451 null
2025-10-23 DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning Runpeng Xie et.al. 2510.19562 null
2025-10-22 olmOCR 2: Unit Test Rewards for Document OCR Jake Poznanski et.al. 2510.19817 null
2025-10-22 Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing Yusu Qian et.al. 2510.19808 null
2025-10-22 Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning Xichen Zhang et.al. 2510.19807 null
2025-10-22 SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration Xichen Zhang et.al. 2510.19767 null
2025-10-22 SEA: Semantic Map Prediction for Active Exploration of Uncertain Areas Hongyu Ding et.al. 2510.19766 null
2025-10-22 Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning Gunshi Gupta et.al. 2510.19732 null
2025-10-22 Semi-Implicit Approaches for Large-Scale Bayesian Spatial Interpolation Sébastien Garneau et.al. 2510.19722 null
2025-10-22 MedReason-R1: Learning to Reason for CT Diagnosis with Reinforcement Learning and Local Zoom Yifan Li et.al. 2510.19626 null
2025-10-22 Demonstrating Real Advantage of Machine-Learning-Enhanced Monte Carlo for Combinatorial Optimization Luca Maria Del Bono et.al. 2510.19544 null
2025-10-22 Quantum Monte Carlo study of low-dimensional Fermi fluids of dipolar atoms Clio Johnson et.al. 2510.19533 null
2025-10-22 The Confusing Instance Principle for Online Linear Quadratic Control Waris Radji et.al. 2510.19531 null
2025-10-22 Optimizing the Unknown: Black Box Bayesian Optimization with Energy-Based Model and Reinforcement Learning Ruiyao Miao et.al. 2510.19530 null
2025-10-22 Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach Sebastian Reboul et.al. 2510.19528 null
2025-10-22 Practical algorithm for simulating thermal pure quantum states Wei-Bo He et.al. 2510.19504 null
2025-10-22 Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning Kevin Huang et.al. 2510.19495 null
2025-10-22 Quantum Machine Learning methods for Fourier-based distribution estimation with application in option pricing Fernando Alonso et.al. 2510.19494 null
2025-10-22 Monte Carlo study of the $O(2)$-invariant $φ^4$ theory with a cubic perturbation in three dimensions Martin Hasenbusch et.al. 2510.19473 null
2025-10-22 Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis Xueqi Ma et.al. 2510.19451 null
2025-10-22 Universal Quantitative Abstraction: Categorical Duality and Logical Completeness for Probabilistic Systems Nivar Anwer et.al. 2510.19444 null
2025-10-21 Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting Howard Chen et.al. 2510.18874 null
2025-10-21 EffiReasonTrans: RL-Optimized Reasoning for Code Translation Yanlin Wang et.al. 2510.18863 null
2025-10-21 Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model Ling Team et.al. 2510.18855 null
2025-10-21 Lyapunov-Aware Quantum-Inspired Reinforcement Learning for Continuous-Time Vehicle Control: A Feasibility Study Nutkritta Kraipatthanapong et.al. 2510.18852 null
2025-10-21 Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning Chenghao Zhu et.al. 2510.18849 null
2025-10-21 MADR: MPC-guided Adversarial DeepReach Ryan Teoh et.al. 2510.18845 null
2025-10-21 PCMS: Parallel Coupler For Multimodel Simulations Jacob S. Merson et.al. 2510.18838 null
2025-10-21 Actor-Free Continuous Control via Structurally Maximizable Q-Functions Yigit Korkmaz et.al. 2510.18828 null
2025-10-21 Search Self-play: Pushing the Frontier of Agent Capability without Supervision Hongliang Lu et.al. 2510.18821 null
2025-10-21 Online SFT for LLM Reasoning: Surprising Effectiveness of Self-Tuning without Rewards Mengqi Li et.al. 2510.18814 null
2025-10-21 Computational Foundations for Strategic Coopetition: Formalizing Interdependence and Complementarity Vik Pant et.al. 2510.18802 null
2025-10-21 Two-loop QCD corrections for real and off-shell diphoton and triphoton production via quark loops Dario Kermanschah et.al. 2510.18801 null
2025-10-21 WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection Guanzhong He et.al. 2510.18798 null
2025-10-21 Beware of the running $n_s$ when producing heavy primordial black holes Sasha Allegrini et.al. 2510.18791 null
2025-10-21 Analysis note: measurement of thrust and track energy-energy correlator in $e^+e^-$ collisions at 91.2 GeV with DELPHI open data Jingyu Zhang et.al. 2510.18762 null
2025-10-21 Verifiable Accuracy and Abstention Rewards in Curriculum RL to Alleviate Lost-in-Conversation Ming Li et.al. 2510.18731 null
2025-10-21 Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options Joongkyu Lee et.al. 2510.18713 null
2025-10-21 Chemistry, Climate, and Transmission Spectra of TRAPPIST-1 e Explored with a Multimodel Sparse Sampled Ensemble Eric T. Wolf et.al. 2510.18704 null
2025-10-21 Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach Chenbei Lu et.al. 2510.18687 null
2025-10-21 Sherlock Your Queries: Learning to Ask the Right Questions for Dialogue-Based Retrieval Dong Yun et.al. 2510.18659 null
2025-10-21 An integrated neural wavefunction solver for spinful Fermi systems Alexander Avdoshkin et.al. 2510.18621 null
2025-10-21 CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent Haojia Lin et.al. 2510.18596 null
2025-10-21 Deep Q-Learning Assisted Bandwidth Reservation for Multi-Operator Time-Sensitive Vehicular Networking Abdullah Al-Khatib et.al. 2510.18553 null
2025-10-21 Improved thermonuclear rate of $^{42}$Ti($p$,$γ$)$^{43}$ V and its astrophysical implication in rp-process S. Q. Hou et.al. 2510.18531 null
2025-10-21 Efficient Model-Based Reinforcement Learning for Robot Control via Online Learning Fang Nan et.al. 2510.18518 null
2025-10-21 Socialized Learning and Emergent Behaviors in Multi-Agent Systems based on Multimodal Large Language Models Sureyya Akin et.al. 2510.18515 null
2025-10-21 Learning to Navigate Under Imperfect Perception: Conformalised Segmentation for Safe Reinforcement Learning Daniel Bethell et.al. 2510.18485 null
2025-10-21 Safe But Not Sorry: Reducing Over-Conservatism in Safety Critics via Uncertainty-Aware Modulation Daniel Bethell et.al. 2510.18478 null
2025-10-21 CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment Xue Jiang et.al. 2510.18471 null
2025-10-21 Uncovering critical temperature dependence in Heusler magnets via explicit machine learning Jean-Baptiste Morée et.al. 2510.18469 null
2025-10-21 DeLoad: Demand-Driven Short-Video Preloading with Scalable Watch-Time Estimation Tong Liu et.al. 2510.18459 null
2025-10-21 Fingerprints of cluster-based Haldane and bound-magnon states in a spin-1 Heisenberg diamond chain Azam Zoshki et.al. 2510.18447 null
2025-10-21 PlanU: Large Language Model Decision Making through Planning under Uncertainty Ziwei Deng et.al. 2510.18442 null
2025-10-21 Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents Guangfu Guo et.al. 2510.18424 null
2025-10-21 On AI Verification in Open RAN Rahul Soundrarajan et.al. 2510.18417 null
2025-10-21 MENTOR: A Reinforcement Learning Framework for Model Enhancement via Teacher-Optimized Rewards in Small Models ChangSu Choi et.al. 2510.18383 null
2025-10-21 Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback Yi-Lun Wu et.al. 2510.18353 null
2025-10-21 PGTT: Phase-Guided Terrain Traversal for Perceptive Legged Locomotion Alexandros Ntagkas et.al. 2510.18348 null
2025-10-21 Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs Jongmin Lee et.al. 2510.18340 null
2025-10-21 The implications of inflation for the last ACT Zhi-Chong Qiu et.al. 2510.18320 null
2025-10-21 MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation Chengshu Li et.al. 2510.18316 null
2025-10-21 Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task Brady Bhalla et.al. 2510.18315 null
2025-10-21 Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models Lehan Wang et.al. 2510.18303 null
2025-10-21 Food4All: A Multi-Agent Framework for Real-time Free Food Discovery with Integrated Nutritional Metadata Zhengqing Yuan et.al. 2510.18289 null
2025-10-21 From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation Ziwei Huang et.al. 2510.18263 null
2025-10-21 NTKMTL: Mitigating Task Imbalance in Multi-Task Learning from Neural Tangent Kernel Perspective Xiaohan Qin et.al. 2510.18258 null
2025-10-21 The Picard-Lagrange Framework for Higher-Order Langevin Monte Carlo Jaideep Mahajan et.al. 2510.18242 null
2025-10-21 Nash Policy Gradient: A Policy Gradient Method with Iteratively Refined Regularization for Finding Nash Equilibria Eason Yu et.al. 2510.18183 null
2025-10-20 Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains Soumya Rani Samineni et.al. 2510.18176 null
2025-10-20 LLMs Encode How Difficult Problems Are William Lugoloobi et.al. 2510.18147 null
2025-10-20 Measuring Reasoning in LLMs: a New Dialectical Angle Soheil Abbasloo et.al. 2510.18134 null
2025-10-20 R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations Connor Mattson et.al. 2510.18085 null
2025-10-20 RL-Driven Security-Aware Resource Allocation Framework for UAV-Assisted O-RAN Zaineh Abughazzah et.al. 2510.18084 null
2025-10-20 Provably Optimal Reinforcement Learning under Safety Filtering Donggeon David Oh et.al. 2510.18082 null
2025-10-20 R2L: Reliable Reinforcement Learning: Guaranteed Return & Reliable Policies in Reinforcement Learning Nadir Farhi et.al. 2510.18074 null
2025-10-20 Fine-tuning Flow Matching Generative Models with Intermediate Feedback Jiajun Fan et.al. 2510.18072 null
2025-10-20 Oxidation State Dynamics and Emerging Patterns in Magnetite Emre Gürsoy et.al. 2510.18061 null
2025-10-20 SPACeR: Self-Play Anchoring with Centralized Reference Models Wei-Jer Chang et.al. 2510.18060 null
2025-10-20 Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models Jiajun Fan et.al. 2510.18053 null
2025-10-20 OPTAGENT: Optimizing Multi-Agent LLM Interactions Through Verbal Reinforcement Learning for Enhanced Reasoning Zhenyu Bi et.al. 2510.18032 null
2025-10-20 Humanoid Goalkeeper: Learning from Position Conditioned Task-Motion Constraints Junli Ren et.al. 2510.18002 null
2025-10-20 Collider Searches for Near-Continuum Dark Matter Steven Ferrante et.al. 2510.17989 null
2025-10-20 Accelerating Bayesian Inference via Multi-Fidelity Transport Map Coupling Sanjan C. Muchandimath et.al. 2510.17946 null
2025-10-20 An Exact Quantile-Energy Equality for Terminal Halfspaces in Linear-Gaussian Control with a Discrete-Time Companion, KL/Schrodinger Links, and High-Precision Validation Sandro Andric et.al. 2510.17945 null
2025-10-20 UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts Fu-Yun Wang et.al. 2510.17937 null
2025-10-20 EvoSyn: Generalizable Evolutionary Data Synthesis for Verifiable Learning He Du et.al. 2510.17928 null
2025-10-20 Rewarding the Journey, Not Just the Destination: A Composite Path and Answer Self-Scoring Reward Mechanism for Test-Time Reinforcement Learning Chenwei Tang et.al. 2510.17923 null
2025-10-20 CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections Keuntae Kim et.al. 2510.17921 null
2025-10-20 Functional Distribution Networks (FDN) Omer Haq et.al. 2510.17794 null
2025-10-20 Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains Austin Xu et.al. 2510.17793 null
2025-10-20 SoftMimic: Learning Compliant Whole-body Control from Examples Gabriel B. Margolis et.al. 2510.17792 null
2025-10-20 UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action Yuhao Yang et.al. 2510.17790 null
2025-10-20 B-Meson Anomalies: Effective Field Theory Meets Machine Learning Alejandro Mir et.al. 2510.17742 null
2025-10-20 Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations Tong Chen et.al. 2510.17733 null
2025-10-20 QueST: Incentivizing LLMs to Generate Difficult Problems Hanxu Hu et.al. 2510.17715 null
2025-10-20 The Marked Edge Walk: A Novel MCMC Algorithm for Sampling of Graph Partitions Atticus McWhorter et.al. 2510.17714 null
2025-10-20 A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning Anjie Liu et.al. 2510.17697 null
2025-10-20 Efficient Algorithms for Mitigating Uncertainty and Risk in Reinforcement Learning Xihong Su et.al. 2510.17690 null
2025-10-20 CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks Xu Zhang et.al. 2510.17687 null
2025-10-20 RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation Yuquan Xue et.al. 2510.17640 null
2025-10-20 Colour coherence in small collision systems Isobel Kolbé et.al. 2510.17570 null
2025-10-20 An Empirical Study of Lagrangian Methods in Safe Reinforcement Learning Lindsay Spoor et.al. 2510.17564 null
2025-10-20 Towards Optimal Control and Algorithmic Structure of Decompression Schedules Benjamin Marsh et.al. 2510.17551 null
2025-10-20 OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction Raghu Vamshi Hemadri et.al. 2510.17532 null
2025-10-20 Plasma Shape Control via Zero-shot Generative Reinforcement Learning Niannian Wu et.al. 2510.17531 null
2025-10-20 Toward Autonomous Neural VMC: An Energy-Variance Convergence Criterion for Quantum Systems Huan-Chen Shi et.al. 2510.17490 null
2025-10-20 Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs Paula Cordero-Encinar et.al. 2510.17472 null
2025-10-20 Estimating Orbital Parameters of Direct Imaging Exoplanet Using Neural Network Bo Liang et.al. 2510.17459 null
2025-10-20 Agentic Reinforcement Learning for Search is Unsafe Yushi Yang et.al. 2510.17431 null
2025-10-20 Leveraging Group Relative Policy Optimization to Advance Large Language Models in Traditional Chinese Medicine Jiacheng Xie et.al. 2510.17402 null
2025-10-20 Finite-Time Bounds for Average-Reward Fitted Q-Iteration Jongmin Lee et.al. 2510.17391 null
2025-10-20 Inference of Deterministic Finite Automata via Q-Learning Elaheh Hosseinkhani et.al. 2510.17386 null
2025-10-20 TabR1: Taming GRPO for tabular reasoning LLMs Pengxiang Cai et.al. 2510.17385 null
2025-10-20 Optimizing Energy Management of Smart Grid using Reinforcement Learning aided by Surrogate models built using Physics-informed Neural Networks Julen Cestero et.al. 2510.17380 null
2025-10-20 When 5G NTN Meets GNSS: Tracking GNSS Signals under Overlaid 5G Waveforms Idir Edjekouane et.al. 2510.17324 null
2025-10-20 Auto-Rubric: Learning to Extract Generalizable Criteria for Reward Modeling Lipeng Xie et.al. 2510.17314 null
2025-10-20 Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks Xinkai Wang et.al. 2510.17277 null
2025-10-20 Characterizing expansivity through $C^*$ -algebras S. Bautista et.al. 2510.17255 null
2025-10-20 From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models Zefan Cai et.al. 2510.17247 null
2025-10-20 Deep Neural Network extraction of Unpolarized Transverse Momentum Distributions I. P. Fernando et.al. 2510.17243 null
2025-10-20 Coinvisor: An RL-Enhanced Chatbot Agent for Interactive Cryptocurrency Investment Analysis Chong Chen et.al. 2510.17235 null
2025-10-20 D2C-HRHR: Discrete Actions with Double Distributional Critics for High-Risk-High-Return Tasks Jundong Zhang et.al. 2510.17212 null
2025-10-20 Trading with the Devil: Risk and Return in Foundation Model Strategies Jinrui Zhang et.al. 2510.17165 null
2025-10-20 ALPINE: A Lightweight and Adaptive Privacy-Decision Agent Framework for Dynamic Edge Crowdsensing Guanjie Cheng et.al. 2510.17162 null
2025-10-20 GACO-CAD: Geometry-Augmented and Conciseness-Optimized CAD Model Generation from Single Image Yinghui Wang et.al. 2510.17157 null
2025-10-20 Decentralized Real-Time Planning for Multi-UAV Cooperative Manipulation via Imitation Learning Shantnav Agarwal et.al. 2510.17143 null
2025-10-20 Rethinking On-policy Optimization for Query Augmentation Zhichao Xu et.al. 2510.17139 null
2025-10-20 Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control Chengxiu Hua et.al. 2510.17122 null
2025-10-20 Learning to Design Soft Hands using Reward Models Xueqian Bai et.al. 2510.17086 null
2025-10-20 Consistent Zero-Shot Imitation with Contrastive Goal Inference Kathryn Wantlin et.al. 2510.17059 null
2025-05-13 ROLeR: Effective Reward Shaping in Offline Reinforcement Learning for Recommender Systems Yi Zhang et.al. 2407.13163 null
2025-01-14 Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement Learning Lanqing Li et.al. 2402.02429 null
2024-12-10 Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone Max Sobol Mark et.al. 2412.06685 null
2024-10-30 Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement Learning Qi Wang et.al. 2305.15260 null
2024-10-03 From Reward Shaping to Q-Shaping: Achieving Unbiased Learning with LLM-Guided Knowledge Xiefeng Wu et.al. 2410.01458 null
2024-05-30 Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RL Yu Luo et.al. 2405.18520 null
2023-06-22 Reward Shaping via Diffusion Process in Reinforcement Learning Peeyush Kumar et.al. 2306.11885 null
2023-02-14 Review of Deep Reinforcement Learning for Autonomous Driving B. Udugama et.al. 2302.06370 null
2023-01-04 Transformer in Transformer as Backbone for Deep Reinforcement Learning Hangyu Mao et.al. 2212.14538 null
2023-01-02 Offline Policy Optimization in RL with Variance Regularizaton Riashat Islam et.al. 2212.14405 null
2022-11-29 Domain Generalization for Robust Model-Based Offline Reinforcement Learning Alan Clark et.al. 2211.14827 null
2022-11-22 Model-based Trajectory Stitching for Improved Offline Reinforcement Learning Charles A. Hepburn et.al. 2211.11603 null
2022-11-09 State Advantage Weighting for Offline RL Jiafei Lyu et.al. 2210.04251 null
2022-11-07 Contrastive Value Learning: Implicit Models for Simple Offline RL Bogdan Mazoure et.al. 2211.02100 null
2021-11-30 Improving Zero-shot Generalization in Offline Reinforcement Learning using Generalized Similarity Functions Bogdan Mazoure et.al. 2111.14629 null
2021-09-24 A Workflow for Offline Model-Free Robotic Reinforcement Learning Aviral Kumar et.al. 2109.10813 null
2020-11-25 An Optimistic Perspective on Offline Reinforcement Learning Rishabh Agarwal et.al. 1907.04543 null
2020-08-20 Conservative Q-Learning for Offline Reinforcement Learning Aviral Kumar et.al. 2006.04779 null
2020-05-19 On-Policy Robot Imitation Learning from a Converging Supervisor Ashwin Balakrishna et.al. 1907.03423 null
2019-11-14 Accelerating Training in Pommerman with Imitation and Reinforcement Learning Hardik Meisheri et.al. 1911.04947 null
2018-11-20 Modelling the Dynamic Joint Policy of Teammates with Attention Multi-agent DDPG Hangyu Mao et.al. 1811.07029 null

Notes:

Function added: