Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-spec arxiv · 2天前
Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision Complex robotic manipulation tasks frequently require a long-term memory of past events and actions. As conditioning on full histories renders policies prone to spurious correlations and degrades perf arxiv · 2天前
FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a arxiv · 2天前
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Professional design requires any-color control: the ability to specify an object's target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, ed arxiv · 2天前
ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a broad spectrum of extracolonic tumors. Early identification and accurate classification of these arxiv · 2天前
Quantifying Overclaiming Propensity in Frontier LLM Agents Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that work a user sees. We quantify the propensity of f arxiv · 2天前
An Empirical Study of Harness Design for Coding Agents Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic syste arxiv · 2天前
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from arxiv · 2天前
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically in arxiv · 2天前
GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different ta arxiv · 2天前
Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights Generative agents are increasingly used to select and narrate video highlights, but they typically operate over unstructured or frame-level representations. Their output is consequently difficult for arxiv · 2天前
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as benchmark task types or conversation types in deployed agents. Exhaustive testing is expensive, arxiv · 2天前
RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents Effective troubleshooting agents in enterprise customer support depend on retrieving actionable guidance from similar historical cases, yet existing retrieval-augmented generation (RAG) systems treat arxiv · 2天前
Large Language Models as Falsifiers for Cyber-Physical Systems Falsification searches for counterexamples to formal specifications in cyber-physical systems (CPS). With specifications written in Signal Temporal Logic (STL), falsification can be formulated as a ro arxiv · 2天前
Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure Semantic cell annotation improves chunking interpretability for spreadsheets in LLM-driven RAG systems, aiding answer generation through enriched context rather than improved retrieval accuracy. We pr arxiv · 2天前