Computing Actual Causes for Neural Network Predictions under Structured Causal Inputs Explaining the predictions of neural networks is a central challenge in trustworthy AI. Existing explanation methods, such as those based on feature attribution or minimal sufficient sets, typically t arxiv · 11小时前
MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Wher arxiv · 12小时前
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provid arxiv · 12小时前
Risky Business: Measuring The Faithfulness-Safety Tension Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output strictly derives from its reasoning trace. We ident arxiv · 12小时前
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards bu arxiv · 12小时前
Can LLMs Test Terminal User Interfaces? Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in developer tools. Yet they lack a dedicated testing methodology. W arxiv · 12小时前
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability. Artificial i arxiv · 12小时前
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks. The security implications of this coo arxiv · 12小时前
AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with no 2D analogue, and post-route engineering chan arxiv · 12小时前
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are cos arxiv · 12小时前
CARE-Bench: Benchmarking Patient-Facing LLM Triage Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We introduce CARE-Bench, arxiv · 12小时前
GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language models (LLMs), treating the model itself as the knowle arxiv · 12小时前
SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and returns machine-consumable artifacts under communication and pow arxiv · 12小时前
When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy ca arxiv · 12小时前
Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction receive bids. This over-distribution weakens auct arxiv · 12小时前