관심도 점수
100
- 검색 slug
- local-llm-tooling
- 실제 연결
- Hacker News, GitHub, arXiv
- 키워드
- and, the, local, for, with, llm, agent, tooling
- 마지막 수집
- 2026-09-21T09:54:05.248Z
출처별 탭
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring procedures for recu...
Multi-Boundary Spinning $AdS_5$ Black Holes, Strongly Coupled Plasma Balls and a dual locally de Sitter spacetime
We construct a new family of exact spinning black hole solutions in five-dimensional General Relativity with a negative cosmological constant, which are dual to strongly coupled spinning plasma balls on four-dimensional Minkowski spacetime. In analogy with the recent discovery of such black solutions in four bulk dimensions, the boundary Lorentz factor acts as a new coordinate in the bulk, giving rise to a new UV region that, in our case, is locally a three-dimensional de Sitter spacetime times the real line. The...
MintAct: A Unified Visual Agent for Digital Environments
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent inst...
APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans against a live payment agent during a public capture-the-flag event, across 14 models from 8 labs, five policy configurations and two replay tracks, with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification. 225,964 evaluations completed. We report five distinct events per evaluation, because collapsing them is how an agent benchmark produces a nu...
Duty Factor Predicts Robust Constrained Quadrupedal Locomotion Across Gait Types
Quadrupedal robots are increasingly deployed in environments where locomotion must remain robust to disturbances and constrained terrain. Gait type, such as walking or trotting, is commonly used to characterize quadrupedal locomotion. However, gait type does not uniquely define locomotion, as parameters such as duty factor, speed, and stance width can vary within a single gait type. In this work, we investigate the relationship between these gait parameters using three distinct quadrupedal locomotion control appro...
Discovery, Characterization, and Potential Origins of a Stream in the Stellar Halo of Nearby LMC-Mass Galaxy NGC 55
We present a previously undetected stellar stream in the halo of the LMC-mass dwarf galaxy NGC 55 (2 Mpc), as part of an ongoing effort to characterize the stellar halos of LMC/SMC-mass dwarfs in the DEEP component of the DECam Local Volume Exploration (DELVE) survey. This structure is aligned with an outflow traced by H$α$ emission, a spur and cloud of neutral hydrogen, and the ultra-diffuse and possibly disrupting satellite galaxy NGC 55-dw1. We investigate possible ex-situ progenitor scenarios for this stream t...
Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent aspect, value fulfillment, and user outcome. The 21 values form six value groups, including Autonomous, Dependable, and Affordable Operation, Bounded Reach, Reviewability, and Equitable Access. Relative to each aspect's...
Generalized Hamiltonian formalism for spatially nonlocal nonlinear differential equations
In this work, we develop a generalized Hamiltonian formalism for spatially nonlocal field theories whose Lagrangian densities depend explicitly on both the local field and its spatially reflected counterpart. Starting from a generalized variational principle, we derive generalized Euler-Lagrange equations and introduce a generalized functional derivative that consistently accounts for reflected-field contributions. The proposed formalism is applied to three spatially nonlocal nonlinear Schrödinger equations. For t...
Traffic Sign Recognition for Autonomous Driving Using Branched YOLOv2 and Geometric Features
Traffic sign recognition (TSR) is an important perception task for autonomous driving and advanced driver-assistance systems, where a system must both localize traffic signs and determine their semantic classes efficiently. This work presents a TSR system based on YOLOv2 for simultaneous detection and classification. Two complementary modifications are studied. First, YOLOv2 is extended with intermediate prediction layers, forming a branched architecture that can terminate inference early for easy cases and reduce...
Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention
Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations. We prove two results formalizing this structure. First (CWAR Reducibility): confident-failure reduction is achievable if and only if retrieval features carry mutual information about success, a condition satisfied by LLM-judge pipelines but substantially weaker in dense-only settings, explaining the AUC-AC gap between regimes. Second (Feature Regime Complementarity): no single ANN sco...