관심도 점수
100
- 검색 slug
- ai-coding-agents
- 실제 연결
- Hacker News, GitHub, arXiv
- 키워드
- and, the, coding, agent, agents, for, code, ai
- 마지막 수집
- 2026-09-21T09:00:31.307Z
출처별 탭
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring procedures for recu...
MintAct: A Unified Visual Agent for Digital Environments
We present MintAct, a family of vision-language models that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use, trained at 2B, 4B, and 8B scales. Through careful design of our environments, data, and training recipes, MintAct models match the performance of per-domain specialists across all of these capabilities. To enable this, we develop a scalable environment and reinforcement learning (RL) infrastructure. On the environment side, we host hundreds of concurrent inst...
Cross-sector generalization of accident-process role classification in occupational accident narratives
Occupational accident narratives contain valuable information about work situations, unfavourable conditions, accident events, and their consequences. Automatically structuring these narratives can facilitate large-scale accident analysis and support occupational risk prevention. However, the terminology and writing styles used to describe accidents vary considerably across sectors and organisations, raising questions about the ability of automated coding systems to generalize beyond their training domain. In this...
Rotating Neutron Star Migrations as a Standardized Test for 3+1 Numerical Relativity
When simulating dynamically unstable neutron stars, truncation errors can introduce perturbations that drive the stellar configuration either toward gravitational collapse into a black hole or toward migration to a dynamically stable, lower-density configuration. This mechanism has become a standard benchmark for validating general-relativistic hydrodynamics codes in the case of spherically symmetric neutron-star models. In this work, we extend the analysis to the more general case of uniformly rotating neutron st...
APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans against a live payment agent during a public capture-the-flag event, across 14 models from 8 labs, five policy configurations and two replay tracks, with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification. 225,964 evaluations completed. We report five distinct events per evaluation, because collapsing them is how an agent benchmark produces a nu...
The AGN Channel in 3D: Scattering Belts and the Importance of Eccentricity in the Black Hole Population
Active galactic nuclei (AGN) are a promising origin for observed gravitational wave mergers. Current population synthesis models are limited to 1D N-body or Monte Carlo methods which rely on statistical approaches to resolving dynamical scatterings. We present three-dimensional hybrid $N$-body simulations of a population of black holes (BHs) surrounding an AGN using a new code in development, AGNBI, where close interactions are directly simulated for both single and binary objects. Our results show binary formatio...
A Sociotechnical Review of Algorithms in Health Systems: Technical, Cost, and Human-Centered Considerations
Artificial intelligence (AI) applications in healthcare are becoming increasingly prevalent, to assist health systems, providers, and patients with tasks such as decision-making, risk prediction, and diagnosis. This increasing computational potential brings AI applications to the forefront of workplace decision making, often without full consideration of subsequent computational, organizational, and social costs. These applications are leveraged to reduce healthcare costs and increase efficiency of daily tasks, wi...
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific...
Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw
Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent aspect, value fulfillment, and user outcome. The 21 values form six value groups, including Autonomous, Dependable, and Affordable Operation, Bounded Reach, Reviewability, and Equitable Access. Relative to each aspect's...
Benchmarking World Models for Continual Learning on Compositional Tasks
A desirable property of a world model is the ability to learn continually across tasks, adapting to new environments without forgetting what the agent has already learnt. In particular, the ability to retain and reuse knowledge obtained from prior experiences underpins an agent's ability to efficiently adapt to novel environments, as the dynamics of the physical world can often be described in recurring mechanisms. However, the world model's measure of adaptation entangles two abilities: the speed and capacity to...
DietrichGebert/ponytail
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
JuliusBrussee/caveman
🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
addyosmani/agent-skills
Production-grade engineering skills for AI coding agents.
headroomlabs-ai/headroom
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
shanraisshan/claude-code-best-practice
from vibe coding to agentic engineering - practice makes claude perfect
obra/superpowers
An agentic skills framework & software development methodology that works.
thedotmack/claude-mem
Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
earendil-works/pi
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
ruvnet/ruflo
🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
x1xhlol/system-prompts-and-models-of-ai-tools
FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI, VSCode Agent, Warp.dev, Windsurf, Xcode, Z.ai Code, Dia & v0. (And other Open Sourced) System Prompts, Internal Tools & AI Models