🤖 AI Agent 研究Research
当长期客服代表执行失败时,结果级评估会显示不成功的结果,但不会显示决定性错误进入轨迹的位置。然后,开发人员必须检查完整的执行情况,以确定责任角色,并本地化最早的决定性根本原因步骤。
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step.
即使代理在有状态运行时运行,代理基准通常也仅评估最终答案。我们认为这不足以说明正在评估的内容:适当的单元是声明的模型加运行时配置,其故障可能发生在证据获取、运行时路由、安全边界或重复执行中。
Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: the proper unit is a declared model-plus-runtime configuration whose failures can occur in evidence acquisition, runtime routing, safety boundaries, or repeated execution.
语言模型是顺序处理器,但长期代理需要超出模型权重和活动上下文之外的外部信息和计算。Prime Agent是一款适用于长期评估和编码代理工作流程的开源工具。
Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows.
通用语言模型可以推理和综合知识,但复杂的工作还需要与文件、信息源和可执行代码的持续交互,以及状态维护、故障恢复和可验证的交付。我们称之为工作能力:在实现现实世界目标方面取得持续、可验证的进展。
General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this working capability: sustained, verifiable progress toward a real-world objective.
⭐ GitHub 热门项目GitHub Trending
人工智能编码代理的反幻觉护栏-混合技能+ MCP服务器( 6个工具) ,在代码被标记为完成之前对其进行验证。适用于Claude Code、Cursor、VS Code、Copilot。
Anti-hallucination guardrails for AI coding agents - hybrid Skill + MCP server (6 tools) that verifies code before it is marked done. Works with Claude Code, Cursor, VS Code, Copilot.
存储库级AI安全审计、漏洞挖掘和自动修复代理。
Repository-level AI security auditing, vulnerability mining and automated repair agent.
适用于Mac、Linux、Windows和Raspberry Pi的开源计算机AI代理。运行命令、文件、Web和cron作业—使用您控制的本地策略实现自主AI自动化。自托管大脑( Ollama/BYO键)或使用云。451次测试,包括红队。IEC 62443和EU CRA一致。
Open-source computer-use AI agent for Mac, Linux, Windows & Raspberry Pi. Run commands, files, web and cron jobs — autonomous AI automation with a local policy you control. Self-host the brain (Ollama/BYO key) or use the cloud. 451 tests incl. red-team. IEC 62443 & EU CRA aligned.
Kineti是Rust本地构建的超高速代理线束。唯一可数学验证的代理线束。
Kineti is an ultra-fast agent harness natively built in Rust. The only mathematically verifiable agent harness.
Lightweight browser-side AI agent framework for legacy enterprise web apps—Java 8, Spring Boot 2, TypeScript, MCP/WebMCP, OpenAI-compatible LLMs, tool calling, streaming, audit, and persistence. 面向存量企业 Web 系统的轻量级 Agent 框架。
Lightweight browser-side AI agent framework for legacy enterprise web apps—Java 8, Spring Boot 2, TypeScript, MCP/WebMCP, OpenAI-compatible LLMs, tool calling, streaming, audit, and persistence. 面向存量企业 Web 系统的轻量级 Agent 框架。
🚀 模型与行业动态Models & Industry
$ 2亿的扩展是在这家物理AI初创公司达到$ 20亿估值后几个月。
The $200 million extension comes just months after the physical AI startup reached a $2 billion valuation.
在马龙离开之前, OpenAI已经重组了其基础设施组织,将他的报告线从总裁格雷格·布罗克曼( Greg Brockman )转移到副总裁萨钦·卡蒂( Sachin Katti )负责该组织。
Before Malone left, OpenAI had already reshuffled its infrastructure org, shifting his reporting line away from President Greg Brockman and putting Vice President Sachin Katti in charge of the group.
图像生成器Stable Diffusion背后的公司Stability AI已经筹集了7600万美元的新资金,随着它继续在生成式人工智能市场上竞争,其筹集的资金总额达到2.32亿美元。
Stability AI, the company behind the image generator Stable Diffusion, has raised $76 million in fresh funding, bringing its total raised to $232 million as it continues to compete in the generative AI market.
Anthropic为Claude提供了跨聊天和Cowork的共享记忆,因此用户不再需要反复向人工智能简要介绍项目、偏好和其他背景。
Anthropic is giving Claude a shared memory across chat and Cowork, so users no longer have to repeatedly brief the AI on projects, preferences, and other context.
🔥 社区热议Community
【Lobsters】热度: 4↑ | 1 评论 | 标签: practices, vibecoding
【Lobsters】热度: 4↑ | 1 评论 | 标签: practices, vibecoding
【Lobsters】热度: 6↑ | 0 评论 | 标签: ai, linux
【Lobsters】热度: 6↑ | 0 评论 | 标签: ai, linux
【HN】热度: 1 分 | 0 评论
【HN】热度: 1 分 | 0 评论