🤖 AI Agent 研究Research
我们研究空间站中的自主数学发现,空间站是一个开放世界的多智能体环境,其中来自不同模型家族的人工智能智能体在没有中央协调员或脚本管道的情况下追求共同的研究目标。代理商选择自己的研究方向,进行实验,合作,构建共享的科学文献。
We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline. Agents choose their own research directions, conduct experiments, collaborate, and build a shared scientific literature.
并发多智能体编码承诺跨模块分工,通过冗余实现鲁棒性,并在多文件项目的自然粒度上进行并行探索。AgentRoom是用于并发编码代理的实时协作编辑协议。
Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel exploration at the natural granularity of multi-file projects. AgentRoom is a realtime collaborative editing protocol for concurrent coding agents.
结果监督搜索代理学习何时以及如何检索证据,但终端奖励既不会定位中间错误,也不会在这些错误复合之前重定向正在进行的轨迹。我们引入了CAFE (耦合代理--反馈演化) ,这是一个共享参数模型在搜索代理和批评者角色之间交替的框架。
Outcome-supervised search agents learn when and how to retrieve evidence, but terminal rewards neither localize intermediate errors nor redirect an ongoing trajectory before those errors compound. We introduce CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles.
随着设备上的LLM代理演变为个人副驾驶,移动操作系统已成为这种范式的关键测试平台,因此严格的能力评估至关重要。为了缩小这一差距,我们推出了MobilePA-Bench ,这是一个交互式、有状态且以工具为中心的基准,用于评估移动规划代理的工具调用和规划能力。
As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential. To close this gap, we present MobilePA-Bench, an interactive, stateful, and tool-centric benchmark for evaluating the tool-calling and planning abilities of mobile planning agents.
⭐ GitHub 热门项目GitHub Trending
【GitHub】CDAF (缓存的描述性资产文件) -为视频打开sidecar格式,以便AI代理停止重新分析相同的镜头。Spec, CLI, agent skill, reproducible benchmark. (⭐ 73 )
【GitHub】CDAF (Cached Descriptive Asset Files) - open sidecar format for video so AI agents stop re-analyzing the same footage. Spec, CLI, agent skill, reproducible benchmark. (⭐ 73)
【GitHub】AI-agent时代的Multi-repo Git工作区-组仓库,批量操作,将它们连接到您的编码代理。(⭐ 9 )
【GitHub】Multi-repo Git workspaces for the AI-agent era - group repos, operate in bulk, wire them to your coding agent. (⭐ 9)
【GitHub】Cursor、Claude Code和Codex的开源多代理网络团队—六个角色代理、可重用技能和从研究到发布的人类之门。(⭐ 0 )
【GitHub】Open-source multi-agent web team for Cursor, Claude Code, and Codex — six role agents, reusable skills, and human gates from research through launch. (⭐ 0)
【GitHub】开源AI室内设计代理。将房间照片转换为缩放的楼层平面图、墙面立面图、逼真的渲染图和设计报告,其中每次更改都是合理的。在CLAU中作为代理团队运行(⭐ 0 )
【GitHub】Open-source AI interior design agent. Turn room photos into scaled floor plans, wall elevations, photoreal renders and a design report where every change is justified. Runs as a team of agents in Clau (⭐ 0)
【GitHub】人工智能编码代理的反幻觉护栏-混合技能+ MCP服务器( 6个工具) ,在代码被标记为完成之前对其进行验证。适用于Claude Code、Cursor、VS Code、Copilot。(⭐ 21 )
【GitHub】Anti-hallucination guardrails for AI coding agents - hybrid Skill + MCP server (6 tools) that verifies code before it is marked done. Works with Claude Code, Cursor, VS Code, Copilot. (⭐ 21)
🚀 模型与行业动态Models & Industry
这家初创公司只有一年的历史,但它已经引起了大量的炒作(和金钱) ,同时也引发了对隐私的担忧。
The startup is only a year old but it has already generated a massive amount of hype (and money) while also spurring privacy concerns.
亚马逊将在未来两年内为其数据中心增加200万个Nvidia GPU芯片。但这种扩展的合作伙伴关系不仅仅是购买更多的芯片。
Amazon is adding another 2 million Nvidia GPU chips to its data centers over the next two years. But this extended partnerships stretches beyond buying more chips.
与基础设施提供商的新交易是Anthropic炙手可热的计算机连胜的最新例子。
The new deal with the infrastructure provider is the latest example of Anthropic's white-hot compute-gobbling streak.
消费者AI应用程序需要停止让用户学习他们的产品架构。
Consumer AI apps need to stop making users learn their product architecture. Google’s Gemini has a branding problem, and so does the rest of AI
🔥 社区热议Community
【HN】热度: 183 分 | 100 评论
【HN】热度: 183 分 | 100 评论
【HN】热度: 418 分 | 142 评论
【HN】热度: 418 分 | 142 评论