🤖 AI Agent 研究Research
编码代理现在通常在SWE-bench系列基准上进行评估,其任务是根据精心策划的GitHub问题构建的:长、结构化和信息丰富。然而,真正的用户请求通常要短得多,结构化程度也较低。
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured.
本文研究自主软件开发,其中基于LLM的编码代理在没有人为干预的情况下将高级需求转换为完整、功能性和可用的软件系统。我们引入了Harness-of-Harness ( HoH )框架,该框架使编码代理能够在自主开发过程中不断改进软件。
This paper studies autonomous software development, in which LLM-based coding agents transform high-level requirements into complete, functional, and usable software systems without human intervention. We introduce Harness-of-Harness (HoH), a framework that enables coding agents to continually improve software during autonomous development.
认知语言代理通过为语言模型配备记忆、工具和决策程序,使代理能够在交互式环境中进行推理和行动,从而取得了重大进展。现有的框架在很大程度上将这些代理作为解决用户指定的、有界任务的系统。
Cognitive language agents have achieved substantial progress by equipping language models with memory, tools, and decision-making procedures, enabling agents to reason and act in interactive environments. Existing frameworks largely cast these agents as systems for solving user-specified, bounded tasks.
法学硕士评委被广泛用于评估代理工具调用系统,但其在结构化、依赖性驱动的工作流程中的可靠性在很大程度上仍未得到检验。我们推出了AgentJudgeBench ,这是系统研究LLM-as-a-judge在工作流程DAG上进行代理工具调用的可靠性的第一个基准,这与更广泛的LLM-as-a-judge开放式文本或偏好评估任务不同。
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation.
⭐ GitHub 热门项目GitHub Trending
【GitHub】零依赖日期和时间MCP服务器—一个文件,纯Python stdlib , DST感知。7个工具:当前时间,转换,偏移,自然语言解析,持续时间,会议窗口,格式。(⭐ 0 )
【GitHub】Zero-dependency date & time MCP server — one file, pure Python stdlib, DST-aware. 7 tools: current time, convert, offset, natural-language parse, durations, meeting windows, format. (⭐ 0)
【GitHub】“我使用Python、Gemini和模型上下文协议构建了GitHub MCP AI Agent。代理理解自然语言的GitHub请求,并动态选择列出存储库等MCP工具(⭐ 0 )
【GitHub】"I built a GitHub MCP AI Agent using Python, Gemini, and the Model Context Protocol. The agent understands natural-language GitHub requests and dynamically selects MCP tools such as listing repositori (⭐ 0)
【GitHub】在Node.js、Python或Java中将任何Web应用程序转换为可安装的、符合标准的MCP服务器—具有跨语言奇偶校验、默认安全、OAuth 2.1+PKCE和代理驱动的技能模式。(⭐ 0 )
【GitHub】Turn any webapp into an installable, standards-compliant MCP server in Node.js, Python, or Java — with cross-language parity, safety-by-default, OAuth 2.1+PKCE, and an agent-driven Skill mode. (⭐ 0)
【GitHub】用于使用Apexnova AI Hub模型的开源连接器,具有OpenCode、Codex和Claude Code。(⭐ 0 )
【GitHub】Open-source connector for using Apexnova AI Hub models with OpenCode, Codex, and Claude Code. (⭐ 0)
【GitHub】自我观察工具: Claude Code的开源代理定义,将Claude变成一个具有不可侵犯规则的生活顾问—没有神经生物学基础就没有建议。(⭐ 0 )
【GitHub】A self-seeing instrument: an open-source agent definition for Claude Code that turns Claude into a life counselor with one inviolable rule — no advice without a neurobiological foundation. (⭐ 0)
🚀 模型与行业动态Models & Industry
在我们新的真实世界人工智能舞台上,我们将专注于数字和物理之间的交集,以及我们将继续看到两者的融合的所有方式。
On our new Real World AI stage, we’ll be focusing on the intersection between the digital and physical, and all the ways we’ll continue to see a blending of the two.
OpenAI的新Astra模型将使用“循环深度” ,这种技术允许模型在大多数推理模型的顺序思维之外运行。
OpenAI’s new Astra model will use “recurrent depth,” a technique that allows the model to operate outside of the sequential thinking that characterizes most reasoning models.
互联网存在信任问题,这不仅仅是因为社交媒体订阅源充斥着人工智能。人工智能生成的文本和图像现在正在进入求职申请、产品评论甚至保险理赔领域。
The internet has a trust problem, and it’s not just because social media feeds are filling up with AI slop. AI-generated text and images are now making their way into job applications, product reviews, and even insurance claims, l
🔥 社区热议Community
【Lobsters】热度: 40↑ | 25 评论 | 标签: vibecoding, visualization
【Lobsters】热度: 40↑ | 25 评论 | 标签: vibecoding, visualization
【Lobsters】热度: 3↑ | 0 评论 | 标签: vibecoding
【Lobsters】热度: 3↑ | 0 评论 | 标签: vibecoding
【HN】热度: 314 分 | 144 评论
【HN】热度: 314 分 | 144 评论
【HN】热度: 372 分 | 164 评论
【HN】热度: 372 分 | 164 评论