🤖 AI Agent 研究Research
随着代理从研究原型转向部署工具,他们的能力越来越依赖于模型-外部执行基础设施,通常称为代理线束。在保持模型权重固定的同时更改此线束可以大大改变任务性能。
As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model weights fixed can substantially alter task performance.
人类学习的许多重要形式都是从一个模糊的目标开始的,比如“成为一个更好的物理学家”或“提高研究水平”。“学员必须解释目标,识别能力差距,决定如何学习,并确定他们是否真的得到了改善。我们推出了ASPIRE ,这是一个模糊目标驱动的自我进化的基准。
Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution.
自主代理开始端到端地进行机器学习( ML )研究。这些代理将模型骨干与用于规划、执行、内存和验证的线束相结合,但这种架构仍然将特定领域的专业知识留在代理之外。
Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents combine a model backbone with a harness for planning, execution, memory, and verification, but this architecture still leaves domain-specific know-how outside the agent.
评估LLM代理对于指导他们的发展至关重要,但它已经变得过于昂贵:代理基准的前沿模型的单次通过可能会花费数百到数千美元,这是在迭代开发周期中反复支付的价格。在这项工作中,我们引入了早期结果预测,这是一种互补的效率轴,可以降低每项任务的成本。
Evaluating LLM agents is essential for guiding their development, yet it has grown prohibitively expensive: a single pass of a frontier model over an agentic benchmark can cost hundreds to thousands of dollars, a price paid repeatedly across iterative development cycles. In this work, we introduce early outcome prediction, a complementary axis of efficiency that instead cuts cost within each task.
⭐ GitHub 热门项目GitHub Trending
【GitHub】开源0到专家编程课程和Claude Code导师技能:预写课程、测试练习、严格的精通门。先行先行,多语言设计。(⭐ 1 )
【GitHub】Open-source 0-to-expert programming curriculum and Claude Code tutor skill: pre-authored lessons, tested exercises, strict mastery gates. Go first, multi-language by design. (⭐ 1)
【GitHub】在Windows上删除Claude Code会话历史记录—批量查看、搜索和删除存档的Claude Code CLI会话。免费,开源,无遥测。(⭐ 0 )
【GitHub】Delete Claude Code session history on Windows — view, search, and remove archived Claude Code CLI sessions in bulk. Free, open source, no telemetry. (⭐ 0)
【GitHub】一组编码代理,每个代理都有自己的模型,在自己的git工作树中的每个作业,在一个终端中。一切都是插件;混合模型是底壳。(⭐ 3 )
【GitHub】A team of coding agents, each on its own model, every job in a git worktree of its own, in one terminal. Everything is a plugin; mixed models are the base case. (⭐ 3)
【GitHub】工具包可重用para desarrollar, diagnosticar y refactorizar Adobe Commerce/Magento开源con Claude Code, optimizando contexto y consumo de tokens. (⭐ 0 )
【GitHub】Toolkit reutilizable para desarrollar, diagnosticar y refactorizar Adobe Commerce / Magento Open Source con Claude Code, optimizando contexto y consumo de tokens. (⭐ 0)
【GitHub】Apache SkyWalking AI Sessionizer -长寿命AI智能体的对话级可观察性(⭐ 3 )
【GitHub】Apache SkyWalking AI Sessionizer - conversation-level observability for long-lived AI agents (⭐ 3)
🚀 模型与行业动态Models & Industry
这家备受瞩目的初创公司的年收入运行率超过1亿美元( $ 100 million )。
The high-profile startup's annual revenue run rate stands at over $100 million.. Accel reportedly in talks to lead $1B round for Thinking Machines at $40B valuation.
Abliteration.AI正在使没有护栏的强大人工智能模型更容易访问,认为为捍卫者提供与不良行为者相同的工具最终可以改善网络安全。
Abliteration.AI is making powerful AI models without guardrails easier to access, arguing that giving defenders the same tools as bad actors could ultimately improve cybersecurity.
对于其用于操作编码和其他代理的新Muse Spark模型, Meta为通过分享提示和模型输出为未来模型开发做出“贡献”的用户提供平均约95%的显式折扣。
For its new Muse Spark model, intended for operating coding and other agents, Meta is offering an explicit discount averaging out to about 95% for users who "contribute" to the development of future models by sharing their prompts and model outputs.
OpenAI声称, Astra代表了“计算机和浏览器使用的新领域” ,它以无与伦比的“速度、准确性和安全性”处理任务。
OpenAI claims that Astra represents "a new frontier on computer and browser use," and that it handles tasks with unmatched "speed, accuracy, and safety."
🔥 社区热议Community
【HN】热度: 102 分 | 32 评论
【HN】热度: 102 分 | 32 评论
【HN】热度: 202 分 | 62 评论
【HN】热度: 202 分 | 62 评论