🤖 AI Agent 研究Research
LLM代理越来越依赖于生成的交互数据来学习如何与外部环境交互。这项工作开发了一个两级框架,将代理数据表示为通用分解对象(环境、任务、交互、验证器) ,并通过精度复杂度分割( ACE )透镜将生成公式化为约束分布设计。
LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. This work develops a two-level framework that represents agentic data as a common factorized object (environment, task, interaction, verifier) and formulates generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens.
现代软件系统积累了数十年的技术债务,使迁移变得昂贵且手动,但现有的基准只测试行为正确性,而不是迁移是否真正发生。这项工作引入了SWE重构平台( SWE Refactor Bench ) ,这是20个全存储库迁移的基准,具有三阶段评估协议,前沿编码代理仅通过520次运行中的5.4%。
Modern software systems accumulate technical debt over decades, making migration expensive and manual, yet existing benchmarks only test behavioural correctness rather than whether the migration actually occurred. This work introduces SWE Refactor Bench, a benchmark of 20 whole-repository migrations with a three-stage evaluation protocol, on which frontier coding agents pass only 5.4% of 520 runs.
编码代理运行跨越数十个模型调用和工具使用的长任务,用户在中期在更便宜和更强大的模型之间切换时面临成本-质量权衡。这项工作研究了切换非本地轨迹如何影响Claude和GPT家族的质量和成本,发现全轨迹升级仅恢复了不到一半的质量差距,作者称之为切换税。
Coding agents run long tasks spanning dozens of model calls and tool uses, and users face a cost-quality trade-off when switching between cheaper and stronger models mid-run. This work studies how handing off a non-native trajectory affects quality and cost across the Claude and GPT families, finding that full-trajectory escalation recovers less than half of the quality gap, a penalty the authors term the handoff tax.
Agent功能不仅仅由模型决定,因为包括内存、规划、操作协议和工具编排在内的线束可以主导基础模型的贡献。这项工作提出了JIT-Agent ,这是一种线束智能模型,可以即时合成任务自适应代理线束,使DeepSeek-V4-Flash在DeepSearchQA和OdysseyBench上超过GPT-5.6。
Agent capability is not determined by the model alone, since the harness, including memory, planning, action protocol, and tool orchestration, can dominate the contribution of the underlying foundation model. This work presents JIT-Agent, a harness intelligence model that synthesizes task-adaptive agent harnesses on the fly, letting DeepSeek-V4-Flash surpass GPT-5.6 on DeepSearchQA and OdysseyBench.
⭐ GitHub 热门项目GitHub Trending
【GitHub】适用于AI编码代理的隐私第一遥测收集器— Claude Code、Codex等。本地优先,开源。(⭐ 0 )
【GitHub】Privacy-first telemetry collector for AI coding agents — Claude Code, Codex, and more. Local-first, open source. (⭐ 0)
【GitHub】用于将Claude Desktop和Claude Code连接到外部AI模型的开源本地AI网关。(⭐ 1 )
【GitHub】Open-source local AI gateway for connecting Claude Desktop and Claude Code to external AI models. (⭐ 1)
【GitHub】开源Codex & Claude Code技能:通过本地Whisper将创意、音频、文章和视频转化为源接地的Telegram帖子和丰富消息,进行严格的验证和安全发布。(⭐ 2 )
【GitHub】Open-source Codex & Claude Code skill: turn ideas, audio, articles and video into source-grounded Telegram posts and Rich Messages with local Whisper, strict validation and safe publishing. (⭐ 2)
🚀 模型与行业动态Models & Industry
在TechCrunch Disrupt 2026上, AI Stage又回到了过去几年由Google for Startups提出的社区中最热门的话题。
At TechCrunch Disrupt 2026, the AI Stage is back to dig into the single hottest topic in the community for the past few years, presented by Google for Startups.
Zoph与Mira Murati共同创立了Thinking Machines Lab ,并担任创业公司的首席技术官,他曾在OpenAI短暂任职,现在在谷歌工作。
Zoph, who co-founded Thinking Machines Lab alongside Mira Murati and also served as the startup's CTO, led a brief stint at OpenAI and is now at Google.
一些全球最大的科技公司和人工智能初创公司聚集在一起,谴责当前的网络安全状况,并宣传一种他们认为可以抵御新一代网络威胁的新解决方案。
Some of the world's largest tech companies and AI startups have come together to decry the current state of cybersecurity and to advertise a new solution that they say can ward off a new generation of cyber threats.
这些更新表明,谷歌希望将AI Mode定位为某种程度上的AI旅行社,因为它不仅仅是帮助用户查找信息,而是实际处理行程规划和预订流程的一部分。
The updates indicate that Google is looking to position AI Mode as an AI travel agent of sorts, as it's moving beyond simply helping users find information to actually handling parts of the trip-planning and booking process.
Hugging Face首席执行官Clem Delangue表示, Microduck是一个“可以通过强化学习教授新技巧的开源机器人”。
Clem Delangue, CEO of Hugging Face, said the Microduck is an “open-source robot you can teach new tricks with reinforcement learning.”
🔥 社区热议Community
【Lobsters】热度: 154↑ | 133 评论 | 标签: vibecoding
【Lobsters】热度: 154↑ | 133 评论 | 标签: vibecoding
【Lobsters】热度: 58↑ | 14 评论 | 标签: vibecoding
【Lobsters】热度: 58↑ | 14 评论 | 标签: vibecoding
【HN】热度: 26 分 | 4 评论
【HN】热度: 26 分 | 4 评论
【HN】热度: 107 分 | 16 评论
【HN】热度: 107 分 | 16 评论
【HN】热度: 47 分 | 2 评论
【HN】热度: 47 分 | 2 评论