🤖 AI Agent 研究Research
商业智能(BI)是企业决策的基石,在Power BI和Tableau等软件中被企业用户广泛使用。在传统的商业智能工作流程中,用户需要通过( 1 )识别相关表格, ( 2 )执行数据转换,以及( 3 )建立加入关系来准备数据,然后才能( 4 )回答他们的业务问题。
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions.
通过强化学习( RL )训练有能力的编码代理需要使用可靠的验证器完成各种任务。为了更好地扩展RL环境,我们提出了CodeMidas ,这是一个代理流水线,将现有代码库中实现的功能转换为可执行的RL环境,使用源代码作为其唯一的特定任务输入。
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input.
计算机使用代理( CUA )沿着两条独立的路线前进:通过代码和命令行进行图形交互和软件开发。我们介绍了RecreationWorld ,这是一个围绕娱乐构建的五个平台框架:给定一个正在运行的引用,代理必须发现其行为并构建一个没有规定工作流程的忠实实现。
Computer-use agents (CUAs) have advanced along two separate lines: graphical interaction and software development through code and the command line. We introduce RecreationWorld, a five-platform framework built around recreation: given a running reference, an agent must discover its behavior and build a faithful implementation with no prescribed workflow.
技能可以通过提供特定于任务的程序指导来提高大型语言模型( LLM )代理的性能,而技能优化可以通过迭代细化进一步提高其有效性。为了应对这些挑战,我们建议将技能表示为图形结构的自然语言工件。
Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. To address these challenges, we propose representing skills as graph-structured natural-language artifacts.
⭐ GitHub 热门项目GitHub Trending
【GitHub】用于查看Claude Code和Codex订阅的配额利用率的开源macOS菜单栏实用程序(⭐ 1 )
【GitHub】An open-source macOS menu-bar utility for viewing quota utilization from Claude Code and Codex subscriptions (⭐ 1)
【GitHub】通过Laya重新排序,在Claude Code、Codex和OpenCode聊天中进行本地VS Code搜索
【GitHub】Local VS Code search across Claude Code, Codex, and OpenCode chats with Laya reranking
【GitHub】LLM代理的扫雷基准。一个相同的电路板,最多可并行九个型号,一个时钟,一个工具层。
【GitHub】A Minesweeper benchmark for LLM agents. One identical board, up to nine models in parallel, one clock, one tool layer.
【GitHub】Codex、Claude Code和OpenCode的开源多Agent开发工作流程。(⭐ 0 )
【GitHub】Open-source multi-agent development workflow for Codex, Claude Code and OpenCode. (⭐ 0)
【GitHub】使用Magnific平台的自定义开源claude代码插件。(⭐ 0 )
【GitHub】A custom open source claude code plugin using the Magnific platform. (⭐ 0)
🚀 模型与行业动态Models & Industry
该小组不会有放慢或重定向OpenAI正在进行的数学研究的余地。
The group won't be given leeway to slow down or redirect OpenAI's ongoing mathematical research.
根据Appfigures的最新估计, Meta的新人工智能代理Muse在美国和加拿大的下载量和日活跃用户数量超过了ChatGPT在移动首次亮相后的同期。
Meta’s new AI agent Muse has racked up more downloads and daily active users in the U.S. and Canada than ChatGPT did over the same period after its mobile debut, according to new estimates from Appfigures.
亚马逊拥有自己的基础模型,以及互联网上最受欢迎的推理平台之一。只要他们没有法律义务向Muse敞开大门,他们为什么要这么做?
Amazon has its own cohort of foundation models, along with one of the most popular inference platforms on the internet. As long as they're under no legal obligation to open the doors to Muse, why would they?
🔥 社区热议Community
【HN】热度: 143 分 | 143 评论
【HN】热度: 143 分 | 143 评论
【HN】热度: 17 分 | 0 评论
【HN】热度: 17 分 | 0 评论
【HN】热度: 407 分 | 184 评论
【HN】热度: 407 分 | 184 评论