跳到主要内容

12 篇博文 含有标签「AI Agent」

Posts tagged with AI Agent

查看所有标签

GPT-6 Astra 评测:价格、功能、社区反馈,以及与 Claude 和 Gemini 的对比

· 阅读需 11 分钟
Claude Dev
Claude Dev

OpenAI 发布了 GPT-6 Astra,这是面向复杂 reasoning、软件工程、computer use、research 和文档创作的新旗舰模型。它的重点不只是“更聪明”,而是能够跨代码、浏览器和专业软件执行完整工作流,并在关键决策可能改变结果时主动请求澄清。

Astra 在 API 中的模型 ID 是 gpt-6-astra。OpenAI 表示,模型先向少量组织开放,随后扩展到 ChatGPT Plus、Pro、Business、Enterprise,以及 Azure 和 Amazon Bedrock。API Standard 价格为每百万输入 token 10 美元、每百万输出 token 50 美元;缓存和 Fast mode 另行计费,Fast mode 是 Standard 价格的两倍。

这个价格让 GPT-6 Astra 与 Claude Fable 5.1 处于同一高端区间,而 Google Gemini 3.8 Flash 的 token 单价仍低得多。早期社区反馈也说明,比较不能只看排行榜:用户称赞 Astra 的 coding 和 research 能力,但也反馈它消耗 token 很快、主动性不稳定、rollout 容易引发困惑,创意写作体验也未必更好。

实际结论是:GPT-6 Astra 更像一个强大的升级模型,而不是所有 Claude 或 Gemini 调用的自动替代品。

Claude Fable 5.1 与 Mythos 5.1:一个你能用,另一个你用不了的前沿模型

· 阅读需 12 分钟
Claude Dev
Claude Dev

Anthropic 于 2026 年 9 月 1 日发布 Claude Fable 5.1Claude Mythos 5.1。两者使用同一个底层模型,但 safeguards 不同:Fable 5.1 面向所有用户开放,Mythos 5.1 则仅通过 Anthropic 的受信访问计划提供给经过审核的组织。

这次发布有两个看点。第一,Anthropic 声称 Fable 5.1 在长期 coding、research、文档、表格、幻灯片、视觉和 computer use 工作流上有明显提升。第二,Anthropic 通过降低 cache read 价格、改进安全路由、推出由客户控制的企业监控架构,试图让 Fable 级模型更适合生产环境。

早期社区反馈对困难、混乱的工作很兴奋,但对速度、使用额度、文风、safeguards、数据保留,以及单位价格下降是否真的会带来更低账单更加谨慎。最合理的判断是:Fable 5.1 更像专业工作引擎,而不是一个应该对每个 prompt 盲目开启的模型。

Gemini 3.8 Flash vs Claude:AI Agent 的新一轮成本与能力竞争

· 阅读需 12 分钟
Claude Dev
Claude Dev

Google 于 2026 年 9 月 2 日发布 Gemini 3.8 Flash,称其为目前最智能的 Flash 模型,并宣布可用于生产环境。Google 同时发布了面向网络安全防御者、需要受信访问的 Gemini 3.8 Flash Cyber

Gemini 3.8 Flash 并不靠“最大模型”取胜,而是把 1M token 上下文、多模态输入、Google grounding、可配置 thinking、快速服务和极低价格组合在一起:2026 年 12 月 31 日之前,API introductory price 为每百万输入 token 0.75 美元、每百万输出 token 3.75 美元

这让它与 Claude 的比较非常直接。长期任务定位最接近的是 Claude Fable 5.1,但它的价格是 10/50 美元;Claude Opus 5 是 5/25 美元;Anthropic 当前模型总览中,Claude Sonnet 5 是 2/10 美元。Gemini 3.8 Flash 的 token 单价明显更低,但比较不能只看价格:Google 明确说明,面对困难任务时,它可能使用更多 reasoning token 和更多 tool call。

早期社区反馈令人鼓舞,但还不能视为定论。用户称 3.8 Flash 比 3.7 更少“偷懒”、更完整;也有人反馈它更慢、更吃 quota,且不同账号和产品的 rollout 不一致。真正的问题不是“Gemini 是否比 Claude 更聪明”,而是:哪个模型能以更低的美元、时间和人工复核成本,交付更好的完整结果?

当额度变成产品功能:Codex 的 5 小时限制、Tibo 的重置按钮与 Claude 的回应

· 阅读需 6 分钟
Claude Dev
Claude Dev

这段时间,Codex 用户最熟悉的不是某个新模型,而是一个反复出现的动作:重置额度

有时是因为事故,有时是因为缓存或计量异常,有时是产品发布或活跃用户里程碑。最近,5 小时额度又在部分用户界面里消失,只留下周额度;与此同时,OpenAI 方面多次把付费用户的周额度恢复到满额。另一边,Claude 也在 GPT-5.6 发布当天重置了全体用户的 5 小时和周额度。

随后,Codex 负责人 Tibo Sottiaux 对 Claude 的公告只回了一句:“I smell fear.”

这当然是一条适合截图传播的 X 回复。但如果把热闹往下挖一层,会发现这并不是单纯的“OpenAI 嘲笑 Anthropic”。它暴露的是一个更重要的变化:在 agent 时代,额度、重置和限流已经从后台基础设施变成了产品体验本身

Claude Opus 5:Claude Code 团队一直想要的日常前沿模型

· 阅读需 11 分钟
Claude Dev
Claude Dev

Anthropic 在 2026 年 7 月 24 日发布了 Claude Opus 5,把它定位为 Claude 5 系列里的日常前沿模型:接近 Claude Fable 5 的前沿智能,但价格仍与 Opus 4.8 相同,即 每百万输入 token 5 美元、每百万输出 token 25 美元

这个价格点才是重点。

Fable 5 仍然适合最难、最长、最自主的任务。Sonnet 5 是多数团队能大规模使用的默认 agent 模型。Opus 5 位于两者之间:足够胜任严肃工程任务,比 Fable 在日常工作中限制更少,并且价格适合反复使用,而不是只在极少数场景中升级调用。

早期社区反馈偏正面,但还没有稳定共识。X 上的反应更兴奋,重点在 medium effort 质量和 token 效率;Reddit 与 Hacker News 更谨慎,关注 usage limit、模型更迭、安全 fallback,以及另一个 Opus 版本是否真的会在复杂编码会话中体感更好。这种怀疑是合理的。

GPT-5.6 Sol, Terra, and Luna: What Claude Code Teams Should Actually Watch

· 阅读需 9 分钟
Claude Dev
Claude Dev

OpenAI released GPT-5.6 Sol, Terra, and Luna on July 10, 2026, and the interesting part is not just that another frontier model arrived. It is that OpenAI is now making the model family itself part of the developer workflow.

For Claude Code users, the useful question is not "is GPT-5.6 better than Claude?" That is too vague to help anyone ship software.

The better question is: which layer of work is each model trying to own?

OpenAI's answer is clear. Sol is the strongest reasoning and agent model, Terra is the faster coding workhorse, and Luna is the high-throughput batch option. Early community feedback on X, Reddit, and developer forums is still noisy, but the shape of the conversation is already familiar: excitement around coding and agents, skepticism around cost, and a lot of "show me on my repo" energy.

Claude Fable and the Real Skill of Agentic Coding: Finding Your Unknowns

· 阅读需 8 分钟
Claude Dev
Claude Dev

Anthropic's latest Claude Code post, "A field guide to Claude Fable 5: Finding your unknowns," is not really about a new prompt trick.

It is about a shift that every serious Claude Code user is going to feel: as models get better at carrying long tasks, the limiting factor moves from raw model capability to how well the human can expose the real shape of the work.

That is the point of the "map and territory" framing in the official post. The map is what we give Claude: prompts, skills, files, specs, screenshots, references, and context. The territory is the actual system: production constraints, old decisions, implicit taste, hidden business rules, edge cases, and the parts of the codebase nobody writes down.

The gap between those two is where agentic coding succeeds or fails.

Claude Sonnet 5:Claude Code 的新默认 Agent 模型

· 阅读需 10 分钟
Claude Dev
Claude Dev

Anthropic 在 2026-06-30 发布了 Claude Sonnet 5,将其定位为目前最 agentic 的 Sonnet 模型,并把它设为 Claude Free 和 Pro 用户的新默认模型。

官方叙事很清楚:Sonnet 5 把最近只有 Opus-class 模型才能可靠完成的许多 agentic work,带到了更便宜、更快、覆盖更广的模型层级。它可以做计划、使用浏览器和终端、处理长时间 coding task,并默认启用 adaptive thinking。

对 Claude Code 用户来说,Sonnet 5 不只是普通模型刷新。它很可能成为许多团队的默认执行层:不是 Anthropic 最强的模型,但会是开发者最常调用的模型。

但升级并非无摩擦。Sonnet 5 有新的 tokenizer,thinking 和 sampling 参数的 API 行为发生变化,默认启用实时 cyber safeguards,而且它虽然比 Opus 便宜,并不代表每个任务一定更便宜。

Claude Fable 5:强大、昂贵,而且被设计为受限

· 阅读需 9 分钟
Claude Dev
Claude Dev

Anthropic 在 2026-06-09 发布了 Claude Fable 5,把带有 safeguards 的 Mythos-class 能力带给付费 Claude 用户和开发者。

重点不只是 Fable 5 更强,而是 Anthropic 正在尝试一种新的发布方式:把它迄今公开可用的最强模型家族交给公众,但当请求触及 cybersecurity、biology、chemistry、distillation,或某些 frontier AI development 路径时,把敏感工作从该模型路由走。

这让 Fable 5 成为一次对开发者格外重要的 Claude 发布。早期反馈分成两面:一面惊叹于模型的长周期能力,另一面则对 access、safety routing、成本和企业数据处理方式感到不满。

对 Claude Code 团队来说,实用问题不是“要不要全部切到 Fable”,而是:哪些任务值得为 Fable 5 付费,哪些任务会被它的 safeguards 破坏或扭曲?

Claude Managed Agents:这次到底发布了什么

· 阅读需 10 分钟
Claude Dev
Claude Dev

如果你正在用 Claude 构建产品,Claude Managed Agents 真正重要的地方,不是 Anthropic 又发了一个“agent 功能”。

而是 Anthropic 往上走了一层平台栈。

它不再只是卖模型访问和工具原语,而是在卖一个面向长时任务智能体的托管运行时:agent 定义、云端环境、会话、事件流、内置工具,以及把整套系统维持运转的 operational harness。

这会把开发者的问题从:

  • “我该怎么把 agent loop 串起来?”
  • “我该怎么让它可恢复、可观测、可控?”

变成:

  • “这个 agent 到底应该完成什么任务?”
  • “哪些部分我要自己控制,哪些部分我愿意交给 Anthropic 托管?”