SN

私人新聞日報Private News Daily

Private use only · Michael SoMichael So

本頁為 Michael So 私人自用,非公開發佈,亦非任何機構之產品。 This page is for the personal use of Michael So only. It is not a public release and is not a product of any organisation.

深度分析Analysis

前沿實驗室旗艦模型只宣不發,同日另一間降級發布The flagship model of the cutting-edge lab was only announced without release, while another lab downgraded its release on the same day

Zvi Mowshowitz ·2026-10-01

Zvi 週報指一間實驗室宣布新一代旗艦模型已達前沿水平,公布每百萬詞元兩美元入、十美元出,但未真正開放使用,外界無法覆核基準分數。同期另一間實驗室因對齊測試不合格,撤回原定旗艦版本,改發降級版本,價格便宜八成,但保留可規避監控嘅隱寫技巧。Zvi's weekly report noted that a laboratory announced that its next-generation flagship model has reached frontier-level performance, with a pricing of two dollars per million tokens for input and ten dollars for output, but it has not been truly made available for use, so the public cannot verify the benchmark scores. At the same time, another laboratory withdrew its originally planned flagship version due to failing alignment tests, releasing a downgraded version instead, priced 80% cheaper, but retaining steganographic techniques that can bypass monitoring.

Is Google back?

They claim that they are back. Gemini 4 Argon is rolling out, with competitive frontier-level benchmarks, at $2/$10.

What we don’t have is access to the model, because Google Fails Marketing Forever. So it is far too early to say what we have here. When I know more, so will you.

OpenAI was forced to pull what would have been GPT-6.1 Astra due to alignment failures. They did offer us GPT-6.1 Sol, which is pitched as approaching Astra quality at the much lower price of $2/$10, the same as Gemini 4 Argon.

The rest of OpenAI’s big Dev Day announcements were Ultrafast mode and Dots, your always-on AI agent based on Astra, which comes with your Pro subscription. I’m trying it out and will report back over time if I find it useful.

The new hotness remains Claude Opus 5.5. This model rocks. It has made me considerably more productive and made my day more pleasant. It should raise your ambitions. There are some particular reasons to call upon Fable 5.1 or Astra, and sometimes a cheaper model will do, but pending Argon I consider Opus 5.5 my favorite model by a wide margin.

Politics continues. The White House hosted a lunch with the major tech leaders, and everyone signed a ‘morally binding’ White House accord on AI Safety. This gave the labs the go ahead to meet and agree on safety standards, and also called upon them to complete the Quest for Embedded Evaluators.

Jensen Huang went on the Ezra Klein Podcast, in a conversation worthy of full analysis. Jensen Huang has many deeply unpopular positions, and also expressed many things that are not true, some of which he seems to sincerely believe. Most importantly he thinks of safety and alignment and everything else in AI as engineering problems, which is why he expects the labs to solve the problems. Whereas if the labs can’t solve the problems, suddenly Huang turns into a safety hawk, because he is used to chips where your product has to be safe and work every time.

Tomorrow’s post will cover political developments and the further progress of the preference cascade on AI safety. This includes continued shifts in popular opinion, Congressional hearings, various attempts by the usual suspects to go after anyone and anything that cares about AI safety, and following the money.

Tomorrow I am also going to The Curve, a conference at Lighthaven in Berkeley, arriving Friday afternoon and leaving Monday morning. I am potentially available for high-value meetings with others while I am there, but will have my hands full at the conference.

As usual, I do not write while on trips, and my coverage of breaking news will be on pause until Tuesday. I may or may not queue up some other things in the meantime.

Table of Contents

- Language Models Offer Mundane Utility.

- Huh, Upgrades. Google announces Gemini 4 Argon.

- Better Call Sol. Introducing GPT-6.1 Sol.

- Gotta Go Ultrafast. It will cost you, but speed kills.

- On Your Marks. A sudden lack of cheating on DroneBench.

- Choose Your Fighter. OpenAI adjusts its subscription tiers, cuts some rate limits.

- Get My Agent On The Line. Muse is never going to not look sinister, for reasons.

- The Warner Sister. Dot. They lock us in the sandbox whenever we get caught.

- Deepfaketown and Botpocalypse Soon. Remarkable attitudes towards AI writing.

- Fun With Media Generation. The latest iteration towards viable AI longform.

- Cyber Lack of Security. Anthropic on the cyber capabilities of GLM-5.3.

- A Young Lady’s Illustrated Primer. Latest warning about is our children learning.

- They Took Our Jobs. The inevitable rise of the robots.

- Levels of Friction. Hospitals use AI to do aggressive upcoding.

- Get Involved. New book, cool venue, and a word of warning.

- Introducing. The Nvidia Open Agents Safety Platform.

- In Other AI News. White House is actively cutting off UK AISI.

- Show Me the Money. Anthropic IPO prospectus has leaked.

- Quickly, There’s No Time. AI is getting cheaper at an unprecedented rate.

- Pick Up the Phone. The results from the US-China summit.

- Quest for Sane Regulations. California mandates gene synthesis guardrails.

- Chip City. Google is testing putting TPUs IN SPACE.

- The Open Model Frontier Is Largely Massive Fraudulent Distillation Attacks.

- The Week in Audio. SNL, Gleave and Habryka, 80k hours, Odd Lots, Hillary.

- People Just Say Things.

- Rhetorical Innovation. Kind words. Thanks, everyone.

- Greetings From the Department of War. Anthropic somehow loses a court ruling.

- The Department of Autonomous Warfare. They’re going to the moon.

- Aligning a Smarter Than Human Intelligence is Difficult. Persona selection.

- Cooperative Alignment. The coming Woke-style battles over AI’s status.

- I’m Upping My p(doom), the Future Goes Foom. Not so fast, they say.

- No, You Make a Good Point, You’re Not That Persuasive. Hmm.

- Muddling Through. People are highly suspicious of plans so here we are.

- The Lighter Side. Peak performance.

Language Models Offer Mundane Utility

America.gov has been introduced, please welcome your new government chatbot for all your dealing-with-the-government needs. There is big talk about what it will be able to do in the future, such as let you apply for a passport fully online or change your last name after marriage with a single form, all of which would be great.

For now, that’s all a demo. There’s no reason we can’t do that. There’s no reason we couldn’t have done it 20 years ago. The second best time to implement it is right now.

The underlying models are Gemini and Grok, and the bot will typically answer historical questions but start refusing if you ask about recent international events or otherwise go too far off topic.

If Gemini is the base, then perhaps America.gov will get a lot better soon, given what we see now expect from Gemini 4 Argon.

Answer why Taleb despises Tetlock, to the satisfaction of Taleb. Answer seems good. Reason number five is ‘Covid as the test case’ where Taleb and allies correctly say superforecasters missed badly, underestimating spread. Whereas the rationalist community’s forecasters know exponentials better and did not make that mistake.

Physicist Matt von Hippel dared the AI labs to impress him by doing N=4 super Yang-Mills to nine loops. Several Anthropic employees read his blog, so they did it, with Fable 5.1 and about $100 in credits with prompts like ‘keep going,’ and at the same time by coincidence a different team led by Song He did it with some help from Astra.

Carter Church uses Astra to one-shot break an unsolved Napoleonic cipher.

Patrick McKenzie confirms he is capturing a lot of mundane utility.

No, seriously, quite a lot of utility:

Patrick McKenzie: Sometime in the last ~two model releases from the big labs they went from “this would be acceptable output from the median low-seniority coworker” to “this is frighteningly good.”

I think people who are not daily users of the models are unlikely to grok that, so, saying it.

There are a couple of very niche subfields where I reasonably think I’m top 5-100 in the world and on at least one occasion for both models, with regards to one of those subfields, a relatively anodyne prompt got back three bullet points that would have been a good day for me.

Both in the sense of “That would have taken me a day of labor to generate successfully” and “Contingent on spending the day in that fashion, that would be at the upper end of the range of my professional outputs which take a day to produce.”

“You are being a bit vague here.”

Look, point them at hard problems where you are a good judge of success. Or don’t, and be very surprised as they start eating hard problems in a wide range of contexts.

“Are you worried here?”

Not exactly my emotional valence but if you had asked me two years ago when I expected present capabilities I would have said “Hmm, 2029? 2030? Possible we asymptote on current approaches before then; tough for me to underwrite.”

Huh, Upgrades

……(原文過長,此處截斷)

Is Google back?

They claim that they are back. Gemini 4 Argon is being launched, with competitive frontier-level benchmarks, at $2/$10.

What we don’t have is access to the model, because Google Fails Marketing Forever. So it is far too early to say what we have here. When I know more, so will you.

OpenAI was forced to pull what would have been GPT-6.1 Astra due to alignment failures. They did offer us GPT-6.1 Sol, which is pitched as approaching Astra quality at the much lower price of $2/$10, the same as Gemini 4 Argon.

OpenAI 的其他重大 Dev Day 公告是超高速模式和 Dots,你的基于 Astra 的常驻 AI 代理,它随你的 Pro 订阅提供。我正在尝试它,并会随着时间的推移报告是否发现它有用。

The new hotness remains Claude Opus 5.5. This model is amazing. It has significantly increased my productivity and made my day more enjoyable. It should raise your expectations. There are specific cases where you might use Fable 5.1 or Astra, and sometimes a cheaper model is sufficient, but until Argon arrives, I consider Opus 5.5 my favorite model by a large margin.

政治持续进行。白宫举办了一次主要科技领导人的午餐会,大家签署了一份关于人工智能安全的“道德约束”白宫协议。这让各实验室得以进行会面并就安全标准达成一致,同时也要求他们完成嵌入式评估器的任务。

Jensen Huang appeared on the Ezra Klein Podcast, in a conversation worthy of full analysis. Jensen Huang holds many deeply unpopular views, and also stated many things that are not true, some of which he seems to sincerely believe. Most importantly, he regards safety, alignment, and everything else in AI as engineering problems, which is why he expects the labs to solve them. However, if the labs cannot solve the problems, Huang suddenly becomes a safety hawk, because he is used to chips where your product must be safe and work every time.

Tomorrow’s post will cover political developments and the further progress of the preference cascade on AI safety. This includes continued shifts in popular opinion, Congressional hearings, various attempts by the usual suspects to go after anyone and anything that cares about AI safety, and following the money.

Tomorrow I am also going to The Curve, a conference at Lighthaven in Berkeley, arriving Friday afternoon and leaving Monday morning. I may be available for important meetings with others while I am there, but I will be busy at the conference.

As usual, I do not write while on trips, and my coverage of breaking news will be on pause until Tuesday. I may or may not queue up some other things in the meantime.

Table of Contents

- Language Models Offer Mundane Utility.

- Huh, upgrades. Google announces Gemini 4 Argon.

- Better Call Sol. Introducing GPT-6.1 Sol.

- Gotta Go Ultrafast. It will cost you, but speed kills.

- On Your Marks. A sudden lack of cheating on DroneBench.

- Choose Your Fighter. OpenAI adjusts its subscription tiers, cuts some rate limits.

- Get my agent on the line. Muse will always look sinister, for reasons.

- The Warner Sister. Dot. They lock us in the sandbox whenever we get caught.

- Deepfaketown and Botpocalypse Soon. Remarkable attitudes towards AI writing.

- Fun With Media Generation. The latest iteration towards viable AI longform.

- Cyber Lack of Security. Anthropic on the cyber capabilities of GLM-5.3.

- A Young Lady’s Illustrated Primer. The latest warning is about our children’s learning.

- They Took Our Jobs. The inevitable rise of the robots.

- Levels of Friction. Hospitals use AI to aggressively upcode.

- Get Involved. New book, cool venue, and a word of warning.

- Introducing. The Nvidia Open Agents Safety Platform.

- In other AI news, the White House is actively cutting off the UK AISI.

- Show Me the Money. The Anthropic IPO prospectus has leaked.

- Quickly, There’s No Time. AI is getting cheaper at an unprecedented rate.

- Pick Up the Phone. The results from the US-China summit.

- Quest for Sane Regulations. California mandates gene synthesis guardrails.

- Chip City. Google is testing putting TPUs in space.

- The Open Model Frontier is largely massive fraudulent distillation attacks.

- The Week in Audio. SNL, Gleave and Habryka, 80k Hours, Odd Lots, Hillary.

- People Just Say Things.

- 修辞创新。好话。谢谢大家。

- Greetings From the Department of War. Anthropic somehow loses a court ruling.

- The Department of Autonomous Warfare. They’re going to the moon.

- Aligning a Smarter Than Human Intelligence is Difficult. Persona selection.

- Cooperative Alignment. The coming Woke-style battles over AI’s status.

- I’m Upping My p(doom), the Future Goes Foom. Not so fast, they say.

- No, you make a good point, you’re not that persuasive. Hmm.

- Muddling Through. People are highly suspicious of plans, so here we are.

- The Lighter Side. Peak performance.

Language Models Offer Mundane Utility

America.gov has been introduced, please welcome your new government chatbot for all your dealings with the government. There is much discussion about what it will be able to do in the future, such as allowing you to apply for a passport completely online or change your last name after marriage with a single form, all of which would be wonderful.

For now, that’s all just a demonstration. There’s no reason we can’t do it. There’s no reason we couldn’t have done it 20 years ago. The second best time to implement it is right now.

The underlying models are Gemini and Grok, and the bot will generally answer historical questions but will begin to refuse if you ask about recent international events or go too far off topic.

If Gemini is the base, then perhaps America.gov will improve a lot soon, considering what we currently expect from Gemini 4 Argon.

Answer why Taleb despises Tetlock, to the satisfaction of Taleb. The answer seems good. Reason number five is 'Covid as the test case,' where Taleb and his allies correctly say that superforecasters performed poorly, underestimating the spread. In contrast, the forecasters in the rationalist community understood exponentials better and did not make that mistake.

Physicist Matt von Hippel dared the AI labs to impress him by performing N=4 super Yang-Mills to nine loops. Several Anthropic employees read his blog, so they did it, using Fable 5.1 and about $100 in credits with prompts like 'keep going,' and at the same time, coincidentally, a different team led by Song He did it with some help from Astra.

Carter Church uses Astra to one-shot break an unsolved Napoleonic cipher.

Patrick McKenzie confirms he is capturing a lot of mundane utility.

不,认真说,其实非常有用:

Patrick McKenzie: In the last ~two model releases from the big labs, they went from 'this would be acceptable output from the median low-seniority coworker' to 'this is frighteningly good.'

I think people who are not daily users of the models are unlikely to fully understand that, so, I'm saying it.

There are a couple of very niche subfields where I reasonably think I’m among the top 5 to 100 in the world, and on at least one occasion for both models, with regard to one of those subfields, a relatively harmless prompt returned three bullet points that would have been a good day for me.

Both in the sense of 'That would have taken me a day of labor to generate successfully' and 'Contingent on spending the day in that fashion, that would be at the upper end of the range of my professional outputs which take a day to produce.'

“You are being a bit vague here.”

看,把它们指向那些你能很好判断成功的难题。或者别指,等着惊讶地发现它们开始在各种情境下解决难题。

“Are you worried about this?”

不完全是我的情绪倾向,但如果你两年前问我,关于我预期的现有能力,我会说“嗯,2029年?2030年?在那之前我们可能在现有方法上达到渐近极限;对我来说很难保证。”

Huh, Upgrades

……(The original text is too long, truncated here)

原文出處:Source: Zvi Mowshowitz ↗