AI's Next Battle may Not be Intelligence, but Experience.
OpenAI dropped a line at its launch event: “Welcome to the AGI era.” Asked whether Astra itself qualifies as AGI, the company’s response was blunt: “I do think we’re there.” The same day, Artificial Analysis slotted GPT-6 Astra into its Intelligence Index — a score of 61, flat against the prior generation’s GPT-5.6 Sol. On the Coding Agent Index, Astra scored 67, barely above Sol’s 65, and not even the top score of its own generation. On one side, a declaration of the AGI era. On the other, a leaderboard that barely moved.
近日,OpenAI 在發布會上丟下一句:「Welcome to the AGI era。」被問到 Astra 本身算不算 AGI 時,官方的回應是:「I do think we’re there。」同一天,Artificial Analysis 把 GPT-6 Astra 排進它的 Intelligence Index——61 分,跟上一代 GPT-5.6 Sol 打平;Coding Agent Index 上,67 分,也只比 Sol 的 65 高一點點,甚至不是本代最高分。一邊是「AGI 元年」的宣言,一邊是動也不動的排行榜。
That gap isn’t a fluke. It’s a signal. For past years, the industry’s script has been simple: AI competes on intelligence, and every new model has to out-think the last one. This leaderboard says something different — the fight over IQ may already be over, ahead of schedule. What comes next is a different battlefield: not how smart the model is, but how well it understands you.
這個落差不是烏龍,是訊號。過去幾年我們習慣的劇本是:AI 產業永遠在比誰更聰明,下一代模型永遠要比上一代聰明一截。但這張排行榜說了另一件事——比智商這場仗,可能已經提前打完了。接下來要打的,換了戰場:不是模型有多聰明,是體驗有多懂你。

From Spec Wars to Experience Wars // 從規格戰到體驗戰 #
Smartphones already ran this course. In the years the spec war raged hottest, every vendor competed on pixel counts, button feel, storage, battery capacity in milliamp-hours — and Apple didn’t fight on that scoreboard. What actually won iPhone the war was iOS’s frictionless experience: unlocking, swiping, switching between apps, every motion designed to disappear, until the phone felt like an extension of the hand. By the end of the spec war, the winner wasn’t the phone with the best specs. It was the one that made people forget they were using a phone at all. This time, it’s the intelligent model’s turn to enter that same cycle.
智慧型手機走過同一條路。規格戰打得最兇的那幾年,各家比的是畫素、按鍵手感、儲存空間、電池毫安,蘋果卻沒有跟著在這張規格表上正面對決。真正讓 iPhone 勝出的,是 iOS 那套絲滑的操作體驗——解鎖、滑動、切換,每一個動作都讓人越用越無感,感覺不到「自己正在操作一支手機」,手機像是身體的延伸。規格戰打到最後,贏的不是規格最好的那支手機,是讓人最忘記自己在用手機的那支。這一次,輪到智慧模型自己走進這個週期。
I ran these numbers once before, in When Intelligence Becomes a Commodity — that piece was about the price line: as inference cost keeps compressing toward the price of electricity, intelligence eventually becomes a metered commodity. This time, the commoditization is happening on a different line — not price, but experience. And there’s already a concrete signal to point to: after Astra shipped, some called it “the best model I’ve ever used,” while others complained they had to push the prompt harder just to get the model to deliver on its actual capability. The model got smarter, and yet people had to work harder to phrase what they wanted, just to cash that intelligence out into a good result. That’s the description tax I wrote about in Product Decoder: Google’s AI Pointer — being forced to restate your intent in words before the AI will understand it. Intelligence keeps climbing. Usability hasn’t kept pace. The two lines are starting to decouple.
我在(當智慧成為大宗商品)裡算過一次帳,那篇談的是「價格」那條線——當推理成本一路壓向電費,智慧終究會變成一種按量計費的大宗商品。這次的商品化,發生在另一條線上:不是價格,是體驗。而且已經有一個很具體的訊號可以當證據——Astra 上線後,有人說它是「用過最好的模型」,也有人抱怨「要把 prompt 講得更精確,才拿得到它真正的實力」。模型明明更聰明了,人卻要更費力地把話講清楚,才能把那份聰明兌現成好結果。這正是我在〈Product Decoder:拆解 Google 的 AI Pointer〉裡寫過的「描述稅」——你被迫用文字,重新講一遍你的意圖,AI 才聽得懂。智商在往上,好不好用卻沒有跟著漲,這兩條線,已經開始脫鉤。
Two Kinds of Assistant // 兩種秘書 #
When Andy first arrived at Runway in The Devil Wears Prada, she understood English and could get things done — whatever Miranda asked, she did — but no matter how fast she moved, she was always a beat behind, and paid for it. The later Andy is different: before Miranda finishes a sentence, the coffee is already in hand; even an unpublished Harry Potter manuscript gets tracked down ahead of time. Andy’s English didn’t improve much between the two versions. What changed is this — “understanding” is just the entry ticket. Anticipating the sentence you haven’t said yet is what actually wins.
《穿著 Prada 的惡魔》裡的 Andy 剛到 Runway 時,聽得懂英文、辦得了事——Miranda 交代什麼,她做什麼,只是做得再快,也常常慢半拍,被念得狗血淋頭。後來的 Andy 不一樣:Miranda 話還沒說完,咖啡已經端在手邊;甚至連還沒出版的《哈利波特》手稿,都提前找人搞定了。兩個階段的 Andy,英文能力沒有差很多,差別在於——「聽得懂」只是門檻,「猜到你還沒說的那句話」才是勝負手。
Most of today’s LLMs are still standing where the first Andy stood: they understand, they can execute, but they’re still waiting for you to spell it out. What will actually be remembered is the system that learns to put things on the table before you open your mouth.
今天的 LLM,多半還站在第一個 Andy 的位置:聽得懂、辦得了事,但還在等你把話講清楚。真正會被記住的,是那個學會在你開口前,就把東西放到桌上的系統。

Where the Battlefield Moves // 戰場搬到哪裡 #
That’s also why the harness — the layer outside the model that handles memory, context, and the rhythm of interaction — is becoming as important as the model itself. But a harness is only a tool, not the answer on its own: a smart model doesn’t automatically mean a good experience. A harness can lift the experience, but how high it can lift it still depends on the capability of the model underneath. Put the same engine in a different chassis and the ride can feel entirely different — but no chassis, however well built, goes far without a good enough engine in the first place. As every vendor’s model IQ approaches the same ceiling, who keeps the user won’t be decided by whose model has the highest exam score. It’ll be decided by model capability and harness stacked together — whoever gets closest to guessing the sentence you haven’t said.
這也是為什麼「Harness」——模型之外,那層負責記憶、上下文、互動節奏的平台——會變得跟模型本身一樣重要。但 Harness 終究只是工具,不是答案本身:模型很聰明,不代表體驗就好;體驗要好,Harness 可以把它往上墊一截,但能墊多高,還是取決於底下那顆模型本身的能力夠不夠。同一顆引擎,裝進不同的車身,體驗可以天差地遠——但沒有一顆夠好的引擎,車身做得再精緻,也跑不遠。當各家模型的智商都逼近同一條天花板,決定使用者留在誰家的,不會只是誰的准考證分數高,而是模型能力與 Harness 疊加起來之後,誰把「猜到你還沒說的那句話」這件事做得最好。
And a better experience may not come from the AI hearing what you said this time more precisely — the real difference is whether it can hold “this time” and “three days ago” in the same frame. The preference you mentioned in passing three days ago, where your last project got stuck, the words and rhythm you tend to use — all of that is context that was never spoken out loud, yet was already there. An AI that can synthesize this request with the memory it’s accumulated over time, and derive an answer that fits you more closely, means you no longer have to spell it out — it already knows you, the way the transformed Andy did: the sentence isn’t even finished, and the thing is already on the table.
而更好的體驗,未必來自 AI 把這次使用者說的話聽得更準——真正的差異,在於它有沒有把「這一次」跟「三天前」放在一起看。三天前你隨口提過的偏好、上一個專案卡住的地方、你習慣的用詞和節奏,都是還沒被說出口、卻早就存在的上下文。能把這次的需求,跟過去累積下來的記憶綜合起來,推導出更貼近、更能滿足使用者的答案,你就不需要再把話講清楚——它已經懂你,像蛻變後的 Andy 那樣,話還沒說完,東西已經在桌上。
At this stage, AI still has to work alongside people — it isn’t replacing them and finishing the job alone. And if it’s a collaboration, experience isn’t a nice-to-have on top. It’s the variable that decides whether the collaboration holds together at all. Whichever AI feels smoother, more seamless, closer to that Jarvis-like state of “it understood before you even said it” — that’s the one more likely to keep its users and win the market’s favor. This fight will run longer, and cut closer to home, than the last one — the one over who’s smarter.
AI 在現階段,終究還是要跟人一起工作,不是把人晾在一邊、自己把事情做完。既然是合作,體驗就不是錦上添花的加分項,是決定這段合作關係撐不撐得住的那個變數。誰的 AI 用起來更順、更絲滑、越接近賈維斯那種「你還沒開口,它已經懂了」的狀態,誰就越可能被使用者留住、被市場青睞——這場仗,會比上一場「比誰更聰明」的仗,打得更久,也更貼身。

© Chung-Hao Lee. All Rights Reserved.
All content on this webpage—including but not limited to text, images, design, code, and multimedia materials—is protected under the international copyright treaties. Unauthorized reproduction, modification, distribution, public transmission, or commercial use is strictly prohibited. Legal action will be taken against infringement.
© 李崇豪。保留所有權利。
本網頁之內容(包括但不限於文字、圖片、設計、程式碼及多媒體素材)均受國際著作權條約保護。未經書面授權,嚴禁任何形式之複製、改作、散布、公開傳輸或商業利用。侵權者將依法追訴。