コンテンツにスキップ
メインサイト ニュース コンソール

[AINews] GPT 5.6 price cut by 20%-80%: Cost of GPT 5.4 Intelligence dropped 13x in 4 months due to GPT 5.6 recursive self-optimization

· Latent Space
播客深度访谈

)、ツールの出力はデフォルトで最大10,000トークンに制限されるため、コンテキストウィンドウが不必要に拡大するのを防ぎます。 Wait, “deferred discovery” -> “遅延検出” or “必要になるまで検出・ロードを遅らせる(deferred discovery)”. “遅延検出” is fine, or “遅延発見”. Let’s use “遅延検出(deferred discovery)”.

“Prompt Caching: To avoid reprocessing the same instructions and history repeatedly, the harness treats all model-visible history as append-only. This preserves the prompt prefix, allowing the system to reuse previously computed data and maintain a high cache hit rate.” -> プロンプトキャッシュ(Prompt Caching): 同じ指示や履歴を繰り返し再処理するのを避けるため、ハーネスはモデルに表示されるすべての履歴を追記型(append-only)として扱います。これによりプロンプトのプレフィックスが保持され、システムは以前に計算されたデータを再利用して高いキャッシュヒット率を維持できます。

“Today, they proved it wasn’t just theory