Skip to content
オープンソース AI

分析: OpenAI が Codex ハーネスをオープンソース化 — ハーネスとは何か、そして「オープン」は何を意味するか

コードモデルを取り巻く実行レイヤーは Apache-2.0 になりました。これにより、チームが組み込める範囲が変わりますが、支払う金額は変わりません。リリースの開発者向け解説。

DigitalNeuron Desk約7分

ひとことで言うと

AIエージェントハーネスとは何か、そしてOpenAIは実際に何をオープンソースにしたのか?

ハーネスはモデルの実行レイヤーで、タスクを保持し、長時間実行のコンテキストを管理し、ツールを呼び出し、イベントをストリーミングし、中断を許可し、承認を人間にルーティングします。OpenAIは、非対話型CLI、SDK、app-serverを含む自社のCodexハーネスをApache-2.0でリリースし、フォークして商用製品に組み込むことができるようにしました。モデルの重みはリリースされていませんが、ハーネスは依然として有料APIを呼び出すため、ライセンスコストはゼロで、実行コストはありません。

要点

  • ハーネス — ではなくモデル — は、ほとんどのチームが悪く再構築していた部分であり、現在は許容的なライセンスで利用可能です。
  • Apache-2.0 は実行コードをカバーします。重み、ホストされた推論、および商標は対象外です。
  • OpenAIは、ハーネスの変更だけでベンチマークが13.3%から38.3%に上昇したと報告しています;方向性を結果として扱い、数値はベンダー報告値として扱ってください。
  • サードパーティモデルとローカルモデルは設定可能ですが、レスポンスワイヤーフォーマット経由でのみ可能です。チャット補完エンドポイントは添付されません。
  • 予算は3層構造:ライセンス(ゼロ)、インフラ(自社所有)、トークン(使用量に応じて唯一スケールするもの)。

ハーネスが何をするか

エンジンとシャシーのメタファーは、これまでのところ正確ですが、具体的な作業を隠しています。ハーネスを分解すると、モデルが自分で行わない六つのことを同時に行っていることがわかります。

  • タスクを保持する。 モデルはリクエストに答えます。ハーネスは目標を数十回のリクエストにわたって維持し、それが達成されたかを判断します。
  • コンテキストを管理する。 長時間の実行はコンテキストウィンドウを超えます。何かが要約し、捨て、再取得し、境界を越えて推論を保持しなければなりません。粗雑にやると、 forty ステップ前に与えた制約を忘れます。
  • ツールを配布する。 ツールを定義し、引数を検証し、サンドボックス内で実行し、モデルが扱える形式で結果を戻します。
  • イベントをストリーミングする。 六分の実行は、終了時にだけではなく、リアルタイムで可読性を保つ必要があります。
  • 中断可能にする。 クリーンに停止し、状態を保持し、後で再開できるようにする — デモと実際に人が走らせるものの違いです。
  • 承認ルーティング。 自動で進められるアクションと人間の承認が必要なアクションを決め、その決定を記録します。

このリストがハーネス作業がベンチマーク数値を動かす理由です。OpenAIは、ハーネスの変更だけでモデルをARC-AGI-3で13.3%から38.3%に引き上げ、ハーネス設計によるトークン消費を六倍削減したと報告しています。ベンダーが自社リリースで報告する数値は usual なdiscountを受け、ベンチマークは狭いものです。しかし重要なのは、同じ重みでもハーネスの違いで結果が大きく変わり、残るエンジニアリングのレバレッジはハーネスにあるということです。

「オープンソース」がここに何を意味するか

このリリースは本物の許容的なライセンスですし、ラベルは依然として重要な境界を隠しています。

ハーネスに対するApache-2.0は、 fork して改変し、私的に保持し、クローズド商用製品に組み込むことができ、帰属と通知義務さえ守れば、白標記製品としても配布可能です。AGPLのようなコピーレストはトリガーしません。AGPLの実行レイヤはネットワーク経由で提供された瞬間に開示を強制するため、白標記のロードマップからは除外されます。

ライセンスがカバーしないことも明確です。重みはリリースされておらず、ハーネスはオープンでも、それに呼びかける知能はオープンではありません。Apache-2.0は商標権を許諾しないので、Codexと名乗る製品を配布するかどうかは、Codexを基にしたかどうかに別問題です。また、クライアントに対する許容的なライセンスは、接続するサービスの利用条件には関係ありません。オープンライセンスとオープン重みの違いが調達レビューに影響を与える場合は、実際に何を意味するかをレビュー前に読むべきです。

最初にチームがぶつかる制約

ハーネスは一つのプロバイダーに固定されていません。カスタムプロバイダーを指すエントリポイントが用意されており、競合他社のAPIや自社ハードウェア上で走るモデルでも利用できます。

この制約を戦略的に読むと、移植性が存在し、ベンダーが制御するワイヤーフォーマットを経由して動作します。ロックインを防ぐと考えるなら、シームは一度設定すれば済むものではなく、継続的な保守が必要だと想定すべきです。

それがどれだけ費用がかかるか

三層に分かれ、スケールするのは一つだけです。

レイヤ費用備考
ハーネスのライセンスゼロApache-2.0。 Fork して組み込むのは自由。
インフラあなたのものRust バイナリと npm パッケージ。既存の環境で走ります。
モデルトークンメータードこのリリースで変わりません。ハーネスが呼び出すだけで、何かが課金します。

OpenAIの現在のラインの公開料金は大体、入力トークンあたり5ドル、出力トークンあたり30ドルです。キャッシュされた入力は約10倍安く、ミッド・ライトティアはさらに低価格です。これだけでも半分の計算です。エージェントの実行は一回のリクエストではなく、累積コンテキストを持つ dozens のリクエストになります。これが 単価は下がる一方で請求は上がる理由です。

二つの結果が導かれます。まず、安定した接頭辞のキャッシュヒット率が請求を支配するので、ハーネスがトーンを入れ替えると、単一のリクエストが失敗することなくコストが数倍に跳ね上がります。請求書を見る前に、キャッシュヒット率を測定すべきです。第二に、席ベースのモデルは狭くなるのではなく広がるわけではありません — Codexの席は2026年6月24日から新規ビジネスワークスペースに提供されていません。現在はAPIキーが実務上の道となります。

埋め込む前に確認すべきこと

ハーネスはあなたの権限を継承します。重要なのはモデルに関するチェックリストではなく、次の点です。

  • サンドボックスの姿勢。 読み取り専用、ワークスペース書き込み、フルアクセスは明確に異なるリスク範囲です。タスクが本当に必要とする最小限の権限を選択してください。
  • 承認ポリシー。 承認を自動化すると、エージェントはレビュー対象から外れます。どのクラスのアクションは絶対に自動化しないかを決めます。
  • プロンプトインジェクション。 エージェントが読む anything には、エージェントに向けた指示が含まれる可能性があります。防御策はより狭い権限設定です。
  • 監査トレイル。 エージェントが共有リポジトリや顧客データに書き込む場合、セッションが終了しても、実行したことと許可した人間の記録が残ります。
  • 帰属義務。 Apache-2.0は、派生製品に通知ファイルを同梱することを要求します。白標記製品も対象です。

この結果はどこに着地するか

GitHub、JetBrains、CiscoはすでにCodexを自社製品に組み込んでおり、米国の税務サービス企業は約7,000件の申告書を処理し、約三分の一の準備時間を削減したと報告しています。これらの導入例のパターンは同じです:エージェントは別々のチャットウィンドウではなく、既に使っているソフトウェア内に現れます。これは実際のプラットフォーム議論であり、コーディングに限定されません。同じ実行レイヤはセキュリティトリアージ、サポート、営業、マーケティングなどにも適用できます。したがって、今後数年間で最も興味深い導入は、ターミナルを一度も開かなかった部門で起こると予想されます。

エンジニアリングチームにとって近期の質問はもっと現実的です。スタックのどこかに、期限切れのループが自社で書かれているはずです。今は、それを自社で保持すべきかどうかを問い直す良い機会です。


What shipped | コンポーネント | 内容 | | --- | --- | | codex exec | 非対話型CLI。1つの指示を受け取り、作業を実行し、構造化された結果を出力する。CIに適したエントリポイント。 | | Codex SDK | TypeScriptとPythonのバインディング。サブプロセスのstdoutを解析する代わりに、セッションオブジェクトを保持する。 | | app-server | 実行サーバ自体 — 埋め込む製品のループ。 |

Licence: Apache-2.0 (リポジトリの LICENSE ファイル、ブログ記事だけではない)。Auth: ChatGPTプランのサインインまたはAPIキー。Not included: モデルの重み、ホステッド推論、商標権。

What a harness does

The engine-and-chassis metaphor in the coverage is accurate as far as it goes, but it hides the specific work. Strip a harness down and it is doing six things at once, none of which the model does for itself:

  • Holding the task. A model answers a request. A harness keeps a goal alive across dozens of requests and decides when it is met.
  • Managing context. Long runs exceed any context window. Something has to summarise, drop, re-retrieve and preserve reasoning across the boundary. Do it crudely and the agent forgets the constraint you gave it forty steps ago.
  • Dispatching tools. Defining the tools, validating arguments, executing them under a sandbox, and feeding results back in a form the model can act on.
  • Streaming events. A run that takes six minutes has to be legible while it happens, not only at the end.
  • Being interruptible. Stopping cleanly mid-run, preserving state, resuming later — the difference between a demo and something a person will leave running.
  • Routing approvals. Which actions proceed automatically, which stop for a human, and how that decision is recorded.

That list is why harness work moves benchmark numbers at all. OpenAI reports that harness changes alone — retained reasoning and context compression — moved a model from 13.3% to 38.3% on ARC-AGI-3, with a sixfold reduction in token consumption from harness design. Vendor-reported figures on a vendor's own release deserve the usual discount, and the benchmark is a narrow one. The direction is still the finding worth keeping: the same weights produce materially different results depending on the machinery around them, and that machinery is where the remaining engineering leverage sits.

What "open source" covers here

This release is genuinely permissive, and the label still hides a boundary that matters for planning.

Apache-2.0 on the harness means you can fork it, modify it, keep the modifications private, and ship the result inside a closed commercial product — including a white‑labelled one — subject to attribution and notice obligations. There is no copyleft trigger, which is the practical difference between this and an AGPL release: an AGPL execution layer would force disclosure the moment you served it over a network, and that alone disqualifies a component from most white‑label roadmaps.

What the licence does not cover is equally clear. The weights were not released; the harness is open, the intelligence it calls is not. Trademark rights are not granted by Apache-2.0, so shipping a product that calls itself Codex is a separate question from shipping one built on Codex. And a permissive licence on a client says nothing about the terms of the service it connects to. If the distinction between an open licence and open weights is doing work in your procurement review, it is worth reading what the labels actually mean before the review, not during it.

The constraint most teams will hit first

The harness is not welded to one provider. A custom provider entry points it at any compatible endpoint — a competitor's API, a gateway, or a model running on your own hardware.

The strategic reading of that constraint is straightforward. Portability exists, and it runs through a wire format the vendor controls and has already changed once. Anyone treating this as insurance against lock‑in should assume the shim is a maintained component, not a one‑time cost.

What it costs

Three layers, and only one of them scales:

LayerCostNotes
Harness licenceZeroApache-2.0. Fork and embed freely.
InfrastructureYoursA Rust binary and an npm package. It runs where you already run things.
Model tokensMeteredUnchanged by this release. The harness makes calls; something bills for them.

Published rates for OpenAI's current line sit at roughly $5 per million input tokens and $30 per million output for the top tier, with cached input an order of magnitude cheaper, and mid and light tiers well below that. Those numbers are only half the calculation. An agent run is not one request — it is dozens, each carrying accumulated context, which is exactly the pattern that makes unit prices fall while bills rise.

Two consequences follow for anyone budgeting a deployment. First, cache hit rate on the stable prefix will dominate the bill; a harness that reshuffles the top of the prompt between turns can quietly multiply cost several times over without failing a single request. Measure it before extrapolating from a rate card. Second, the seat‑based route is narrowing rather than widening — Codex seats have not been available to new business workspaces since 24 June 2026, which leaves API keys as the practical path for most teams starting now.

Before you embed it

A harness inherits your permissions. The checklist that matters is short and mostly not about the model:

  • Sandbox posture. Read‑only, workspace‑write and full‑access are meaningfully different blast radii. Pick the narrowest one the task actually needs.
  • Approval policy. Automating approvals is where an agent stops being reviewable. Decide which classes of action are never automatic.
  • Prompt injection. Anything the agent reads may contain instructions aimed at it. The defence is a smaller permission set, not a better system prompt.
  • Audit trail. If the agent writes to a shared repository or a customer's data, the record of what it did and which human allowed it needs to survive the session.
  • Attribution obligations. Apache-2.0 requires the notice file to travel with derivative products, white‑labelled ones included.

Where this lands

GitHub, JetBrains and Cisco have already embedded Codex in their own products, and a US tax‑services firm reports processing 7,000 filings with roughly a third less preparation time. The pattern in those deployments is the same: the agent appears inside the software people already use, rather than in a separate chat window. That is the actual platform argument, and it is not confined to coding — the same execution layer works for security triage, support, sales and marketing operations, which is why the interesting adoption over the next year is likely to be in departments that have never opened a terminal.

For engineering teams the near‑term question is more prosaic. You almost certainly have a homegrown version of this loop somewhere in your stack, written under deadline. It is now worth asking whether it should keep being yours.

よくある質問

ハーネスをオープンソース化することで、Codexは自由に実行できるようになりますか?
いいえ。ハーネスはコピー、変更、埋め込みが自由です。実行ごとに、ホストされたAPIであっても、自分で運用するハードウェアであっても、トークン課金されるモデルにトークンが送信されます。
ハーネスはOpenAI以外のモデルを駆動できますか?
はい、ローカルサーバーを含む、Responsesワイヤーフォーマットを話す任意の СЭндпойнтを指すカスタムモデルプロバイダーエントリを通じて可能です。古いChat Completionsフォーマットのサポートは2026年2月に削除されたため、多くのOpenAI互換ゲートウェイには翻訳シムが必要です。
Apache-2.0 は、コピーレフト ライセンスが許可しないことを何を許可しているか?
フォークして、変更を加えて、変更を非公開に保ち、帰属表示と通知義務はありますが、ソースを公開する必要なく、クローズドまたはホワイトラベルの商用製品に組み込むことができます。
ハーネスはエージェントフレームワークと同じものですか?
重複するが、同一ではない。フレームワークは主にステップの構成方法を説明する。ハーネスはそれらを実際に実行するランタイムである — コンテキスト管理、ツールディスパッチ、ストリーミング、割り込み、承認、セッション状態。

出典

  1. Codex — repository and licence — OpenAI
  2. Codex as a platform: build on the open agent harness — OpenAI
  3. Advanced configuration — custom model providers — OpenAI
  4. API pricing — OpenAI
タグagentsharnesslicensingapache-2.0developer toolsinference cost

あわせて読みたい

オープン・ウェイトとオープンソースAI:ラベルが実際に何を意味するのか

オープンな重みとは、学習済みのモデルファイルをダウンロードし、ライセンスに従って自分で実行できるという意味です。オープンソースは、使用・学習・改変・再配布が制限なく行えるというより厳しい法的基準です。多くの広く使われるモデルはオープンな重みですが、オープンソースとは限りません。

約4分

分析: OpenAIはハーネスを開源したのではなく、モデルを開源したのではなく、そしてそれが戦略である

ハーネスはモデルを取り巻くコードです:コンテキストを組み立て、ツール呼び出しループを実行し、イベントをストリーミングし、長いセッションをコンパクトにし、不可逆なアクションを人間の承認の下で保持します。OpenAIはCodexのハーネス — codex exec, the app-server and the SDK — をApache-2.0の下でリリースしたので、どの企業でも自社のソフトウェアに同じエージェントループを組み込むことができ、その背後にあるモデルの料金を支払い続けることができます。

約7分

分析:オープンソースモデルは本当にどれだけ遅れているのか?

一般的なベンチマークでは、最先端のオープンソースモデルがフロンティアの商用モデルに近づいており、多くの日常的なタスクでは差がほとんど分からない。残るギャップは、長期的な信頼性、ツールの使用、非常に長いコンテキスト、安全性チューニング、そして自分で実行する運用作業に現れる。

約5分