feat(wrap): add headroom wrap kimi for Kimi CLI (#1426)

## Description

Adds `headroom wrap kimi`, routing Kimi CLI through the Headroom proxy.

Kimi CLI speaks an OpenAI-compatible `/chat/completions` API (its
`kosong` backend wraps `AsyncOpenAI`) and lets the base URL be
overridden via `KIMI_BASE_URL`. This wrapper points it at the local
proxy. Kimi's own OAuth bearer is forwarded upstream unchanged, so —
unlike the Copilot subscription path — no extra login or token exchange
is needed.

## Type of Change

- [x] New feature (non-breaking change that adds functionality)

## Changes Made

- `headroom/providers/kimi/`: new slice; `build_launch_env` sets
`KIMI_BASE_URL` with the per-project base-URL prefix, mirroring the
aider/vibe slices.
- `headroom/cli/wrap.py`: `kimi` subcommand; falls back to the
`kimi-cli` binary when `kimi` is not on `PATH`; `--kimi-api-url`
overrides the upstream coding endpoint (default
`https://api.kimi.com/coding/v1`).
- `tests/test_cli/test_wrap_kimi.py`: 8 tests for the wrap command.
- `README.md`: Kimi CLI row in the agent-compatibility matrix.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [ ] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ pytest tests/test_cli/test_wrap_kimi.py -q
........                                                                 [100%]
8 passed in 0.36s

$ ruff check headroom/providers/kimi headroom/cli/wrap.py tests/test_cli/test_wrap_kimi.py
All checks passed!

$ ruff format --check headroom/providers/kimi headroom/cli/wrap.py tests/test_cli/test_wrap_kimi.py
4 files already formatted
```

## Real Behavior Proof

- Environment: macOS; Kimi CLI (`kimi` / `kimi-cli`); `headroom proxy`
started with `--openai-api-url https://api.kimi.com/coding/v1`.
- Exact command / steps: start `headroom proxy --port 8787
--openai-api-url https://api.kimi.com/coding/v1`, then `curl -s
http://localhost:8787/v1/chat/completions` with the Kimi OAuth bearer
and a one-line `kimi-for-coding` chat request (`"Reply with exactly:
PONG"`).
- Observed result: `HTTP 200`; `choices[0].message.content == "PONG"`
from `kimi-for-coding`; the OAuth bearer was forwarded and accepted
upstream; the per-project path `/p/<name>/v1/chat/completions` also
returned `HTTP 200`.
- Not tested: Windows/Linux PATH discovery; the `--learn` / `--memory`
live paths beyond flag wiring.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes

## Additional Notes

- `ruff check` and `ruff format --check` pass locally; `mypy` was run on
the new `headroom/providers/kimi` slice only (clean), so the full-tree
`mypy headroom` box is left unchecked and is left to CI.
- The slice deliberately reuses `codex.proxy_base_url` and
`with_project_prefix`, identical to the aider/vibe wrappers, so
per-project savings attribution works without Kimi sending custom
headers.
- Kimi's separate search/fetch services are out of scope for
`KIMI_BASE_URL` and continue to hit Kimi directly; only the LLM
`/chat/completions` traffic is compressed.

Co-authored-by: JerrettDavis <mxjerrett@gmail.com>
This commit is contained in:
kaz 2026-07-16 05:40:58 +08:00 committed by GitHub
parent 46d4378cf7
commit eac49656a1
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
5 changed files with 344 additions and 2 deletions

View file

@ -240,10 +240,11 @@ shows an **Output Tokens Saved** card next to input compression, labelled
| Mistral Vibe | ✅ | starts proxy + launches |
| Oh My Pi | ✅ | injects config · starts proxy + launches |
| Cortex Code | Library only | 6065% savings (library mode; no `wrap`) |
| Kimi CLI | ✅ | OAuth bearer forwarded — log in once |
| ZCode | ✅ | starts proxy and prints base URLs for ZCode settings |
Any OpenAI-compatible client works via `headroom proxy`. MCP-native: `headroom mcp install`.
Undo durable wrapping with `headroom unwrap <tool>` (supports: `claude`, `copilot`, `codex`, `grok`, `omp`, `opencode`, `openclaw`, `zcode`).
Undo durable wrapping with `headroom unwrap <tool>` (supports: `claude`, `copilot`, `codex`, `grok`, `kimi`, `omp`, `opencode`, `openclaw`, `zcode`).
Registry authors can use the canonical [`server.json`](server.json) in the repo root instead of reconstructing the `headroom mcp serve` contract from prose.
### GitHub Copilot CLI subscription mode

View file

@ -120,6 +120,7 @@ from headroom.providers.grok_build.config import (
inject_grok_provider_config,
restore_grok_provider_config,
)
from headroom.providers.kimi import build_launch_env as _build_kimi_launch_env
from headroom.providers.mistral_vibe import build_launch_env as _build_mistral_vibe_launch_env
from headroom.providers.omp import build_launch_env as _build_omp_launch_env
from headroom.providers.omp import inject_models_override as _inject_omp_models_override
@ -1832,7 +1833,7 @@ def _codex_toml_value(value: Any) -> str:
return json.dumps(value)
if isinstance(value, bool):
return "true" if value else "false"
if isinstance(value, (int, float)):
if isinstance(value, int | float):
return str(value)
raise TypeError(f"unsupported Codex config override: {type(value).__name__}")
@ -5709,6 +5710,93 @@ def vibe(
)
# =============================================================================
# Kimi CLI
# =============================================================================
@wrap.command(context_settings={"ignore_unknown_options": True})
@click.option(
"--port", "-p", default=8787, type=click.IntRange(1, 65535), help="Proxy port (default: 8787)"
)
@click.option(
"--no-context-tool",
"--no-rtk",
"no_rtk",
is_flag=True,
help="Skip CLI context-tool setup (no effect for kimi)",
)
@click.option(
"--code-graph",
is_flag=True,
help="Enable code graph indexing via codebase-memory-mcp (optional)",
)
@click.option("--no-proxy", is_flag=True, help="Skip proxy startup (use existing proxy)")
@click.option("--learn", is_flag=True, help="Enable live traffic learning")
@click.option("--memory", is_flag=True, help="Enable persistent cross-session memory")
@click.option("--verbose", "-v", is_flag=True, help="Verbose output")
@click.option(
"--kimi-api-url",
default="https://api.kimi.com/coding/v1",
help="Upstream Kimi coding endpoint (default: https://api.kimi.com/coding/v1)",
)
@click.option("--prepare-only", is_flag=True, hidden=True)
@click.argument("kimi_args", nargs=-1, type=click.UNPROCESSED)
def kimi(
port: int,
no_rtk: bool,
code_graph: bool,
no_proxy: bool,
learn: bool,
memory: bool,
verbose: bool,
kimi_api_url: str,
prepare_only: bool,
kimi_args: tuple,
) -> None:
"""Launch Kimi CLI through Headroom proxy.
\b
Sets KIMI_BASE_URL to route Kimi's OpenAI-compatible /chat/completions
traffic through Headroom. Kimi's own OAuth bearer is forwarded upstream,
so no extra login is required run `kimi` once to authenticate first.
\b
Examples:
headroom wrap kimi # Start proxy + kimi
headroom wrap kimi -- -m kimi-for-coding # Pass args to kimi
headroom wrap kimi --port 9999 # Custom proxy port
headroom wrap kimi --kimi-api-url https://api.moonshot.ai/v1
"""
if prepare_only:
return
kimi_bin = shutil.which("kimi") or shutil.which("kimi-cli")
if not kimi_bin:
click.echo("Error: 'kimi' (or 'kimi-cli') not found in PATH.")
click.echo("Install Kimi CLI: https://github.com/MoonshotAI/kimi-cli")
raise SystemExit(1)
env, env_vars_display = _build_kimi_launch_env(
port, os.environ, project=_project_name_from_cwd()
)
_launch_tool(
binary=kimi_bin,
args=kimi_args,
env=env,
port=port,
no_proxy=no_proxy,
tool_label="KIMI",
env_vars_display=env_vars_display,
learn=learn,
memory=memory,
agent_type="kimi",
code_graph=code_graph,
openai_api_url=kimi_api_url,
)
# =============================================================================
# Grok CLI
# =============================================================================

View file

@ -0,0 +1,5 @@
"""Kimi CLI-specific provider helpers."""
from .runtime import build_launch_env
__all__ = ["build_launch_env"]

View file

@ -0,0 +1,34 @@
"""Runtime helpers for Kimi CLI integrations."""
from __future__ import annotations
import os
from collections.abc import Mapping
from headroom.providers.codex import proxy_base_url as codex_proxy_base_url
from headroom.proxy.project_context import with_project_prefix
def build_launch_env(
port: int,
environ: Mapping[str, str] | None = None,
project: str | None = None,
) -> tuple[dict[str, str], list[str]]:
"""Build environment variables for Kimi CLI through the local proxy.
Kimi CLI (``kimi`` / ``kimi-cli``) talks to its managed coding endpoint with
an OpenAI-compatible ``/chat/completions`` client (``kosong``'s ``Kimi``
provider wraps ``AsyncOpenAI``). Its base URL is overridable via the
``KIMI_BASE_URL`` environment variable, so we point it at the local proxy.
The proxy forwards the request including Kimi's own OAuth ``Authorization``
bearer (passthrough auth mode) to the real upstream configured by
``--openai-api-url`` (``https://api.kimi.com/coding/v1``).
``project`` (the wrap launch directory) is encoded as a ``/p/<name>``
base-URL prefix because the Kimi base-URL override cannot carry custom
headers; the proxy strips it and attributes savings per project.
"""
env = dict(environ or os.environ)
base_url = with_project_prefix(codex_proxy_base_url(port), project)
env["KIMI_BASE_URL"] = base_url
return env, [f"KIMI_BASE_URL={base_url}"]

View file

@ -0,0 +1,214 @@
"""Tests for `headroom wrap kimi` command."""
from __future__ import annotations
from pathlib import Path
from typing import Any
from unittest.mock import patch
import pytest
from click.testing import CliRunner
from headroom.cli import wrap as wrap_mod
from headroom.cli.main import main
@pytest.fixture
def runner() -> CliRunner:
return CliRunner()
def test_wrap_kimi_launch(
runner: CliRunner,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Kimi launches with correct configuration."""
monkeypatch.chdir(tmp_path)
monkeypatch.delenv("HEADROOM_CONTEXT_TOOL", raising=False)
captured: dict[str, Any] = {}
def fake_launch_tool(**kwargs: Any) -> None: # noqa: ANN003
captured.update(kwargs)
with patch.object(wrap_mod.shutil, "which", return_value="kimi"):
with patch.object(wrap_mod, "_launch_tool", side_effect=fake_launch_tool):
with patch.object(wrap_mod, "_project_name_from_cwd", return_value=None):
result = runner.invoke(
main, ["wrap", "kimi", "--port", "9000", "--", "-m", "kimi-for-coding"]
)
assert result.exit_code == 0, result.output
env = captured["env"]
assert isinstance(env, dict)
assert captured["tool_label"] == "KIMI"
assert captured["agent_type"] == "kimi"
assert captured["args"] == ("-m", "kimi-for-coding")
assert captured["openai_api_url"] == "https://api.kimi.com/coding/v1"
assert env["KIMI_BASE_URL"] == "http://127.0.0.1:9000/v1"
def test_wrap_kimi_with_project_name(
runner: CliRunner,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Project name is encoded in KIMI_BASE_URL when run from a project directory."""
project_dir = tmp_path / "my-project"
project_dir.mkdir()
monkeypatch.chdir(project_dir)
monkeypatch.delenv("HEADROOM_CONTEXT_TOOL", raising=False)
captured: dict[str, Any] = {}
def fake_launch_tool(**kwargs: Any) -> None: # noqa: ANN003
captured.update(kwargs)
with patch.object(wrap_mod.shutil, "which", return_value="kimi"):
with patch.object(wrap_mod, "_launch_tool", side_effect=fake_launch_tool):
result = runner.invoke(main, ["wrap", "kimi", "--port", "7000"])
assert result.exit_code == 0, result.output
env = captured["env"]
assert env["KIMI_BASE_URL"] == "http://127.0.0.1:7000/p/my-project/v1"
def test_wrap_kimi_cli_fallback(
runner: CliRunner,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Falls back to the `kimi-cli` binary when `kimi` is not on PATH."""
monkeypatch.chdir(tmp_path)
monkeypatch.delenv("HEADROOM_CONTEXT_TOOL", raising=False)
captured: dict[str, Any] = {}
def fake_launch_tool(**kwargs: Any) -> None: # noqa: ANN003
captured.update(kwargs)
def fake_which(name: str) -> str | None:
return "/usr/local/bin/kimi-cli" if name == "kimi-cli" else None
with patch.object(wrap_mod.shutil, "which", side_effect=fake_which):
with patch.object(wrap_mod, "_launch_tool", side_effect=fake_launch_tool):
with patch.object(wrap_mod, "_project_name_from_cwd", return_value=None):
result = runner.invoke(main, ["wrap", "kimi"])
assert result.exit_code == 0, result.output
assert captured["binary"] == "/usr/local/bin/kimi-cli"
def test_wrap_kimi_not_found(
runner: CliRunner,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Error message when neither kimi nor kimi-cli is found."""
monkeypatch.chdir(tmp_path)
monkeypatch.delenv("HEADROOM_CONTEXT_TOOL", raising=False)
with patch.object(wrap_mod.shutil, "which", return_value=None):
result = runner.invoke(main, ["wrap", "kimi"])
assert result.exit_code == 1
assert "Error: 'kimi' (or 'kimi-cli') not found in PATH" in result.output
assert "https://github.com/MoonshotAI/kimi-cli" in result.output
def test_wrap_kimi_custom_port(
runner: CliRunner,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Custom --port is passed to _launch_tool and appears in KIMI_BASE_URL."""
monkeypatch.chdir(tmp_path)
monkeypatch.delenv("HEADROOM_CONTEXT_TOOL", raising=False)
captured: dict[str, Any] = {}
def fake_launch_tool(**kwargs: Any) -> None: # noqa: ANN003
captured.update(kwargs)
with patch.object(wrap_mod.shutil, "which", return_value="kimi"):
with patch.object(wrap_mod, "_launch_tool", side_effect=fake_launch_tool):
with patch.object(wrap_mod, "_project_name_from_cwd", return_value=None):
result = runner.invoke(main, ["wrap", "kimi", "--port", "9999"])
assert result.exit_code == 0, result.output
assert captured["port"] == 9999
assert captured["env"]["KIMI_BASE_URL"] == "http://127.0.0.1:9999/v1"
def test_wrap_kimi_custom_api_url(
runner: CliRunner,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""--kimi-api-url overrides the upstream endpoint passed to _launch_tool."""
monkeypatch.chdir(tmp_path)
monkeypatch.delenv("HEADROOM_CONTEXT_TOOL", raising=False)
captured: dict[str, Any] = {}
def fake_launch_tool(**kwargs: Any) -> None: # noqa: ANN003
captured.update(kwargs)
with patch.object(wrap_mod.shutil, "which", return_value="kimi"):
with patch.object(wrap_mod, "_launch_tool", side_effect=fake_launch_tool):
with patch.object(wrap_mod, "_project_name_from_cwd", return_value=None):
result = runner.invoke(
main,
["wrap", "kimi", "--kimi-api-url", "https://api.moonshot.ai/v1"],
)
assert result.exit_code == 0, result.output
assert captured["openai_api_url"] == "https://api.moonshot.ai/v1"
def test_wrap_kimi_no_proxy(
runner: CliRunner,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""--no-proxy flag prevents proxy startup."""
monkeypatch.chdir(tmp_path)
monkeypatch.delenv("HEADROOM_CONTEXT_TOOL", raising=False)
captured: dict[str, Any] = {}
def fake_launch_tool(**kwargs: Any) -> None: # noqa: ANN003
captured.update(kwargs)
with patch.object(wrap_mod.shutil, "which", return_value="kimi"):
with patch.object(wrap_mod, "_launch_tool", side_effect=fake_launch_tool):
with patch.object(wrap_mod, "_project_name_from_cwd", return_value=None):
result = runner.invoke(main, ["wrap", "kimi", "--no-proxy"])
assert result.exit_code == 0, result.output
assert captured["no_proxy"] is True
def test_wrap_kimi_learn_memory(
runner: CliRunner,
tmp_path: Path,
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""--learn and --memory flags are passed to _launch_tool."""
monkeypatch.chdir(tmp_path)
monkeypatch.delenv("HEADROOM_CONTEXT_TOOL", raising=False)
captured: dict[str, Any] = {}
def fake_launch_tool(**kwargs: Any) -> None: # noqa: ANN003
captured.update(kwargs)
with patch.object(wrap_mod.shutil, "which", return_value="kimi"):
with patch.object(wrap_mod, "_launch_tool", side_effect=fake_launch_tool):
with patch.object(wrap_mod, "_project_name_from_cwd", return_value=None):
result = runner.invoke(main, ["wrap", "kimi", "--learn", "--memory"])
assert result.exit_code == 0, result.output
assert captured["learn"] is True
assert captured["memory"] is True