headroom/crates/headroom-proxy/tests/integration_conversations.rs
chopratejas 866d346bd0 fix: C4 — /v1/responses streaming + Conversations API in Rust
Phase C PR-C4 of the Realignment. Completes the Responses surface
in the Rust proxy and lays the Conversations passthrough rails.

# /v1/responses streaming pipeline

C3 left an INFO-WARN breadcrumb (`responses_streaming_passthrough_until_c4`)
on every SSE-Accept request because the byte-level SSE framer +
ResponseState machine wired by C1 was not formally tied into the
endpoint's contract. C4 closes the loop:

- Replaces the C3 fallback warn with a structured-log INFO event
  `event = "responses_streaming_pipeline_active"`. The bytes flow
  unchanged (compression on streamed response output is OUT of scope
  per live-zone-only contract); telemetry runs in the parallel task.
- New emergency-rollback toggle `--enable-responses-streaming`
  (env: `HEADROOM_PROXY_ENABLE_RESPONSES_STREAMING`, default `true`).
  When `false`, the OpenAI-Responses SSE state machine is skipped
  and `event = "responses_streaming_state_machine_skipped"` is
  emitted; bytes still pass through. Other providers' state
  machines are unaffected.
- Request-side compression (the C3 live-zone dispatcher) continues
  to run on streaming requests — `should_intercept` does not gate on
  Accept, so SSE clients get the same body compression as
  non-streaming clients.

# Conversations API surface (`/v1/conversations*`)

New `crates/headroom-proxy/src/handlers/conversations.rs` mounts
explicit axum routes (no regex per build constraints) for OpenAI's
stateful thread surface:

- `POST   /v1/conversations`
- `GET    /v1/conversations/{id}`
- `POST   /v1/conversations/{id}` (metadata update)
- `DELETE /v1/conversations/{id}`
- `POST   /v1/conversations/{id}/items`
- `GET    /v1/conversations/{id}/items`
- `GET    /v1/conversations/{id}/items/{item_id}`
- `DELETE /v1/conversations/{id}/items/{item_id}`

For PR-C4 each handler is passthrough-with-instrumentation: forward
upstream byte-equal via `forward_http`, emit
`event = "conversations_passthrough_pr_c4"` with route name +
extracted IDs. Compression of stored conversation items is C5+/B-phase
territory and explicitly NOT done here. Bodies are streamed (no
buffering) — `is_compressible_path` does not match
`/v1/conversations*` so the compression gate's else-branch streams
the body via `reqwest::Body::wrap_stream`.

New toggle `--enable-conversations-passthrough` (env:
`HEADROOM_PROXY_ENABLE_CONVERSATIONS_PASSTHROUGH`, default `true`).
When `false`, the per-route handlers are NOT mounted — requests
still reach upstream via the catch-all (no per-route logs); a
single WARN at app-build time confirms the rollback.

# Tests added

- `tests/integration_responses_streaming.rs` (4 tests):
  request bytes byte-equal upstream on streaming; client receives
  bytes that round-trip through the same SseFramer + ResponseState
  the proxy spawns; rollback path still passes bytes;
  below-threshold streaming request round-trips byte-equal.
- `tests/integration_conversations.rs` (10 tests):
  every CRUD endpoint passthrough byte-equal through wiremock;
  4xx upstream errors surface verbatim (no swallowing);
  passthrough disabled still falls through to catch-all byte-equal.
- `tests/sse_openai_responses.rs` (+2 tests):
  `chunk_boundary_invariance_pr_c4` — every single-byte split point
  produces the same final state (cache-safety streaming property);
  `minimal_upstream_response_pr_c4` — empty `[DONE]` upstream never
  panics the state machine.

# Structured-log events introduced

- `responses_streaming_pipeline_active` (INFO) — replaces the C3 WARN
- `responses_streaming_pipeline_disabled` (WARN, only when toggle off)
- `responses_streaming_state_machine_skipped` (INFO, in proxy.rs)
- `conversations_passthrough_pr_c4` (INFO, per request)
- `conversations_passthrough_disabled` (WARN, app-build time)

# Config keys introduced

- `enable_responses_streaming: bool` (default `true`)
- `enable_conversations_passthrough: bool` (default `true`)

# Verification

- `cargo build --workspace --all-features`: green
- `cargo test --workspace --all-features`: 793 + 14 + ... all green
- `cargo clippy --workspace --all-features -- -D warnings`: zero
- `cargo fmt --all --check`: green
- `make ci-precheck-rust`: green
- `make ci-precheck-python` (against main repo .venv): 176 passed

No Python files modified; Conversations compression deferred to C5+
and follow-up B-phase work that retires
`headroom/proxy/responses_converter.py`.
2026-05-03 00:14:55 -07:00

348 lines
11 KiB
Rust

//! Integration tests for the Conversations API
//! (`/v1/conversations*`) — Phase C PR-C4.
//!
//! Per spec PR-C4: the Conversations endpoints are
//! passthrough-with-instrumentation. Every request must reach
//! upstream byte-equal, and every response must reach the client
//! byte-equal. Compression of stored items is C5+/B-phase territory;
//! these tests pin the byte-fidelity contract through the entire
//! conversations CRUD surface.
mod common;
use common::start_proxy_with;
use serde_json::{json, Value};
use sha2::{Digest, Sha256};
use std::sync::{Arc, Mutex};
use wiremock::matchers::{method, path};
use wiremock::{Mock, MockServer, ResponseTemplate};
fn sha256_hex(bytes: &[u8]) -> String {
let mut hasher = Sha256::new();
hasher.update(bytes);
hasher
.finalize()
.iter()
.fold(String::with_capacity(64), |mut acc, b| {
use std::fmt::Write as _;
let _ = write!(acc, "{b:02x}");
acc
})
}
#[track_caller]
fn assert_byte_equal(inbound: &[u8], received: &[u8]) {
assert_eq!(
inbound.len(),
received.len(),
"byte length mismatch: client={}, upstream={}",
inbound.len(),
received.len()
);
assert_eq!(
sha256_hex(inbound),
sha256_hex(received),
"SHA-256 mismatch (client vs. upstream-received)"
);
}
/// Mount a capture-on-path handler that records the request body.
async fn mount_capture(
upstream: &MockServer,
method_name: &str,
path_str: &str,
response_body: &'static str,
) -> Arc<Mutex<Option<Vec<u8>>>> {
let captured: Arc<Mutex<Option<Vec<u8>>>> = Arc::new(Mutex::new(None));
let captured_clone = captured.clone();
Mock::given(method(method_name))
.and(path(path_str))
.respond_with(move |req: &wiremock::Request| {
*captured_clone.lock().unwrap() = Some(req.body.clone());
ResponseTemplate::new(200).set_body_string(response_body)
})
.mount(upstream)
.await;
captured
}
#[tokio::test]
async fn create_conversation_passthrough_byte_equal() {
let upstream = MockServer::start().await;
let captured = mount_capture(
&upstream,
"POST",
"/v1/conversations",
r#"{"id":"conv_abc","object":"conversation"}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |c| {
c.enable_conversations_passthrough = true;
})
.await;
let payload = json!({"metadata": {"user_id": "u1"}});
let body = serde_json::to_vec(&payload).unwrap();
let resp = reqwest::Client::new()
.post(format!("{}/v1/conversations", proxy.url()))
.header("content-type", "application/json")
.body(body.clone())
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let resp_bytes = resp.bytes().await.unwrap().to_vec();
let resp_parsed: Value = serde_json::from_slice(&resp_bytes).unwrap();
assert_eq!(resp_parsed["id"], json!("conv_abc"));
let got = captured.lock().unwrap().clone().expect("body captured");
assert_byte_equal(&body, &got);
proxy.shutdown().await;
}
#[tokio::test]
async fn get_conversation_passthrough() {
let upstream = MockServer::start().await;
let _captured = mount_capture(
&upstream,
"GET",
"/v1/conversations/conv_xyz",
r#"{"id":"conv_xyz","object":"conversation","metadata":{}}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |_| {}).await;
let resp = reqwest::Client::new()
.get(format!("{}/v1/conversations/conv_xyz", proxy.url()))
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let body: Value = resp.json().await.unwrap();
assert_eq!(body["id"], json!("conv_xyz"));
proxy.shutdown().await;
}
#[tokio::test]
async fn delete_conversation_passthrough() {
let upstream = MockServer::start().await;
let _captured = mount_capture(
&upstream,
"DELETE",
"/v1/conversations/conv_to_delete",
r#"{"id":"conv_to_delete","deleted":true}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |_| {}).await;
let resp = reqwest::Client::new()
.delete(format!("{}/v1/conversations/conv_to_delete", proxy.url()))
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let body: Value = resp.json().await.unwrap();
assert_eq!(body["deleted"], json!(true));
proxy.shutdown().await;
}
#[tokio::test]
async fn update_conversation_metadata_byte_equal() {
let upstream = MockServer::start().await;
let captured = mount_capture(
&upstream,
"POST",
"/v1/conversations/conv_42",
r#"{"id":"conv_42","object":"conversation"}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |_| {}).await;
let payload = json!({"metadata": {"tag": "session-2026"}});
let body = serde_json::to_vec(&payload).unwrap();
let resp = reqwest::Client::new()
.post(format!("{}/v1/conversations/conv_42", proxy.url()))
.header("content-type", "application/json")
.body(body.clone())
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let got = captured.lock().unwrap().clone().expect("body captured");
assert_byte_equal(&body, &got);
proxy.shutdown().await;
}
#[tokio::test]
async fn create_items_byte_equal_through_proxy() {
let upstream = MockServer::start().await;
let captured = mount_capture(
&upstream,
"POST",
"/v1/conversations/conv_1/items",
r#"{"object":"list","data":[{"id":"msg_1"}]}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |_| {}).await;
// Multi-item payload — the kind of body that could grow large
// in production. Bytes must round-trip identically.
let payload = json!({
"items": [
{"type": "message", "role": "user",
"content": [{"type": "input_text", "text": "first turn"}]},
{"type": "message", "role": "assistant",
"content": [{"type": "output_text", "text": "first reply"}]}
]
});
let body = serde_json::to_vec(&payload).unwrap();
let resp = reqwest::Client::new()
.post(format!("{}/v1/conversations/conv_1/items", proxy.url()))
.header("content-type", "application/json")
.body(body.clone())
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let got = captured.lock().unwrap().clone().expect("body captured");
assert_byte_equal(&body, &got);
proxy.shutdown().await;
}
#[tokio::test]
async fn list_items_passthrough() {
let upstream = MockServer::start().await;
let _captured = mount_capture(
&upstream,
"GET",
"/v1/conversations/conv_1/items",
r#"{"object":"list","data":[]}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |_| {}).await;
let resp = reqwest::Client::new()
.get(format!("{}/v1/conversations/conv_1/items", proxy.url()))
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let body: Value = resp.json().await.unwrap();
assert_eq!(body["object"], json!("list"));
proxy.shutdown().await;
}
#[tokio::test]
async fn get_item_passthrough() {
let upstream = MockServer::start().await;
let _captured = mount_capture(
&upstream,
"GET",
"/v1/conversations/conv_1/items/item_42",
r#"{"id":"item_42","type":"message"}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |_| {}).await;
let resp = reqwest::Client::new()
.get(format!(
"{}/v1/conversations/conv_1/items/item_42",
proxy.url()
))
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let body: Value = resp.json().await.unwrap();
assert_eq!(body["id"], json!("item_42"));
proxy.shutdown().await;
}
#[tokio::test]
async fn delete_item_passthrough() {
let upstream = MockServer::start().await;
let _captured = mount_capture(
&upstream,
"DELETE",
"/v1/conversations/conv_1/items/item_42",
r#"{"id":"item_42","deleted":true}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |_| {}).await;
let resp = reqwest::Client::new()
.delete(format!(
"{}/v1/conversations/conv_1/items/item_42",
proxy.url()
))
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let body: Value = resp.json().await.unwrap();
assert_eq!(body["deleted"], json!(true));
proxy.shutdown().await;
}
#[tokio::test]
async fn upstream_error_surfaces_verbatim() {
// No-silent-fallbacks: if upstream returns 4xx/5xx, we forward
// it verbatim — never swallow + return 500.
let upstream = MockServer::start().await;
Mock::given(method("GET"))
.and(path("/v1/conversations/missing"))
.respond_with(
ResponseTemplate::new(404)
.set_body_string(r#"{"error":{"message":"conversation not found"}}"#)
.insert_header("content-type", "application/json"),
)
.mount(&upstream)
.await;
let proxy = start_proxy_with(&upstream.uri(), |_| {}).await;
let resp = reqwest::Client::new()
.get(format!("{}/v1/conversations/missing", proxy.url()))
.send()
.await
.unwrap();
assert_eq!(resp.status(), 404);
let body: Value = resp.json().await.unwrap();
assert_eq!(body["error"]["message"], json!("conversation not found"));
proxy.shutdown().await;
}
#[tokio::test]
async fn passthrough_disabled_falls_through_to_catch_all() {
// When `enable_conversations_passthrough = false`, the per-route
// axum handlers are NOT mounted, but the request still reaches
// upstream via the catch-all. Bytes still round-trip equal.
let upstream = MockServer::start().await;
let captured = mount_capture(
&upstream,
"POST",
"/v1/conversations",
r#"{"id":"conv_fallthrough","object":"conversation"}"#,
)
.await;
let proxy = start_proxy_with(&upstream.uri(), |c| {
c.enable_conversations_passthrough = false;
})
.await;
let payload = json!({"metadata": {"x": 1}});
let body = serde_json::to_vec(&payload).unwrap();
let resp = reqwest::Client::new()
.post(format!("{}/v1/conversations", proxy.url()))
.header("content-type", "application/json")
.body(body.clone())
.send()
.await
.unwrap();
assert_eq!(resp.status(), 200);
let got = captured.lock().unwrap().clone().expect("body captured");
assert_byte_equal(&body, &got);
proxy.shutdown().await;
}