Measuring OpenAI Compatibility for LangGraph Endpoints

Community Article
Published September 14, 2026

An OpenAI-compatible endpoint is easy to demo and surprisingly hard to specify.

The happy path is familiar: put a FastAPI route in front of a LangGraph application, accept an OpenAI-shaped Chat Completions request, call the graph, and serialize its final message into a choices array. Add server-sent events for streaming and the official SDK can often consume it. Existing applications can then change a base URL and treat the graph as another model.

That proves transport compatibility for one request. It does not prove that the endpoint honors the request surface, preserves response fields, streams valid semantics, or survives the gateway that public clients actually use.

This article presents a conformance architecture we used for a production LangGraph-based agent endpoint. The implementation remains private, but the measurement design is general.

TL;DR

  • Pin one OpenAPI revision instead of testing a moving branch.
  • Classify every claimed request field as supported, rejected, declared no-op, or silent.
  • Validate raw responses and stream frames, then reconstruct the response the client sees.
  • Compare the wrapper with its model path, and test the direct service and public gateway separately.
  • Plant a defect in every instrument before trusting its green score.
  • Report schema validity and behavior certification separately.

The target is a supported subset that is explicit, testable, and difficult to overstate, rather than a claim about the entire OpenAI API.

I track three evidence planes separately: contract coverage describes what the endpoint claims to handle, wire conformance checks what crosses the API boundary, and behavioral conformance tests whether the request had the promised effect. A score from one plane cannot stand in for either of the others.

At release, the endpoint classified 37 of 37 request fields and showed 0 candidate failures in 18 public-path differential checks. A later experimental feature, built on that same endpoint, still shipped a stream that produced a second logical answer and attached a citation to the wrong text.

OpenAI recommends the Responses API for new projects while continuing to support Chat Completions. Chat Completions is still a useful compatibility boundary for existing clients and gateways. The same measurement principles apply to /v1/responses, but the contract changes: typed events such as response.created and response.output_item.added replace Chat Completions chunks, and the endpoint has a different pinned request schema. Validate its own event vocabulary and lifecycle rather than reusing Chat Completions assertions unchanged.

Why expose a graph as a model?

An OpenAI-compatible endpoint turns a LangGraph application into a virtual model. Existing clients can change a base URL and keep the Chat Completions request shape they already understand. A gateway can route to the agent beside ordinary model backends. Internal nodes can search, call tools, select models, retrieve state, or apply policy without requiring every client to know the graph topology.

The abstraction is valuable precisely because the two sides are different:

client contract: one request -> one response object or one ordered stream
graph runtime:   state -> nodes -> branches -> tools -> model calls -> state

The compatibility adapter is responsible for collapsing the second line into the first without losing client-visible meaning. That includes more than JSON field names. A parameter must reach the node or model it controls. A request rejection must keep its status and error shape; a model refusal remains completion content. Tool-call fragments must assemble correctly. A graph-level progress event must not become an illegal Chat Completions field. A gateway-injected usage event must not create a second answer.

At that boundary, endpoint compatibility becomes an agent-systems problem rather than a serializer task.

Compatibility already has several meanings in the ecosystem

The phrase "OpenAI-compatible" can describe several useful but different properties:

  • an OpenAI SDK can connect after changing base_url;
  • the server accepts familiar request names;
  • selected endpoints follow OpenAI response shapes;
  • tool calls and streaming work for a documented subset;
  • the full claimed surface matches both schema and behavior.

Current open-source serving documentation shows this range. vLLM's OpenAI-compatible server documentation lists supported APIs while explicitly noting an unsupported field and an ignored Chat Completions parameter. Hugging Face Text Generation Inference, now in maintenance mode, described its Messages API as fully compatible while its guidance documentation identifies a behavioral difference in tool_choice="auto".

Those disclosures are useful. They let users judge the actual contract instead of guessing from the label. The client layer adds another boundary: LangChain documents that ChatOpenAI targets the official OpenAI specification and does not preserve nonstandard response fields from third-party providers. A conformance ledger makes the standard surface and any extensions machine-testable for a custom agent endpoint.

Compatibility policy also varies by provider. Anthropic says most unsupported fields in its OpenAI compatibility layer are silently ignored. Groq documents unsupported fields that cause a 400 response. OpenRouter ignores unsupported provider parameters by default, with strict parameter support available as an opt-in routing constraint. A field ledger makes your choice explicit instead of inheriting one of those policies accidentally.

Conformance tooling is also emerging. openai-compatible-tester targets practical failures in OpenAI-style model APIs, while Open Responses offers schema-based compliance tests for the Responses API. Our narrower problem was the adapter around an agent graph: proving field handling and client-visible behavior across graph execution, model calls, and a gateway. Use an outside harness as a cross-check; keep wrapper-specific instruments for the seams only your system has.

1. Pin the contract

OpenAI publishes a machine-readable OpenAPI 3.1 specification and a human-readable Chat API reference. The repository describes endpoints, authentication, parameters, and request and response schemas, and OpenAI's official SDKs are generated from the same specification.

Do not test against a moving branch and record the result as permanent. Pin a commit and verify the downloaded artifact. Our denominator came from openai/openai-openapi@4bb21ba8:

spec_repository: openai/openai-openapi
spec_commit:     4bb21ba8e9213c3d955b69dc3f76dd7537439828
spec_file:       openapi.json
spec_version:    2.3.0
sha256:          3d6223349eadfd937624b9e6b8abf596ec2f680a1a367889cf6a6f924e568127

That snapshot contained 338 operations. Our endpoint did not claim all of them. Its declared surface centered on Chat Completions, Models, and image generation. The Chat Completions request schema supplied a concrete denominator of 37 top-level request fields.

The pin and field count are reproducible without our private harness:

curl -fsSLo openapi.json \
  https://raw.githubusercontent.com/openai/openai-openapi/4bb21ba8e9213c3d955b69dc3f76dd7537439828/openapi.json
sha256sum openapi.json

python - <<'PY'
import json

doc = json.load(open("openapi.json"))
schemas = doc["components"]["schemas"]

def properties(schema):
    if "$ref" in schema:
        return properties(schemas[schema["$ref"].split("/")[-1]])
    out = dict(schema.get("properties", {}))
    for key in ("allOf", "anyOf", "oneOf"):
        for child in schema.get(key, []):
            out.update(properties(child))
    return out

print(len(properties(schemas["CreateChatCompletionRequest"])))
PY

The expected outputs are the SHA-256 and field count shown above: 3d622334…e568127 and 37.

The pin makes three things reproducible:

  • which fields existed when the score was produced;
  • which response schemas and enums were applied;
  • what changed when the specification is refreshed later.

Treat a spec refresh as a migration. Diff the parameter and endpoint ledgers, run the full suite, and make new unclassified fields fail the gate rather than silently entering production.

Pinning does not mean treating the document as flawless. The published 3.1 artifact still uses the older nullable keyword in places, a mismatch tracked in openai-openapi issue 483. A 2020-12 JSON Schema validator treats nullable as an unknown annotation, so type: string plus nullable: true still rejects a legal null. Normalize that construct explicitly before validation and red-proof both type and enum handling for null. Version or hash the normalization rules alongside the source pin; after normalization, those rules are part of the effective contract.

2. Separate the public profile from the implementation

LangGraph is an orchestration runtime for stateful workflows, streaming, persistence, and human-in-the-loop execution. Its state and event models do not have to look like OpenAI's wire protocol. That separation is useful: the graph should represent the application, while an adapter owns the public contract.

A minimal architecture looks like this:

OpenAI client
    |
    v
request validator and compatibility adapter
    |
    v
LangGraph state / routing / tools / model calls
    |
    v
response assembler
    |
    +--> non-stream serializer
    |
    +--> SSE serializer
    |
    v
optional public gateway

Keep OpenAI-specific concerns at the edges:

  • request field names, types, enums, and refusal behavior;
  • conversion from messages and content parts into graph state;
  • tool-call and structured-output translation;
  • response fields, errors, finish reasons, usage, and annotations;
  • SSE framing and [DONE] termination.

The graph can route across tools or models internally without exposing that topology. Compatibility is judged by what the client sent and what the client received.

For open-weight models, this boundary often includes one more translation. Hugging Face chat templates convert role/content messages into the token sequence a particular model expects. If the model accepts only one initial system message, the adapter must define how the agent's base instructions and the client's system messages are merged and ordered. That is observable request behavior, so it belongs in the compatibility profile and its tests.

This edge adapter complements a native agent-serving API rather than replacing it. The LangGraph Agent Server API reference exposes assistants, threads, stateful and stateless runs, schedules, persistent storage, A2A, and MCP endpoints. Those surfaces carry lifecycle semantics that Chat Completions does not. For example, LangGraph interrupts require a checkpointer and a thread ID, and resume through a new invocation carrying a resume command. The pinned Chat Completions request has no field for that thread ID or resume command. Use the OpenAI-compatible edge when existing clients need to treat the graph as a model; retain a native agent surface when clients need to manage the graph as an agent.

3. Give every request field an explicit state

For each field in the pinned request schema, assign one state:

State Meaning Example behavior
supported The service implements the declared values and modes, with a named evidence type temperature forwarding reaches the selected model call
unsupported-rejected The service cannot honor the behavior and returns an OpenAI-shaped 400 asking for multiple choices when only one is implemented
declared-noop The field is accepted for client compatibility, validated, and documented as having no effect a routing hint that has no routing target yet
SILENT The service accepts the field, does not honor it, and neither refuses it nor declares it a no-op defect verdict emitted by the instrument

The distinction between a declared no-op and a silent field is observable contract state. Both may produce the same model answer. Only one tells operators and clients what happened.

Under our final classification rules, the baseline contained 11 silent fields. On the released checkpoint, the offline parameter ledger measured:

37 / 37 classified
26 supported
 6 declared-noop
 5 unsupported-rejected
 0 SILENT

This number did not claim 100% of OpenAI's API. It described the 37-field Chat Completions request profile at one pinned specification revision.

Here is the full release-era classification. "Supported" describes the adapter behavior measured by the stated evidence; forwarding evidence does not prove how a selected model used the value.

Release state Fields at the measured checkpoint Evidence and restrictions
Supported: request mapping model, messages Mapping and preservation at the upstream boundary
Supported: forwarding verified frequency_penalty, presence_penalty, logit_bias, max_completion_tokens, max_tokens, reasoning_effort, response_format, seed, stop, temperature, top_p, verbosity Values observed in the selected upstream request; backend effect not inferred
Supported: return behavior verified logprobs, top_logprobs, tools, tool_choice, parallel_tool_calls, functions, function_call Request forwarding plus the expected response shape or effect
Supported: server behavior verified stream, stream_options, n, modalities, audio Response behavior; n=1 supported and n>1 rejected
Declared no-op metadata, store, user, safety_identifier, prompt_cache_key, prompt_cache_options Validated and accepted with no claimed effect
Unsupported-rejected moderation, prediction, prompt_cache_retention, service_tier, web_search_options OpenAI-shaped request rejection naming the field

This table classifies field names, not every value combination. Each production ledger entry should also record its supported values, modes, nested variants, and restrictions. For a graph with several model calls, state whether a generation parameter controls the final-answer call, every internal call, or a graph-level budget.

Validate no-ops too

Accepted does not mean unvalidated. If a no-op field is an enum, reject unknown enum values. If it is an object, validate its nested shape. If it has size limits, enforce them.

Otherwise converting a refusal into a no-op can accidentally turn this:

{"prompt_cache_retention": "24h"}

into permission for all of these:

{"prompt_cache_retention": "forever"}
{"prompt_cache_retention": 7}
{"prompt_cache_retention": {}}

A compatibility adapter should be permissive about supported semantics, not about invalid wire values.

Reject unknown fields deliberately

Configure request parsing so unknown top-level fields do not disappear into an untyped dictionary. If the pinned OpenAI contract does not contain the field and the field is not a declared vendor extension, return a shaped error.

This converts specification drift into a visible failure. When a future OpenAI revision adds a field, the next pin refresh adds it to the ledger rather than allowing it to enter as an accidental no-op.

4. Validate the raw response, not only the SDK object

The schema lane sends requests to the endpoint and validates the raw JSON against the pinned response schema. For streaming, it validates every JSON-bearing SSE data event, checks framing rules, and requires one terminal [DONE].

At minimum, check:

  • HTTP status and content type;
  • required fields and literal discriminator values;
  • nested object and array types;
  • enum values such as finish reasons;
  • nullability;
  • error envelope shape;
  • every stream event's schema;
  • exactly one terminal [DONE] with no later data on successful completion.

Pseudocode:

def validate_nonstream(response, schema):
    assert response.status_code == 200
    assert response.headers["content-type"].startswith("application/json")
    payload = response.json()
    jsonschema.validate(payload, schema)


def validate_stream(response, chunk_schema):
    saw_done = False

    for event in parse_sse(response.iter_lines()):
        if event.data == "[DONE]":
            assert not saw_done
            saw_done = True
            continue

        assert not saw_done
        payload = json.loads(event.data)
        jsonschema.validate(payload, chunk_schema)

    assert saw_done

Cancellation, disconnect, and failed streams need separate termination assertions. Do not force an intentionally aborted stream to masquerade as a normal [DONE] completion.

Keep evidence payload-safe. A schema error message can contain the offending instance value, which may include user or model text. Emit the validator keyword, expected schema condition, JSON path, schema path, and observed type. Do not place the raw value in a report, metric, or CI artifact.

JSON Schema permits extra object properties unless the schema closes them. A schema-valid score therefore cannot prove that an adapter did not leak a nonstandard field. Check extras separately, either against an explicit allowlist or in the differential lane.

The official Python SDK is useful as a client check, but it does not strictly validate responses by default. Its _strict_response_validation option is underscore-prefixed and described in the client source as provisional. If you use SDK models as a second validator, call their Pydantic strict mode on the raw payload rather than relying on ordinary object construction. We verified this example with openai==3.13.0 and pydantic==2.13.5:

from openai.types.chat import ChatCompletion, ChatCompletionChunk

ChatCompletion.model_validate(nonstream_payload, strict=True)
ChatCompletionChunk.model_validate(stream_chunk, strict=True)

This is a second opinion from the client types. The pinned OpenAPI document remains the denominator because it also supplies endpoint coverage, request fields, error schemas, and version drift.

5. Measure effects, not only status codes

For each supported field, the parameter matrix needs an oracle: evidence that the value reached or changed the intended boundary.

Examples:

  • a generation control reaches the upstream model request unchanged;
  • a tool definition produces a correctly shaped tool call;
  • strict structured output returns schema-conforming JSON;
  • a routing value selects the declared backend and the response reports the public value actually used;
  • a rejected field returns the expected status, error type, parameter name, and code;
  • a declared no-op is validated and identified as a no-op in the ledger.

A 200 is not proof that a field was honored. It may be proof that the server accepted and discarded it. Name the evidence precisely: adapter forwarding verified, output constraint verified, routing behavior verified, request rejection shape verified, or declared no-op.

One practical result format is one JSON object per check:

{
  "check": "chat.temperature.nonstream",
  "path": "direct",
  "verdict": "pass",
  "evidence": {
    "request_field_observed_upstream": true,
    "response_schema": "ChatCompletion"
  }
}

Useful verdicts include:

pass
fail
unsupported-rejected
declared-noop
SILENT
reference-nonconformant

Keep declared-noop separate from pass. It counts as handled, but it must not inflate the supported count.

6. Differential-test the wrapper and the reference path

Send the same request to two paths:

  1. the underlying model route;
  2. the LangGraph-based compatibility endpoint.

Define differential equivalence from the compatibility profile. Fields compare exactly only where the adapter promises preservation. Intentional graph transformations require explicit mapping rules and separate assertions. For example, an agent may consume an upstream tool call, execute it internally, and return a final answer with a different finish reason; equality would test the wrong contract.

Use two related test shapes:

  • Adapter-preservation tests: replay a controlled upstream response and verify which fields survive, which are mapped, and which are intentionally consumed.
  • End-to-end graph tests: verify the declared public behavior instead of equality with a separately generated model response.

Generated content and truly volatile values may differ where the profile says so. Narrow preservation exemptions might include IDs, timestamps, model-name mapping, and token counts when the wrapper legitimately adds context.

The comparison should catch:

  • fields present upstream but lost by the wrapper;
  • wrong object or role values;
  • changed finish reasons;
  • missing usage or tool-call fields;
  • degraded error types, parameters, or codes;
  • extension fields leaking into the standard namespace.

The reference path is not automatically correct. Validate both sides independently against the pinned schema.

Use a truth table:

Candidate Reference Differential verdict
valid valid and contract-equivalent pass
valid invalid reference-nonconformant
invalid valid fail
invalid invalid fail
valid valid but a declared equivalence rule is violated fail

This prevents a broken upstream or gateway from becoming the authority over the published contract.

7. Test the direct service and the public gateway separately

In our deployment, the direct endpoint emitted required fields whose values were null. The public gateway removed those keys while normalizing the response. The same behavior appeared when the underlying model route passed through the gateway, locating the difference at that boundary rather than in the LangGraph service.

This is a deployment-specific observation, not a universal claim about the gateway product.

It still changes the client contract.

Run the same schema and behavior checks against:

in-process application
live direct service
public gateway route

Label every result with its path. Do not let a direct-service pass overwrite a public-path failure in the summary.

8. Validate the assembled stream

Per-event validation is necessary and insufficient.

LangGraph applications often hold, transform, or combine graph events before emitting OpenAI-shaped chunks. A public gateway may inject usage events or normalize deltas. Every event can be legal while their concatenation violates the intended behavior.

Reconstruct what a client sees:

assembled_text = ""
tool_calls = {}
extension_annotations = []
finish_reasons = []

for chunk in chunks:
    choice = chunk["choices"][0] if chunk["choices"] else None
    if not choice:
        continue

    delta = choice.get("delta", {})
    assembled_text += delta.get("content") or ""
    merge_tool_calls(tool_calls, delta.get("tool_calls"))
    extension_annotations.extend(read_declared_extension(chunk))
    finish_reasons.append(choice.get("finish_reason"))

assert_single_logical_answer(assembled_text)
assert_finish_sequence(finish_reasons)
assert_annotation_spans(assembled_text, extension_annotations)

In the pinned Chat Completions schema, annotations belongs to the non-stream response message and is not declared on ChatCompletionStreamResponseDelta. Our streamed citations put annotations on the delta and documented it as an extension; schema validation accepted it because extra properties were allowed. If we built that boundary again, we would keep the extension outside the standard delta namespace or buffer citations for the non-stream response path.

Our live failure involved a gateway-shaped usage chunk containing an empty delta choice after a held search answer. The service treated it as another answer boundary, emitted a fallback, and then emitted the real answer. Citation offsets remained valid relative to the real answer and became wrong relative to the client's assembled text. The OpenAI streaming cookbook documents the reference usage-only event with an empty choices array, but a deployed intermediary can emit a different plausible shape.

Every chunk could pass schema validation. The assembled response was still incorrect, and the shifted citation falsely attributed the fallback sentence to a source.

The regression fixture can be small and synthetic. Replay a held answer followed by the intermediary-shaped usage event, then assert the exact client-visible text:

events = [
    held_answer("A sourced fact."),
    {"choices": [{"index": 0, "delta": {}}], "usage": usage},
]

chunks = list(emit_stream(events, include_usage=True))
text = assemble_text(chunks)
assert text == "A sourced fact."
assert "couldn't attach verified citations" not in text
assert_annotation_spans(text, read_extensions(chunks))

Add gateway-realistic variants for at least:

  • usage-only events with choices: [];
  • empty-delta choices;
  • stream_options.include_usage;
  • multiple tool-call fragments;
  • empty content and role-only deltas;
  • declared annotation extensions split across model or graph events;
  • disconnect and cancellation before and after finalization.

9. Make model-to-parser protocols fail closed

Agent graphs often ask a model to emit temporary structure: tool arguments, citation markers, routing labels, or evidence tags. A parser converts that structure into the public response. If your wrapper has no temporary model-to-parser protocol, skip this section and continue with the instrument checks in §10.

Do not assume the model follows the temporary protocol because the prompt is clear.

In our search path, a fake model emitted perfect citation markers by construction. The real model sometimes malformed or omitted them. The parser correctly failed closed, but live behavior remained unusable until the model-to-parser seam was exercised directly.

For citation-bearing answers, test three separate claims:

  1. Membership: every citation points to a source returned by the search step.
  2. Span integrity: every citation offset is nonempty, in bounds, and points to the intended answer text after all markup is removed.
  3. Source support: the cited page actually supports the claim in that span.

The first two can be mechanical. The third requires a judged certification set or a reliable semantic evaluator with its own falsification evidence.

Keep the feature experimental until all three are established on the live public path.

10. Red-proof every instrument

Before accepting a green conformance score, plant a nonconformance that the instrument is supposed to catch.

Examples:

  • delete a required response field;
  • change object to an invalid stable scalar;
  • corrupt message.role;
  • drop a request field before the upstream call;
  • accept an invalid enum value;
  • remove part of the error envelope;
  • emit duplicate [DONE] markers;
  • let a usage event create a second logical answer;
  • shift a citation offset by one character;
  • place a planted secret in a schema-invalid field and prove reports never contain it.

The expected sequence is:

known-good implementation  -> pass
planted defect             -> fail for the intended reason
restored implementation    -> pass

If the planted defect stays green, the evaluator has found a defect in itself. Fix that before using its score to authorize a release.

11. Put the fast subset in the deploy gate

Conformance has three useful cadences:

Per change

Run offline request validation, response-schema checks, parameter classifications, and assembled-stream regression fixtures. These should be deterministic and fast enough to block deployment.

Per deploy

Probe the exact live revision through the direct service and public gateway. Confirm health, revision, representative non-stream and stream shapes, and any routing behavior that only exists in deployment configuration.

Scheduled

Run the broader live matrix, differential checks, official-client checks, and judged feature certifications. Publish per-check results and alert on regression.

Do not publish one blended "compatibility percentage." A useful dashboard keeps at least these separate:

declared request coverage
schema-valid direct
schema-valid public
behavior-certified direct
behavior-certified public
silent count
candidate failures (our endpoint, rather than the reference, broke the pinned contract)
reference nonconformances
instrument red-proof status

12. What the result means

At the measured release checkpoint, all 37 Chat Completions request fields were classified and none was silent. Live checks separately established the deployed revision, health, required response fields, and 0 candidate failures across 18 public-path differential checks: 15 passed and 3 found the reference path nonconformant.

Those are separate evidence planes. The 37-field ledger was an offline measurement on the released checkpoint; it was not itself a live certification of every field.

Later feature work found a live streamed-search defect despite green schema checks and offline parser tests. That did not invalidate the ledger. It demonstrated why the ledger, schema score, and behavior certification must remain separate.

The endpoint was more measurable after the work, not magically complete.

A practical build order

If you are adding an OpenAI-compatible surface to a LangGraph application, I would implement conformance in this order:

  1. Pin the upstream OpenAPI specification and record its hash.
  2. Define the endpoints and fields you claim.
  3. Generate an endpoint and parameter ledger from the pin.
  4. Reject unrecognized fields unless they are declared extensions.
  5. Validate every accepted field, including no-ops.
  6. Validate raw non-stream responses and every SSE event.
  7. Reconstruct streams and test client-visible invariants.
  8. Differential-test the wrapper and underlying model path.
  9. Repeat the checks through the public gateway.
  10. Plant one defect per evaluator and prove it turns red.
  11. Put the fast deterministic subset in the deploy gate.
  12. Refresh the pin on a schedule and treat new fields as unclassified failures.

This architecture does not require LangGraph internals to imitate OpenAI. It requires a disciplined adapter at the boundary and evidence that the adapter keeps its promises.

Compatibility should be falsifiable

Go beyond asking, "Does the OpenAI SDK return an object?" Ask:

For the exact interface we claim, which requests are honored, which are refused, which are accepted without effect, what arrives at the client, and which test has demonstrated that it can catch each class of failure?

Once those answers exist, "OpenAI-compatible" becomes a versioned engineering contract rather than a marketing adjective.

If you are publishing an OpenAI-compatible endpoint, publish the contract you actually tested.

That is the standard I would use for any LangGraph endpoint intended to be a drop-in replacement: pin the contract, classify the surface, validate the wire, test the behavior, reconstruct the stream, observe the gateway, and make every green instrument prove it knows how to turn red.

Community

Sign up or log in to comment