← All articles
Published 01 October 2026

Introduction

If your app still uses the legacy Completions, Chat Completions, or Realtime endpoints, now is the time to plan a controlled migration to the newer Responses and Conversations APIs. The Responses API unifies chat, structured output, tool usage and streaming into a single, item-based interface; the Conversations API provides a durable conversation object you can attach to for long-running state. This playbook walks through the practical steps, endpoint/request mappings, code snippets, and a rollout checklist so you can migrate incrementally and safely. (developers.openai.com)

1) High-level migration strategy

2) Endpoint and payload mappings (quick reference)

Legacy (Chat Completions) - Endpoint: POST /v1/chat/completions - Key fields: messages[], model, functions[], temperature - Response: choices[0].message.content, finish_reason

Responses API - Endpoint: POST /v1/responses - Key fields: input (array of items), model, text.format (for structured output), stream (true for streaming)
- Response: response.output (array of typed items) and a response.id you can reuse/store; streaming emits items as they are produced. (developers.openai.com)

Realtime (WebRTC/WebSocket) - If you use Realtime for voice or low-latency streaming, migrate session tooling to GPT‑Live or to the Responses API with delegation for tools — follow the Realtime -> GPT‑Live migration guide for event and tool mapping. (developers.openai.com)

3) Simple code examples

Note: these are minimal examples showing the mapping. Replace your API keys and model names accordingly.

Old Chat Completions (Node fetch)

// legacy chat completions
await fetch("https://api.openai.com/v1/chat/completions", {
  method: "POST",
  headers: { "Authorization": `Bearer ${API_KEY}`, "Content-Type": "application/json" },
  body: JSON.stringify({
    model: "gpt-4o",
    messages: [{role:"system", content:"You are a helpful assistant."}, {role:"user", content:"Summarize the report."}]
  })
});

New Responses API (Node fetch)

// responses API
await fetch("https://api.openai.com/v1/responses", {
  method: "POST",
  headers: { "Authorization": `Bearer ${API_KEY}`, "Content-Type": "application/json" },
  body: JSON.stringify({
    model: "gpt-5.1",
    input: [
      { type: "system", text: "You are a helpful assistant." },
      { type: "user", text: "Summarize the report." }
    ]
  })
});

Read the answer from response.output (items in the response object). (developers.openai.com)

Preserving conversation state using previous_response_id (Python)

# first call
res = requests.post("https://api.openai.com/v1/responses", json={
  "model": "gpt-5.1",
  "input": [{"type":"user","text":"Hello"}],
  "store": True
}, headers={"Authorization": f"Bearer {API_KEY}"})
response_id = res.json()["id"]

# follow-up (server-side chaining without resending full history)
res2 = requests.post("https://api.openai.com/v1/responses", json={
  "model": "gpt-5.1",
  "previous_response_id": response_id,
  "input": [{"type":"user","text":"Follow-up question"}]
}, headers={"Authorization": f"Bearer {API_KEY}"})

Using store: true and previous_response_id lets the Responses API manage context without you resending the entire history. Alternatively, use Conversations.create() to create a durable conversation object if you need a long-lived conversation resource. (developers.openai.com)

4) Structured outputs and tools

5) Streaming and realtime differences

6) Testing, validation and rollout plan

  1. Feature-flag the new path: route a small percent of traffic to Responses.
  2. Output parity tests: compare outputs of legacy calls vs Responses on a broad sample of prompts. Focus on structured outputs, finish reasons, and tool invocation behavior. (developers.openai.com)
  3. Edge-case testing: long histories, multimodal inputs, very large tool result payloads, streaming interruptions.
  4. Monitoring: add metrics for latency, error rate, token usage and semantic drift (does the new model produce different behavior?).
  5. Staged cutover: migrate internal/low-risk flows first, then user-facing ones. Keep the legacy endpoint available for rollback until you have confidence.

7) Cost, token accounting and observability

8) Tools and resources

9) Rollback plan (keep it simple)

Conclusion

Migrating to the Responses and Conversations APIs is primarily a mapping and state-design exercise: replace endpoints and field names, adopt the Responses output model and streaming semantics, and choose a state strategy (stateless history, previous_response_id chaining, or Conversations). Take an incremental, test-driven approach: inventory usages, run parallel A/B tests, adjust parsing for typed output items, migrate tool definitions to delegation/tools, and monitor costs and behavioral drift. Use the official migration guide and Conversations docs as your reference during the migration and lean on open-source helpers where helpful. (developers.openai.com)

If you want, I can: - Generate a checklist tailored to your codebase (Node/Python/Realtime) and produce concrete diffs for specific files. - Produce runnable migration scripts that replace common field names and add response_id chaining. Which would you prefer?