Docs

How do you use the Plugsky chat completions API?

Plugsky's chat completions endpoint is a drop-in replacement for OpenAI's: POST https://plugsky.com/v1/chat/completions with the same request and response shape. It supports streaming, function calling, JSON mode, vision and the plugsky-fusion router, so existing OpenAI SDK code works after changing base_url and the model name.

Key facts

EndpointPOST https://plugsky.com/v1/chat/completions
Request shapemodel, messages and optional temperature, top_p, n, stop, max_tokens, penalties, tools, response_format
Response shapeSame id, object, created, model, choices, finish_reason and usage fields as OpenAI
Streamingstream: true returns the OpenAI SSE delta format
Toolstools and tool_choice accept OpenAI-format function definitions
JSON moderesponse_format: {type: json_object}
RoutingSet model to plugsky-fusion to route per request
Product statusLive

TL;DR

  • Same request and response JSON as OpenAI — change the base URL and model.
  • Streaming, function calling, JSON mode and vision all use OpenAI shapes.
  • Model is required: pick a Plugsky model or plugsky-fusion for routing.
  • Errors use the OpenAI error schema, so existing retry logic keeps working.
  • Start free with plugsky-micro or plugsky-lite, then move up tiers.

How it works, step by step

  1. Generate an API key in the Plugsky dashboard.
  2. Point your OpenAI client base_url at https://plugsky.com/v1.
  3. Send model and messages — the minimum required fields.
  4. Add tools or response_format when you need function calling or JSON.
  5. Turn on stream: true for token-by-token SSE responses.
  6. Watch usage in the response body or the dashboard usage view.
1Generate an API keyin the Plugskydashboard.2Point your OpenAIclient base_url athttps://plugsky.com/v1.3Send model andmessages — theminimum required4Add tools orresponse_formatwhen you need5Turn on stream:true fortoken-by-token SSE6Watch usage in theresponse body orthe dashboard usage

Try it yourself

Open the OpenAI-compatible API tester →

The endpoint and request shape

The endpoint is POST https://plugsky.com/v1/chat/completions. Only model and messages are required. Messages carry one of four roles — system, user, assistant or tool — and optional fields cover sampling (temperature, top_p, seed), length (max_tokens), repetition (presence_penalty, frequency_penalty) and stopping (stop).

For generation count use n; for safety attribution use user. Because the schema matches OpenAI, prompts, SDK calls and validation code port without changes.

Streaming, tools and JSON mode

Set stream: true to receive server-sent events with the same delta structure as OpenAI. Function calling uses a tools array of JSON-Schema definitions plus tool_choice set to none, auto or a named function; the model returns structured tool calls you execute and send back as a tool-role message. JSON mode is a single field: response_format: {type: json_object}.

  • Keep streaming and tools separate per request for simpler error handling.
  • Validate JSON-mode output against your schema before writing it to storage.
  • Use seed for best-effort determinism in tests, not as a correctness guarantee.

Vision and model choice

Vision models accept content as an array of text and image_url parts, matching OpenAI's vision format. Model choice is explicit: the catalogue spans 30+ models from free chat models to frontier reasoning, plus plugsky-fusion if you want the platform to pick per request. You can also set a fixed model per API key when you want a predictable default.

Errors, limits and compatibility notes

Errors use the OpenAI schema with message, type, code and param fields, and rate limits return 429 with a Retry-After header. Request bodies are capped at 16 MB, and context overflows return a 400 that includes exact token counts.

Honest caveats: legacy /v1/completions and the stateful /v1/responses endpoint are listed as coming soon, so new builds should target chat completions. Audio, images and fine-tuning are also roadmap items — check the docs before planning those workloads.

Honest comparison

CapabilityPlugsky chat completionsTypical OpenAI-compatible providerBuilding a compatibility layer in-house
Request/response shapeIdentical to OpenAIUsually compatibleYou maintain parity
StreamingOpenAI SSE deltasCommonYou implement SSE
Function callingOpenAI tools formatVaries by providerYou adapt per model
Model choice30+ models behind one endpointVariesYou integrate each
Error schemaOpenAI error shapeVariesYou normalise errors
Migration effortChange base_url and modelOften a rewriteMonths of work

Frequently asked questions

Is the response shape exactly the same as OpenAI?

Yes. The response keeps the same id, object, created, model, choices, message, finish_reason and usage fields, so drop-in clients work unchanged.

Does streaming work the same way?

Yes. Pass stream: true and you receive SSE events with the same delta structure as OpenAI.

Which models support function calling?

Plugsky chat models support the OpenAI tools format; check the live model catalogue for the current capability matrix per model.

Can I send images?

Yes. Vision-capable models accept content as an array of text and image_url objects, using OpenAI's vision format.

How do I get JSON output?

Set response_format to {"type": "json_object"} and describe the shape you expect in the prompt, then validate the result against your schema.

What happens when I hit a rate limit?

The API returns 429 with a Retry-After header; SDKs retry with exponential backoff automatically.

Is /v1/completions available?

No. The legacy completions endpoint is listed as coming soon in the docs; use /v1/chat/completions for new integrations.

Cite this page

Plugsky (2026). “Chat Completions API Reference (OpenAI-Compat)”. Plugsky. Available at: https://plugsky.com/docs/chat-completions (last updated 2026-09-25).