A — Z Reference

Plugsky Documentation

Everything you need to ship AI to production. API reference, SDK compatibility for every language, every model, auth & quotas, deployment topologies, security controls, migration guide from OpenAI, billing, SLAs, and a real-world examples library.

Getting started

Five minutes from zero to your first model response. The Plugsky API is 100% OpenAI-compatible — change your base_url and your existing code, SDK, and prompts keep working.

Install

# Python (OpenAI SDK — works as-is)
pip install openai

# Node.js / TypeScript
npm install openai

# Go
go get github.com/sashabaranov/go-openai

# Or use raw HTTP — no SDK required

Your first API call

python
from openai import OpenAI

client = OpenAI(
    api_key="sk-live-…",                          # your Plugsky key
    base_url="https://plugsky.com/v1",         # the only line that changes
)

resp = client.chat.completions.create(
    model="plugsky-pro",
    messages=[{"role": "user", "content": "Say hello in 5 languages"}],
    max_tokens=200,
)
print(resp.choices[0].message.content)

That's it. Same request shape, same response shape, same streaming, same function-calling, same JSON mode. Full OpenAI-compatibility reference →

✓
Drop-in. Your existing OpenAI Python, Node, Go, Java, .NET, and cURL clients all work without code changes. Tools like LangChain, LlamaIndex, Vercel AI SDK, and the OpenAI Playground are first-class supported.

Core concepts

  • Model. The inference engine. plugsky-micro through plugsky-frontier, plus third-party providers behind the same API (NVIDIA NIM, OpenCode, OpenRouter, and fal.ai FLUX for images).
  • Request. A single API call. Token-counted. Priced.
  • Thread / Run. Stateful multi-turn conversation (Assistants API).
  • Tool. A function you expose to the model for function-calling.
  • Knowledge base / Vector store. Indexed documents the model can retrieve from (RAG).
  • Endpoint / Region. Where the model runs. me-central-1 (UAE), eu-west-1, us-east-1, plus customer VPC and on-prem.

Authentication

Plugsky uses bearer-token API keys. Keys are project-scoped, role-scoped, and rotatable without downtime.

API keys

Generate keys in the Dashboard → API keys. Every key is tied to your workspace and inherits its plan. Create separate keys per project or environment and revoke them independently. MCP usage can be limited with fine-grained scopes (mcp.read, mcp.write, mcp.web, mcp.media, mcp.memory, mcp.compute) — see the MCP section.

Environment variables

bash
export PLUGSKY_API_KEY="sk-live-…"
export PLUGSKY_BASE_URL="https://plugsky.com/v1"
export PLUGSKY_PROJECT="prj_8x2…"   # optional, defaults to your first project
export PLUGSKY_REGION="me-central-1"

Scopes & roles

Each key has a comma-separated scope list. Examples: chat:write,embeddings:write,files:read. Use the narrowest scope that works — production keys should never have admin.

OAuth 2.0 (3rd-party apps)

Plugsky supports OAuth 2.1 with Dynamic Client Registration (RFC 7591) and PKCE (S256) for apps that connect on behalf of a workspace. Discovery: /.well-known/oauth-authorization-server · /.well-known/oauth-protected-resource. Endpoints: POST /oauth/register, GET|POST /oauth/authorize, POST /oauth/token. Browser consent opens the dashboard (/dashboard#mcp-oauth) and minted tokens act as scoped API keys.

⚠
Rotate safely. Create a new key in the dashboard, deploy it, then revoke the old one — revocation is immediate. Manage keys →

API reference

Every endpoint, every parameter, every status code. Compatible with OpenAI's /v1/* namespace; Plugsky-specific extensions live under /v1/plugsky/*.

Chat completions

POST/v1/chat/completionslive
The primary inference endpoint. Supports streaming, function-calling, JSON mode, structured outputs, vision, and tool use. Requires model and messages; optional temperature, top_p, n, stream, stop, max_tokens, presence_penalty, frequency_penalty, tools, tool_choice, response_format, seed, user.
python
from openai import OpenAI
client = OpenAI(api_key="sk-live-…", base_url="https://plugsky.com/v1")

stream = client.chat.completions.create(
    model="plugsky-pro",
    messages=[{"role": "user", "content": "Write a haiku about GCC summers"}],
    stream=True,
    temperature=0.7,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Models

GET/v1/modelslive
List every available model with tier, provider, features and context_window. Public — no key required. This is the live source of truth for the model table below.
GET/v1/models/{id}live
Retrieve one model with full metadata: context_window, provider, upstream model, features and description. Unknown ids return 404 model_not_found.

Legacy completions

POST/v1/completionslive
Pre-chat raw completion endpoint. Backwards-compatible. Prefer /v1/chat/completions for new integrations.

Embeddings

POST/v1/embeddingslive
Vector embeddings for semantic search, RAG, clustering, and recommendations — 2048 dimensions (plugsky-embed, plugsky-embed-multilingual, plugsky-embed-nim).
python
resp = client.embeddings.create(
    model="plugsky-embed",
    input=["Plugsky is a sovereign AI cloud", "GCC banks run on us"],
)
print(len(resp.data[0].embedding), "dimensions")  # 2048

Rerank

POST/v1/reranklive
Rerank documents by relevance to a query. Send {query, documents[], top_n}; get back a scored list (embedding cosine similarity) sorted best-first. Ideal after a first-stage retrieval step.

Image generation

POST/v1/images/generationslive
OpenAI-compatible image generation powered by FLUX. Aliases: dall-e-2 → flux-schnell (fast, ~2s), dall-e-3 → flux-dev (higher quality), plugsky-flux → flux-schnell. Sizes 256×256 to 1792×1024, n up to 10, response_format url or b64_json.
POST/v1/images/editslive
Instruction-based image editing via FLUX Kontext — multipart image upload (PNG/JPEG/WebP) or image_url, plus prompt. Returns the edited image URL.

Audio (ASR / TTS / translation)

POST/v1/audio/transcriptionslive
Speech-to-text with whisper-plugsky (whisper-large-v3-turbo) — strong Arabic & multilingual support, up to 25 MB per file.
POST/v1/audio/translationslive
Multilingual audio → English text (whisper-large-v3, multipart file upload). Same shape as transcriptions: returns {"text"}.
POST/v1/audio/speechlive
Text-to-speech via ElevenLabs Turbo v2.5. OpenAI-compatible voice names (alloy, echo, fable, onyx, nova, shimmer) plus custom ElevenLabs voices. Multilingual including Arabic & English. Returns MP3.

Moderation

POST/v1/moderationslive
Classify text for harmful content with per-category flags and scores (hate, harassment, violence, sexual, self-harm) — powered by an NVIDIA content-safety model.

Files

POST/v1/fileslive
Upload files for fine-tuning, batch, or vector-store ingestion. Up to 512 MB per file. PDF, DOCX, MD, TXT, JSON, JSONL, CSV, audio, video, images.
GET/v1/fileslive
List your uploaded files.
GET/v1/files/{id}live
Retrieve file metadata (name, bytes, purpose, status).
GET/v1/files/{id}/contentlive
Download the raw file content.
DELETE/v1/files/{id}live
Delete a file.

Batch API

POST/v1/batcheslive
Create async batches from uploaded JSONL files. Jobs are recorded and status-tracked; processing runs on Plugsky managed infrastructure.
GET/v1/batcheslive
List batch jobs.
GET/v1/batches/{id}live
Retrieve a batch's status, counts and output file.
POST/v1/batches/{id}/cancellive
Cancel a batch.

Fine-tuning

POST/v1/fine_tuning/jobslive
Create a supervised fine-tuning (SFT) job. Base models: plugsky-micro, plugsky-lite, plugsky-pro. Full job API live — training runs on Plugsky managed infrastructure; contact us to enable GPU training for your workspace.
GET/v1/fine_tuning/jobslive
List fine-tuning jobs.
GET/v1/fine_tuning/jobs/{id}live
Retrieve a job (status, hyperparameters, result model).
GET/v1/fine_tuning/jobs/{id}/eventslive
List the job's event log (created, queued, …).
POST/v1/fine_tuning/jobs/{id}/cancellive
Cancel a job.

Assistants

POST/v1/assistantslive
Create an assistant (name, model, instructions) — the reusable configuration for stateful runs.
GET/v1/assistantslive
List assistants.
GET/v1/assistants/{id}live
Retrieve an assistant.
POST/v1/assistants/{id}live
Update an assistant (name, instructions, model, tools).
DELETE/v1/assistants/{id}live
Delete an assistant.

Threads & runs (executable assistants)

POST/v1/threadslive
Create a thread — a stateful conversation container.
GET/v1/threads/{id}live
Retrieve a thread (metadata). POST updates it, DELETE removes it with its messages and runs.
POST/v1/threads/{id}/messageslive
Add a message (role: user / assistant, content string or parts).
GET/v1/threads/{id}/messageslive
List thread messages in order.
POST/v1/threads/{id}/runslive
Create a run — executes the assistant (or the supplied instructions/model) against the thread and stores the reply as an assistant message. Returns status, output and usage.
GET/v1/threads/{id}/runs/{rid}live
Retrieve a run (status, output, usage).

Function calling / tools

Pass a tools array with JSON-Schema function definitions. The model returns a structured tool_calls payload you execute, then return the result. Streaming and parallel tool calls supported. Generate your function schema →

Responses

POST/v1/responseslive
OpenAI Responses-compatible endpoint. input accepts a string or message array plus instructions; responses are stored and retrievable. Function tools supported.
GET/v1/responses/{id}live
Retrieve a stored response (output items, usage).
POST/v1/responses/{id}/cancellive
Cancel a response — returns 400 for responses that already completed synchronously.

Vector stores (file search)

POST/v1/vector_storeslive
Create a vector store. GET lists your stores.
GET/v1/vector_stores/{id}live
Retrieve a store with file and chunk counts. DELETE removes it.
POST/v1/vector_stores/{id}/fileslive
Attach an uploaded file — it is chunked and embedded automatically. Body: {file_id}.
POST/v1/vector_stores/{id}/searchlive
Search a store. Body: {query, max_results} → scored text chunks, best-first.

Plugsky extensions

GET/v1/plugsky/usagelive
Per-key, per-model, per-day usage. Returns the same shape used by the Dashboard.
POST/v1/plugsky/routelive
Smart routing — pass model="auto" and Plugsky picks the best model for the prompt, your cost target, and your latency target.

MCP & agent surface

The MCP surface exposes Plugsky's tools to ChatGPT, Claude, Cursor and any MCP client. Full details in the MCP section and on the dashboard.

POST/mcplive
JSON-RPC 2.0 MCP endpoint — 20 tools, 12 prompts, 10 resources: chat, fusion, web search & fetch, transcripts, audio/vision/screen/site ingest, image generation, memory, browser sessions, code runner.
POST/oauth/registerlive
OAuth 2.1 + Dynamic Client Registration (RFC 7591/7592) with PKCE (S256). Discovery at /.well-known/oauth-authorization-server and /.well-known/oauth-protected-resource. Browser consent flow supported.
GET/api/mcp/usagelive
Per-tool MCP metering (calls, errors, latency) for your workspace. Companion export: GET /api/mcp/events.

Platform API (/api/*)

The dashboard is powered by a documented REST surface under /api/*, authenticated with your session token. Automation groups include: auth (signup, login, 2FA, refresh, OAuth), keys, usage + usage/budget, dashboard (summary, savings, onboarding), workspaces, team (members, invites), models (config, fusion presets, templates, preview, cost, compare), playground (run, search, tools, browser, conversations), skills, share, integrations, health/models and enterprise. All endpoints return JSON and share the same error schema as /v1/*.

Models

All models across free, paid, and embedding tiers — served behind one OpenAI-compatible endpoint, billed from one invoice, governed by one set of policies. The table is populated live from GET /v1/models (the single source of truth). Hover the ⓘ icon for full details.

ModelContextBest forTier
Loading models…

Smart routing & Model Fusion

Two ways to get cost savings automatically. Set model="plugsky-fusion" to use the dashboard's default chain (sequential, parallel, cost-saver — your choice). For classifier-based routing use POST /v1/plugsky/route (it defaults to model="auto"); model="auto" is not accepted by /v1/chat/completions. Typical savings: 60-80% on production traffic. See Model Fusion in the dashboard for the full UI.

python
# Option 1: use your configured Fusion chain
resp = client.chat.completions.create(
    model="plugsky-fusion",   # runs the workspace's default chain
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.model)             # which model actually answered (e.g. "plugsky-micro")

# Option 2: smart routing — dedicated endpoint
import requests
resp = requests.post(
    "https://plugsky.com/v1/plugsky/route",
    headers={"Authorization": "Bearer sk-live-…"},
    json={
        "model": "auto",           # classifier picks the best model
        "route_hint": "cost",      # cost | quality | latency
        "messages": [{"role": "user", "content": "Hello!"}],
    },
)
print(resp.json()["model"])    # the model Plugsky chose

SDKs

The OpenAI SDK works as-is — change the base URL and key. Pick your language below; official and community OpenAI SDKs plus the ecosystem frameworks are all supported.

OPENAI SDK
Python
3.8+. pip install openai. Set base_url to Plugsky. Streaming, async, type hints.
OPENAI SDK
Node.js / TypeScript
18+. npm install openai. Set baseURL to Plugsky. Browser, Bun, Deno, Cloudflare Workers.
COMMUNITY SDK
Go
1.21+. go get github.com/sashabaranov/go-openai. Set BaseURL to Plugsky. Context-aware.
COMMUNITY SDK
Java / Kotlin
JDK 11+. Maven & Gradle. Coroutines, Reactor, sync.
COMMUNITY SDK
Rust
1.74+. cargo add openai. tokio, async-std, sync.
OPENAPI
cURL & raw HTTP
Any HTTP client. OpenAPI 3.0 spec published at /v1/openapi.json.

Framework integrations

  • LangChain — ChatOpenAI(base_url="https://plugsky.com/v1")
  • LlamaIndex — OpenAI(base_url=…)
  • Vercel AI SDK — openai("…", { baseURL: "https://plugsky.com/v1" })
  • Haystack — OpenAIGenerator(api_base=…)
  • Semantic Kernel — OpenAIChatCompletion(…endpoint=…)
  • AutoGen — OpenAIWrapper(base_url=…)
  • OpenAI Playground — Custom base URL field. Tested daily.

Example: streaming with back-pressure

python
import openai
client = openai.OpenAI(api_key="sk-live-…", base_url="https://plugsky.com/v1")

with client.chat.completions.stream(
    model="plugsky-pro",
    messages=[{"role":"user","content":"Tell me a 500-word story about Plugsky"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_completion()
    print("\n--- usage:", final.usage)

Official software & installers

First-party desktop, CLI, and self-hosted apps — all open-source under MIT / Apache-2.0, branded Plugsky, and pre-configured to talk to the Plugsky API. Install with one command.

Plugsky CLI

Native terminal AI coding agent — Bun-compiled single binary with built-in agent loop, file editing, shell sandbox, MCP client, RAG indexing, and auto model routing. No provider config needed — just plugsky login.

bash
# macOS / Linux (one-line install)
curl -fsSL https://plugsky.com/install | sh

# Authenticate
plugsky login

# One-shot agentic task
plugsky "explain this project"

# Interactive TUI (REPL)
plugsky

Quick reference:

plugsky "add tests"One-shot agentic task
plugskyOpen interactive TUI
plugsky chat "hi"Streaming reply, no tools
plugsky modelsList model ladder
plugsky indexBuild RAG index
plugsky -m auto "task"Auto model routing
plugsky --resumeResume last session

Approval modes: suggest (prompts for edits + shell) · auto-edit (auto edits, prompts shell) · full-auto (no prompts)

View on GitHub → · Full user guide → · Dashboard CLI docs →

Plugsky Desktop (Jan-based)

Native AI chat client for macOS, Linux, Windows. Built on the open-source Jan project (Apache-2.0). Branded Plugsky, pre-configured with plugsky-pro as the default model.

bash
curl -fsSL https://plugsky.com/install-desktop | bash
plugsky-desktop

Available for macOS (Apple Silicon + Intel), Linux (x64), Windows (x64).

Plugsky Web (Open WebUI-based)

Self-hosted AI chat UI in your browser. Built on the open-source Open WebUI project. Docker or pip install, both backends supported. The plugsky-fusion model is pre-listed.

bash
curl -fsSL https://plugsky.com/install-web | bash
plugsky-web start
# → opens http://localhost:8080

Latest releases (Plugsky CLI v0.1.2 — 2026-07-05)

All three are MIT-licensed. See NOTICE for full upstream attribution.

Operations

Rate limits & quotas

All plans include unlimited usage within fair-use rate limits — no per-token charges, no per-request charges, no overage fees. The only limit is the per-minute request rate (RPM) for your tier. Increase limits from the Dashboard or by emailing support@plugsky.com.

PlanMonthly feeFair-use RPMConcurrentAPI keysSeats
Free$0305Unlimited0
Trial (14 days)$0605Unlimited1
Hobby$8 / mo3005Unlimited0
Starter$20 / mo60010Unlimited1
Builder$60 / mo1,50050Unlimited5
Scale$120 / mo5,000200Unlimited20
EnterpriseAnnual contractCustom (10K+)CustomUnlimitedUnlimited

Same flat rate on every model — no separate pricing for plugsky-frontier vs plugsky-micro. Every model in the catalog is included on every paid plan. Hit a 429? The response includes a Retry-After header. The SDKs retry with exponential backoff automatically.

MCP (Model Context Protocol)

Plugsky ships a remote MCP endpoint and a local stdio server, so ChatGPT, Claude, Cursor, VS Code, Windsurf and Zed can use your models, web search, media ingest, memory, browser sessions and code runner directly.

  • Remote (Streamable HTTP): https://plugsky.com/mcp with Authorization: Bearer sk-live-… — 20 tools, 12 prompts, 10 resources (chat, fusion, web search & fetch, YouTube transcripts, video/audio/screen/site ingest, image generation, memory, browser sessions, code runner).
  • OAuth 2.1 + DCR: modern clients self-register and authenticate with PKCE — no manual key needed. Discovery at /.well-known/oauth-authorization-server.
  • Fine-grained scopes: mcp.read, mcp.write, mcp.web, mcp.media, mcp.memory, mcp.compute — enforced per tool group.
  • Metering: per-tool usage at GET /api/mcp/usage (raw events: GET /api/mcp/events).
  • Local (stdio): npx -y @plugsky/mcp with PLUGSKY_API_KEY. Filter tools with --groups=chat,models,rag,usage,tools,files.
  • Ready-made configs: Dashboard → AI → MCP Servers (Claude Desktop, Cursor, Windsurf, VS Code, ChatGPT connectors).

Retries & idempotency

All POST endpoints accept an Idempotency-Key header. Re-sending the same key returns the cached result for 24 hours. This makes your POSTs safe to retry without double-billing or double-creating resources.

bash
curl -X POST https://plugsky.com/v1/chat/completions \
  -H "Authorization: Bearer $PLUGSKY_API_KEY" \
  -H "Idempotency-Key: $(uuidgen)" \
  -H "Content-Type: application/json" \
  -d '{"model":"plugsky-pro","messages":[{"role":"user","content":"hello"}]}'

Errors & status codes

CodeMeaningWhat to do
400Bad request — malformed JSON, invalid paramValidate locally before sending
401Invalid or missing API keyCheck Authorization header
403Key lacks required scopeCheck key role in Dashboard
404Model or resource not foundList /v1/models to see what's available
409Conflict (duplicate idempotency key with different body)Generate a fresh key per logical request
429Rate limit hitHonor Retry-After
413Request body too largeTrim the conversation history or split the request
402Project budget cap reachedRaise the cap in Dashboard → Usage
400 (context)Context window exceededMessage includes exact token counts — reduce history or max_tokens
500Internal errorRetry with backoff. Open a ticket if persistent.
502Upstream provider failure (after auto-failover)Retry with backoff. Status page: /status
503Upstream provider downSmart-routed requests automatically failover

All API errors use the OpenAI error schema: {"error":{"message","type","code","param"}} — JSON only, never HTML. Oversized bodies are rejected with 413.

Webhooks

Billing webhooks are live (Stripe, provisioned automatically). Platform event webhooks (batch.completed, fine_tuning.completed, quota.warning, and similar) are on the roadmap — today, poll state with GET /v1/batches/{id}, GET /v1/fine_tuning/jobs/{id}, GET /v1/plugsky/usage, or stream MCP tool events with GET /api/mcp/events.

Logs & observability

Every request is logged with: timestamp, model, tokens, latency, status, key ID, project ID, region, request ID, optional user tag. Export to Datadog, Splunk, Grafana, New Relic, OpenTelemetry, or your SIEM.

Status & SLAs

Live status: /status. Public incident history. Uptime SLAs:

  • Builder / Scale: 99.9% monthly uptime, 10% credit on miss
  • Enterprise: 99.95% monthly uptime, 25% credit, 99.99% on multi-region deployments

Deployment topologies

Same API, four deployment models. Pick one, or combine them across teams.

DEFAULT
Plugsky Cloud
Multi-tenant SaaS. me-central-1 (UAE) primary, eu-west-1, us-east-1, ap-southeast-1. Fastest to start.
ENTERPRISE
VPC deployment
Plugsky runs inside your AWS / GCP / Azure / OCI VPC. Your network, your peering, your KMS, your keys.
REGULATED
On-premises
Helm chart or air-gap installer for your data center. GPU pool: H100, H200, MI300X, or CPU-only.
SOVEREIGN
Air-gapped
No internet at all. Bundle ships on physical media. Monthly model refresh by courier.
FINANCIAL
Bring-your-own-cloud
Plugsky control plane runs in our cloud; inference runs in your cloud account. You pay your hyperscaler directly.
SAAS
White-label
Your brand on the dashboard, your domain, your colors. Resell AI under your own SKU.

Decision matrix

You need…Use
Ship in 1 day, no compliancePlugsky Cloud
Data stays in me-central-1Plugsky Cloud (region pinned)
No data leaves your AWS / Azure / GCP accountVPC deployment
SAMA / CBUAE / NSD audit trailVPC deployment + customer-managed keys
Air-gap, no internetOn-prem or air-gapped
Resell AI under your brandWhite-label

Security & compliance

Security model

  • Encryption in transit: TLS 1.3 only, HSTS, modern ciphers
  • Encryption at rest: AES-256-GCM, customer-managed keys available
  • Network isolation: per-tenant VPC, security groups, no shared kernel
  • Tenant isolation: logical (RBAC + scoped keys) or physical (your own cluster)
  • Secret hygiene: keys never logged, never returned in responses, hashed at rest with Argon2id
  • Pen tests: quarterly by HackerOne + an external firm. Reports under NDA.
  • Bug bounty: up to $25,000. security@plugsky.com

Data residency

Choose per-request or pin globally. Regions: me-central-1 (UAE — default GCC), sa-central-1 (Riyadh — Enterprise), eu-west-1 (Dublin), eu-central-1 (Frankfurt), us-east-1, us-west-2, ap-southeast-1 (Singapore). Data never leaves the pinned region. Deep dive →

Compliance & certifications

  • SOC 2 Type II — annually audited, report under NDA
  • ISO 27001 — InfoSec management
  • ISO 27701 — Privacy management
  • ISO 27017 / 27018 — Cloud & PII
  • GDPR — EU data protection
  • HIPAA — Healthcare (Enterprise + BAA)
  • PCI DSS — Card data safety (no inference on card data unless on-prem)
  • FedRAMP Moderate — In process, available on Enterprise
  • UAE PDPL — Federal Data Protection Law
  • DIFC DPL — Dubai International Financial Centre
  • SAMA CSF — Saudi Central Bank cyber framework
  • NSD — National Security Directive alignment (Enterprise on-prem)

DPA & legal

Standard Contractual Clauses (SCCs) baked into the DPA. Sub-processor list published and updated within 30 days of any change. Read the DPA →

Audit logs

Every key action — creation, rotation, scope change, deletion — is logged with actor, timestamp, IP, and request body hash. Exportable to your SIEM (Splunk, Sentinel, QRadar, Chronicle) via webhook or Kinesis/Firehose.

BYOK / HSM

Bring Your Own Key. Plugsky never sees your key — you import it into our HSM integration (AWS KMS, GCP KMS, Azure Key Vault, HashiCorp Vault, Thales Luna, AWS CloudHSM). Key rotation, revocation, and audit all yours.

PII handling

Three modes: no-PII (strict filter, PII auto-redacted), detect-only (PII tagged but not modified), passthrough (your responsibility). Default is detect-only for inference, no-PII for embeddings. Run the residency checklist →

Billing

Plans

Flat monthly fee per workspace. Every model in the catalog included on every paid plan — no per-token charges, no per-request charges, no overage fees, no surprise bills. Cancel or downgrade anytime.

PlanMonthly feeWhat's includedBest for
Free$0Free AI models (plugsky-micro, plugsky-lite), Playground + Playground Beta, 30 RPM (fair use), Unlimited API keys, 1 user (no team seats). Paid features (Agents, Tools, RAG, Marketplace, Integrations, Teams) are locked.First call, evaluation, hobby chat
Trial$0 / 14 daysFull-access trial: models up to plugsky-plus, Tools, RAG, Integrations, 60 RPM, Unlimited API keys, no card requiredEvaluating the full platform
Hobby$8 / moEvery model, 300 RPM, add team seats from $1.68/mo, Unlimited keysHobby projects, indie builders
Starter$20 / moEvery model, 600 RPM, 1 seat, Unlimited keysSolo devs, side projects
Builder$60 / moEvery model incl. vision, 1,500 RPM, 5 seats, Unlimited keysProduction teams
Scale$120 / moEvery model including frontier, 5,000 RPM, 20 seats, Unlimited keys, SSOHigh-volume SaaS
EnterpriseAnnual contractUnlimited usage, 10K+ RPM, unlimited seats, on-prem, BYOK, 99.99% SLA, DPA, BAA, dedicated engineerBanks, gov, regulated

Annual billing saves 20% (Hobby: 50%). Enterprise plans are annual contracts priced to your deployment, security, and volume requirements — mustafa@plugsky.com or mustafa@plugsky.com · WhatsApp +973 3659 9909.

Free plan — what you get

The Free plan is permanent ($0, no card) and is built for trying the platform and light use:

  • Free AI models — plugsky-micro and plugsky-lite (chat + JSON, unlimited fair use, 30 requests/min).
  • Playground and Playground Beta — full access to chat, compare models and prompt tools.
  • API access — Unlimited API keys, OpenAI-compatible endpoint, free models only.
  • Usage & Analytics, models catalog, billing and account settings.

Locked until you upgrade: paid models (plugsky-plus and above), Agent Cloud, Tools & Connectors (function calling), Knowledge / RAG, Model Fusion, Marketplace, Integrations and multi-seat Teams.

Every new account also starts with a 14-day full-access trial — no card required — which unlocks models up to plugsky-plus, Tools, RAG and Integrations so you can evaluate everything before choosing a plan.

Usage & metering

Unlimited usage on every plan. No per-token charges, no per-request charges, no overage fees. The only limit is the fair-use RPM for your tier (60 trial · 300 · 600 · 1,500 · 5,000 · 10K+ on Enterprise). Token counts are still returned in every response (usage.prompt_tokens, usage.completion_tokens) for observability — but they don't drive billing.

Invoices & taxes

Monthly billing on the 1st. PDF invoices emailed automatically. VAT-compliant for UAE (5%), KSA (15%), EU (reverse charge), and US (no sales tax on SaaS in most states). Wire transfer, ACH, SEPA, and major credit cards. Annual contracts: priced per deployment, paid upfront.

Quotas & limits

Hard $ caps at the project level prevent runaway spend. Soft warning alerts at 50%, 80%, 95%. Hard block at 100% (auto-reject 402). You can set overage_behavior=allow with a finance-approved key to allow overage up to 3× the cap with auto-billing.

Migration guides

From OpenAI

  1. Generate a Plugsky key in the Dashboard
  2. In your code, change base_url to https://plugsky.com/v1
  3. No model renames needed — OpenAI (gpt-4o, gpt-4o-mini, o1/o3), Claude (claude-3-5-sonnet, claude-3-opus), Gemini (gemini-2.5-flash) and DeepSeek model names are accepted and auto-mapped to Plugsky models (gpt-4o → plugsky-pro, gpt-4o-mini → plugsky-lite, text-embedding-3-* → plugsky-embed)
  4. Run your existing evals — should pass unchanged
  5. Switch DNS / cut over when ready

Need a per-language migration walkthrough? Full guide → or generate code for your stack →

From Anthropic

Roadmap: a native Anthropic-compatible /v1/messages endpoint is planned. Today, point the Anthropic SDK’s OpenAI-compatibility mode at https://plugsky.com/v1 — Claude model names (claude-3-5-sonnet-*, claude-3-opus-*, claude-3-5-haiku-*) are accepted and auto-mapped to Plugsky models (→ plugsky-pro / plugsky-max / plugsky-lite).

From Azure OpenAI

Point your Azure SDK at https://plugsky.com/v1 (Azure SDK supports custom endpoints). Models keep their Azure names with the azure/ prefix. Existing content filters and Azure-specific features have Plugsky equivalents — see the full compatibility matrix →

From AWS Bedrock

Use the Bedrock SDK's endpoint_url parameter. Roadmap: a native Bedrock Converse endpoint is planned. Today, use the Bedrock SDK’s OpenAI-compatible mode with endpoint_url=https://plugsky.com/v1 and model IDs such as plugsky-pro.

Reference

Glossary

TermDefinition
TokenThe atomic unit of metering. ~4 chars in English. ~1.5 chars in Arabic.
Context windowMax tokens a model can see in a single call (input + output).
EmbeddingA fixed-length vector representation of text. Used for semantic search.
RAGRetrieval-Augmented Generation: retrieve relevant docs, stuff into prompt, generate.
Function callingModel returns a structured tool call instead of free text. You execute, return result.
Fine-tuningContinue training a base model on your data. SFT (supervised) or DPO (preference).
DistillationTrain a small model to mimic a large one. Cheaper, faster, similar quality.
AgentModel + tools + memory + planning loop. Autonomous multi-step task execution.
Vector storeIndexed embeddings for fast similarity search. Plugsky includes one out of the box.
BYOKBring Your Own Key. You control the encryption keys. We can't read your data.
SovereignHosted entirely inside one jurisdiction, with no foreign access. PDPL-compliant.

Changelog

/changelog — full release history. Subscribe to RSS or the model.deprecated webhook.

Support

Frequently asked questions

Is Plugsky OpenAI-compatible?

Yes — the API mirrors OpenAI's shape (chat completions, embeddings, models), so existing SDKs work after a base URL and API key change.

How do I get an API key?

Create a free account, open Dashboard → API Keys, and create a key. The free plan includes two free models with no credit card.

Where do I report an API issue?

Include the request id from the error response and open a support ticket from the dashboard; tickets are tracked and answered by the team.