Docs

How do you use the Plugsky embeddings API?

The embeddings endpoint, POST /v1/embeddings, returns OpenAI-compatible vectors for semantic search, clustering and RAG. Send up to 2,048 inputs per request, choose float or base64 output, and optionally truncate dimensions. Plugsky's plugsky-embed family returns 2,048-dimension vectors and supports multilingual retrieval, including Arabic.

Key facts

EndpointPOST https://plugsky.com/v1/embeddings
Modelplugsky-embed (family includes multilingual and NIM variants)
Vector size2,048 dimensions for the default embed model
Batch inputUp to 2,048 strings per request
Encodingencoding_format: float or base64
DimensionsOptional truncation for faster search and lower storage
RAG pairingCollections, documents and query via the RAG API
Product statusLive

TL;DR

  • POST /v1/embeddings with an OpenAI-compatible request and response shape.
  • Batch up to 2,048 inputs per call to cut request overhead.
  • Use float for numerical work and base64 to shrink payloads.
  • Truncate dimensions when storage or search speed matters more than recall.
  • Pair embeddings with the RAG API for chunking, storage and citations.

How it works, step by step

  1. Create an API key and point your OpenAI client at api.plugsky.com/v1.
  2. Call embeddings.create with model plugsky-embed and one or more inputs.
  3. Choose encoding_format: float for direct maths or base64 for transport.
  4. Optionally set dimensions to truncate vectors to your storage budget.
  5. Store vectors in pgvector, Pinecone, Qdrant, Weaviate or your own store.
  6. Query nearest neighbours, then feed the retrieved text into chat completions.
1Create an API keyand point yourOpenAI client at2Callembeddings.createwith model3Chooseencoding_format:float for direct4Optionally setdimensions totruncate vectors to5Store vectors inpgvector, Pinecone,Qdrant, Weaviate or6Query nearestneighbours, thenfeed the retrieved

Original data

POST https://aEndpoint2,048 dimensioVector sizeUp to 2,048 stBatch inputencoding_formaEncodingSource: Plugsky facts table · updated 2026-09-25

Try it yourself

Open the embedding model comparison →

Endpoint, request and response

The endpoint is POST https://plugsky.com/v1/embeddings. A request takes model, input (a string or an array of strings), optional encoding_format and optional dimensions. The response is the OpenAI list shape: a data array where each item carries an index and an embedding vector, plus the model name and token usage.

There is no server-side embedding cache — every call computes a fresh vector. If you re-embed the same content often, cache vectors in your own layer (Redis, your database or a vector store) keyed by content hash.

Choosing a model and dimensions

The plugsky-embed family covers general and multilingual retrieval; the live docs example prints 2,048 dimensions for the default model. Multilingual and NIM-served variants exist, so confirm the current model list and dimensions on the model catalogue before you fix a schema.

  • Multilingual corpora and Arabic content: test cross-lingual retrieval, not just English similarity.
  • Large archives: try truncated dimensions first — the recall loss is often acceptable for big storage savings.
  • Bring-your-own embeddings: supported on collections when you supply vectors at upload time.

Embeddings inside a RAG pipeline

For most teams the fastest path is the RAG API rather than wiring embeddings by hand. Create a collection, upload documents and query it: Plugsky chunks the files (500-token chunks with 50-token overlap by default), embeds them and returns ranked chunks with source citations. Query results can then be composed into a chat completion prompt. If you already own a vector database or want custom chunking, use the embeddings endpoint directly and manage storage yourself.

Operational notes and limits

Embeddings are billed from the same flat plan as chat on self-serve plans — there is no separate per-token meter, only your plan's fair-use rate limit. Keep batch sizes reasonable: 2,048 inputs is the ceiling per request, and very large strings should be chunked upstream so one document cannot dominate a batch.

On Enterprise contracts you can bring Cohere, Voyage, BGE or custom embedding models through the model router. For residency-sensitive workloads, embeddings can run in the same region-locked data plane as your chat traffic.

Honest comparison

CapabilityPlugsky embeddingsStandalone embedding vendorSelf-hosted embedding model
API shapeOpenAI-compatible /v1/embeddingsVendor-specificYou build the server
Model choiceplugsky-embed family plus BYO on EnterpriseTheir modelsWhatever you deploy
Batch inputUp to 2,048 inputs per requestVariesYour batch logic
Dimension controlOptional truncationVariesRetrain or post-process
RAG integrationSame platform as chat and RAG APISeparate systemsYou glue it together
Ops overheadManagedManagedGPUs, serving, upgrades

Frequently asked questions

Are embeddings cached?

No. Every request computes a fresh embedding, so cache vectors in your own layer if you re-embed the same content.

What languages are supported?

Plugsky's embedding family supports multilingual retrieval with strong cross-lingual behaviour, including Arabic; test on your own corpus to confirm quality.

Can I bring my own embedding model?

Yes, on Enterprise contracts. Supported options include Cohere, Voyage, BGE and custom models routed through the platform.

How do I use embeddings for RAG?

Either upload documents to the RAG API, which chunks and embeds them automatically, or embed chunks yourself and store them in your own vector database.

What vector databases work?

pgvector is the platform default. Pinecone, Qdrant and Weaviate are supported on Enterprise, and you can bring your own store when you call the embeddings endpoint directly.

How many inputs can one request contain?

Up to 2,048 strings per request; chunk larger documents before embedding them.

Can I control vector size?

Yes. The dimensions parameter truncates output vectors when you prefer smaller storage and faster search over maximum recall.

Cite this page

Plugsky (2026). “Embeddings API — OpenAI-Compatible Vectors”. Plugsky. Available at: https://plugsky.com/docs/embeddings (last updated 2026-09-25).