Key facts
| Endpoint | POST https://plugsky.com/v1/embeddings |
| Model | plugsky-embed (family includes multilingual and NIM variants) |
| Vector size | 2,048 dimensions for the default embed model |
| Batch input | Up to 2,048 strings per request |
| Encoding | encoding_format: float or base64 |
| Dimensions | Optional truncation for faster search and lower storage |
| RAG pairing | Collections, documents and query via the RAG API |
| Product status | Live |
TL;DR
- POST /v1/embeddings with an OpenAI-compatible request and response shape.
- Batch up to 2,048 inputs per call to cut request overhead.
- Use float for numerical work and base64 to shrink payloads.
- Truncate dimensions when storage or search speed matters more than recall.
- Pair embeddings with the RAG API for chunking, storage and citations.
How it works, step by step
- Create an API key and point your OpenAI client at api.plugsky.com/v1.
- Call embeddings.create with model plugsky-embed and one or more inputs.
- Choose encoding_format: float for direct maths or base64 for transport.
- Optionally set dimensions to truncate vectors to your storage budget.
- Store vectors in pgvector, Pinecone, Qdrant, Weaviate or your own store.
- Query nearest neighbours, then feed the retrieved text into chat completions.
Original data
Try it yourself
Open the embedding model comparison →
Endpoint, request and response
The endpoint is POST https://plugsky.com/v1/embeddings. A request takes model, input (a string or an array of strings), optional encoding_format and optional dimensions. The response is the OpenAI list shape: a data array where each item carries an index and an embedding vector, plus the model name and token usage.
There is no server-side embedding cache — every call computes a fresh vector. If you re-embed the same content often, cache vectors in your own layer (Redis, your database or a vector store) keyed by content hash.
Choosing a model and dimensions
The plugsky-embed family covers general and multilingual retrieval; the live docs example prints 2,048 dimensions for the default model. Multilingual and NIM-served variants exist, so confirm the current model list and dimensions on the model catalogue before you fix a schema.
- Multilingual corpora and Arabic content: test cross-lingual retrieval, not just English similarity.
- Large archives: try truncated dimensions first — the recall loss is often acceptable for big storage savings.
- Bring-your-own embeddings: supported on collections when you supply vectors at upload time.
Embeddings inside a RAG pipeline
For most teams the fastest path is the RAG API rather than wiring embeddings by hand. Create a collection, upload documents and query it: Plugsky chunks the files (500-token chunks with 50-token overlap by default), embeds them and returns ranked chunks with source citations. Query results can then be composed into a chat completion prompt. If you already own a vector database or want custom chunking, use the embeddings endpoint directly and manage storage yourself.
Operational notes and limits
Embeddings are billed from the same flat plan as chat on self-serve plans — there is no separate per-token meter, only your plan's fair-use rate limit. Keep batch sizes reasonable: 2,048 inputs is the ceiling per request, and very large strings should be chunked upstream so one document cannot dominate a batch.
On Enterprise contracts you can bring Cohere, Voyage, BGE or custom embedding models through the model router. For residency-sensitive workloads, embeddings can run in the same region-locked data plane as your chat traffic.
Honest comparison
| Capability | Plugsky embeddings | Standalone embedding vendor | Self-hosted embedding model |
|---|---|---|---|
| API shape | OpenAI-compatible /v1/embeddings | Vendor-specific | You build the server |
| Model choice | plugsky-embed family plus BYO on Enterprise | Their models | Whatever you deploy |
| Batch input | Up to 2,048 inputs per request | Varies | Your batch logic |
| Dimension control | Optional truncation | Varies | Retrain or post-process |
| RAG integration | Same platform as chat and RAG API | Separate systems | You glue it together |
| Ops overhead | Managed | Managed | GPUs, serving, upgrades |
Frequently asked questions
Are embeddings cached?
No. Every request computes a fresh embedding, so cache vectors in your own layer if you re-embed the same content.
What languages are supported?
Plugsky's embedding family supports multilingual retrieval with strong cross-lingual behaviour, including Arabic; test on your own corpus to confirm quality.
Can I bring my own embedding model?
Yes, on Enterprise contracts. Supported options include Cohere, Voyage, BGE and custom models routed through the platform.
How do I use embeddings for RAG?
Either upload documents to the RAG API, which chunks and embeds them automatically, or embed chunks yourself and store them in your own vector database.
What vector databases work?
pgvector is the platform default. Pinecone, Qdrant and Weaviate are supported on Enterprise, and you can bring your own store when you call the embeddings endpoint directly.
How many inputs can one request contain?
Up to 2,048 strings per request; chunk larger documents before embedding them.
Can I control vector size?
Yes. The dimensions parameter truncates output vectors when you prefer smaller storage and faster search over maximum recall.
Plugsky (2026). “Embeddings API — OpenAI-Compatible Vectors”. Plugsky. Available at: https://plugsky.com/docs/embeddings (last updated 2026-09-25).