free usage
credits
Serverless
From $7.50 per query hour
The easiest way to start and scale your RAG app.
Dedicated
From $0.60 per instance hour
For organizations with established workloads.
Enterprise
Custom pricing
Dedicated hardware for at-scale teams w/ advanced security needs.
How does PostgresML
pricing work?
Pay-per-use
Only use what you need, and pay as you go with no up-front costs.
Committed use discounts
Commit to certain levels of usage for a fixed monthly cost and get a discounted rate. Scale your configuration up or down at any time with the click of a button.
Serverless pricing
Storage is charged per GB/mo, and all requests by CPU or GPU millisecond of compute required to perform them.
Vector & Relational Database
| Name | Pricing |
|---|---|
| Tables & index storage | $0.25/GB per month |
| Retrieval, filtering, ranking & other queries | $7.50 per hour |
| Embeddings | Included w/ queries |
| LLMs | Included w/ queries |
| Fine tuning | Included w/ queries |
| Machine learning | Included w/ queries |
Serverless models
Serverless databases come with predefined models and a flexible pricing structure.
Embedding Models
| Name | Parameters (M) | Max input tokens | Dimensions | Strengths |
|---|---|---|---|---|
| intfloat/e5-small-v2 | 33.4 | 512 | 384 | Good quality, low latency |
| mixedbread-ai/mxbai-embed-large-v1 | 335 | 512 | 1024 | High quality, higher latency |
| Alibaba-NLP/gte-base-en-v1.5 | 137 | 8192 | 768 | Supports up to 8,000 input tokens |
| Alibaba-NLP/gte-large-en-v1.5 | 434 | 8192 | 1024 | Highest quality, 8,000 input tokens |
Instruct Models
| Name | Parameters (B) | Active Parameters (B) | Context size | Strengths |
|---|---|---|---|---|
| meta-llama/Llama-3.2-1B-Instruct | 1 | 1 | 128 | Lowest latency |
| meta-llama/Llama-3.2-3B-Instruct | 3 | 3 | 128 | Low latency |
| meta-llama/Meta-Llama-3.1-405B-Instruct | 405 | 405 | 128k | Highest quality |
| meta-llama/Meta-Llama-3.1-70B-Instruct | 70 | 70 | 128k | High quality |
| meta-llama/Meta-Llama-3.1-8B-Instruct | 8 | 8 | 128k | Low latency |
| microsoft/Phi-3-mini-128k-instruct | 3.8 | 3.8 | 128k | Low latency |
| mistralai/Mixtral-8x7B-Instruct-v0.1 | 56 | 12.9 | 32k | MOE high quality |
| mistralai/Mistral-7B-Instruct-v0.2 | 7 | 7 | 32k | Low latency |
Summarization Models
| Name | Parameters (B) | Context size | Strengths |
|---|---|---|---|
| google/pegasus-xsum | 568 | 512 | 8k |
Cost estimator
Detailed estimate 🤌
Vector Database
| unit | total | |
| import | - | |
| storage | - /month | |
| read queries | - /month | |
| write queries | - /month | |
| total | - /month |
| unit | total | |
| import | - | |
| storage | - /month | |
| read queries | - /month | |
| write queries | - /month | |
| total | - /month |
Embeddings
| ADA-V2 | total | |
| import | - | |
| read tokens | - /month | |
| write tokens | - /month | |
| total | - /month |
| units | total | |
| import | included | - |
| read tokens | included | - /month |
| write tokens | included | - /month |
| total | - /month |
Text Generation
| model | total | |
| gpt-3.5-turbo-0125 | - /month |
| model | total | |
| mixtral-8x7B | - /month |
All-in Rag
| import | - | |
| total | - /month |
| import | - | |
| total | - /month |
Frequently asked questions
💭
What does serverless mean on PostgresML?
add removeOn PostgresML you can build and scale Postgres without having to manage servers or GPUs. Your database will respond to your application’s demand automatically, and scale up or down as needed. Your charges will be based purely on your usage, and measured down to the millisecond.
Does PostgresML charge per token?
add removePostgresML does not charge per token. We charge by the amount of time a query runs. Queries that generate or process more tokens will often run longer, but queries that use smaller models will run more quickly. You’re only charged for the resources you use.
Does PostgresML charge for storage?
add removePostgresML charges $0.25 per gigabyte per month for storage. This includes fault tolerant RAID configurations for high availability as well as backups for disaster recovery.
How is PostgresML so inexpensive?
add removeOur approach to GPU memory management is inherently more efficient because at PostgresML, we move full AI capability to the database rather than moving the data to the models.
How does the cost estimator work?
add removePostgresML estimates costs based on typical workloads and real world benchmarks. Workload prediction is difficult which can make future cost estimation even harder. Please contact our team if you would like help estimating the size of your workload and the associated costs. We’re happy to help if you have any questions.
What can I do with my free credits?
add removeAnything you want with PostgresML. We’ll send you an email when your free credits expire as a reminder that you may start incurring charges in the future.
How does billing work?
add removeBy default, you will be billed monthly based on your usage. You will receive an invoice with total charges three days before your elected payment method is automatically billed. If you incur significantly increased utilization before your normal billing cycle, we will notify you with an off cycle invoice to help you control costs and maintain service.
Does PostgresML provide technical support?
add removeServerless plans have access to our community Discord. Dedicated plans offer a private Slack or MS teams channel for direct communication with our team. PostgresML provides custom SLAs for enterprise plans. Contact us for details.
Still have questions?
Contact us for more details about PostgresML plans and pricing.
Get started
with
$100 in
free credits
PostgresML
PostgresML 2024 Ⓒ All rights reserved.
This site uses cookies for usage analytics to improve our service. By continuing to browse this site, you agree to this use. See our Privacy Policy
Pinecone +
OpenAI