Together AI offers transparent pricing for AI infrastructure, including token-based serverless inference and dedicated GPU instances. Rates vary by model and hardware, with discounts for reserved capacity. Additional services cover code sandboxes, interpreters, and model fine-tuning.
Highlights
Token-based serverless inference pricing varies by model, with cached output rates often lower.
Dedicated single-tenant GPU instances are available for hardware like H100 and HGX B200.
Reserved GPU capacity offers discounted hourly rates for commitments from 7 to 180+ days.
Code sandboxes and interpreters are priced per hour or session, with filesystem storage fees.
Supervised fine-tuning is priced per 1M tokens, with tiers for standard and specialized needs.
LLM API PricingGPU Cloud ComputingServerless InferenceModel Fine-TuningNVIDIA H100NVIDIA HGX B200
Discover Similar Content
givemeanode.com
givemeanode
Give your LLM a supercomputer. Thousands of GPUs on the only PaaS designed for agent-driven machine learning: nodes by the minute, batch jobs, sweeps,...
featherless.ai
Serverless LLM Hosting - Featherless.ai
Instantly run any Llama model from HuggingFace without setting up any servers. Over 12,200+ models available. Starting at $10/month for unlimited acce...
opencomputer.dev
OpenComputer – The background agent cloud
Full Linux VM sandboxes with hardware-level isolation. Long-running, checkpoint and fork, resize at runtime.