AI Infrastructure プロダクト比較
同カテゴリのプロダクトを主要指標で並べて比較しています。
ポジショニングマップ
| 項目 | Cerebrium | Milvus | Qdrant | LiteLLM | Open WebUI | BentoML | Humanloop | Together AI | Fermyon | AssemblyAI | Vapi | Fireworks AI | Weights & Biases | Qovery | LiveKit | Firecrawl | Pinecone | Koyeb | Mistral AI | Replicate | Lambda | Runpod | Baseten |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 国 | United States | United States | Germany | United States | Global | Global | United States | United States | United States | United States | United States | United States | United States | France | United States | United States | United States | France | France | United States | United States | United States | United States |
| ローンチ | 2021 | 2019 | 2021 | 2023 | — | 2019-01-01 | 2020-06-01 | 2022-01-01 | 2021-01-01 | 2019 | 2020-01-01 | 2022-01-01 | 2017-01-01 | 2020-01-01 | 2021-01-01 | 2022-01 | 2019-01 | 2021-01-01 | 2023-04-01 | 2019 | 2012-01-01 | 2021 | 2019 |
| 比較準備 | ready | ready | ready | ready | ready | ready | 要補完: 定量指標 | ready | ready | ready | ready | ready | ready | ready | ready | ready | ready | ready | ready | ready | ready | ready | ready |
| 料金 | 従量課金 / Free credits: 無料 / Usage: $0.000306/月 | 要確認 | 従量課金 / Free: 無料 | 要確認 / Open source: 無料 / Enterprise: 要確認 | 要確認 | 要確認 | 定額/従量 | 従量課金 / Serverless inference: $0.03/月 | 従量課金 / Free: 無料 | 従量課金 / Universal-2: $0.15/月 / Universal-3.5 Pro: $0.21/月 | 従量課金 / Free: 無料 | 従量課金 / tokens 0/usage | 要確認 / Free: 無料 | 要確認 / Free: 無料 | 従量課金 / Build: 無料 / Ship: $50/月 | 従量課金 / Free: 無料 / Hobby: $16/月 | 従量課金 / Starter: 無料 | 従量課金 / Free: 無料 / Pro: $29/月 | 従量課金 / Experiment: 無料 / API models: 要確認 / input tokens 0/1M tokens / output tokens 0/1M tokens | 従量課金 | 従量課金 / NVIDIA H100 SXM (1 GPU): $4.29/月 / NVIDIA B200 (16 GPU cluster): $9.86/月 | 従量課金 / Pods: 要確認 / Serverless: 要確認 | 従量課金 / Basic: 無料 / Pro: 要確認 |
| 定量指標 | 調達 $8,500,000 | 調達 $113,000,000 / users 300 / ★ 45,900 | 調達 $78,000,000 | ★ 57,541 | users 481,000 / ★ 150,000 | ★ 8,800 | 公式情報なし | 調達 $800,000,000 | ★ 6,505 | 調達 $115,000,000 | 従業員 50 | ★ 110 | ★ 11,221 | 従業員 40 | ★ 20,191 | 調達 $16,200,000 / users 1,250,000 | users 9,000 | 従業員 16 / users 100,000 | 調達 $640,000,000 / ★ 10,000 | 調達 $57,800,000 | 調達 $480,000,000 | ARR $120,000,000 / 調達 $20,000,000 / users 500,000 | 調達 $1,990,000,000 |
| 情報基準日 | 2026-08-30 | 2026-08-29 | 2026-08-28 | 2026-08-29 | 2026-08-25 | 2026-08-24 | 2026-08-24 | 2026-08-20 | 2026-08-29 | 2026-08-10 | 2026-08-09 | 2026-08-08 | 2026-08-07 | 2026-08-07 | 2026-08-04 | 2026-07-27 | 2026-07-27 | 2026-07-21 | 2026-07-13 | 2026-07-12 | 2026-07-18 | 2026-07-12 | 2026-07-12 |
| Best fit | serverless GPU/CPU / snapshotによる低遅延 | Apache 2.0のOSS / 分散構成とhardware-aware search | Rust性能 / OSSとCloud | OpenAI互換の導入容易性 / provider差分の吸収 | provider中立性 / self-hostingとoffline | OSSとmanaged inferenceの接続 / Pythonicな開発体験 | prompt・eval・observabilityを一体化 / SDKとhuman review | research-to-production / open model coverage | WebAssembly特化 / OSSとCloudの接続 | 要確認 | API-first / provider交換性 | PyTorch・Meta・Google出身の専門チーム / 推論とcustomizationの統合 | 統合workflow / SDKとdocs | Kubernetes control plane / BYOKとpolicy governance | OSSとCloudの一貫性 / 低遅延realtime stack | OSSとCloudの両輪 / AI向け出力形式 | API / docs | GPUからCPUまでの統合deploy / scale-to-zeroとglobal deployment | 研究人材と公開モデルの組み合わせ / self-hostingとAPIの選択肢 | 要確認 | instanceからSuperclusterまでのAI compute階層 / NVIDIA GPUとhigh-speed interconnectへの特化 | Pod、Serverless、Clusterを同一platformで扱える / 公式にARR $120M超と500,000 developers超を公表 | architecture別runtimeとmulti-cloud capacity management / TrussとChainsを含むPython中心のdeveloper workflow |
| Limitation | GPU供給・価格依存 / 公開売上情報が限定的 | 分散運用の学習コスト / benchmark条件への依存 | 市場競争 / embedding依存 | 互換性維持の複雑さ / self-hosting運用負担 | 運用責任 / 商用価格が要確認 | pricingと製品境界が要確認 / 複数層の学習コスト | 公開価格と財務情報が限られる / enterprise導入が重い | capital and GPU supply dependence / high operational complexity | エコシステムの成熟度 | 要確認 | 音声品質と遅延への依存 / 料金の複雑さ | GPU原価とモデル更新への依存 / 知名度の高い競合が多い | platform依存 / 複雑化 | Kubernetesの複雑さ / 価格透明性は要確認 | モデル品質を直接支配しない / 専門性の高い運用 | Web変化への依存 / infrastructure cost | 非OSS | GPU供給とMistral統合への依存 | モデル競争の速さ / 商用規模の開示が限定的 | 要確認 | GPU・電力・data centerへの大きな先行投資 / 汎用hyperscalerよりサービス範囲が狭い可能性 | GPU supplyとhardware reliabilityに依存する / workload別のcost・latency設計が利用者に残る | GPU原価と供給に左右される運用構造 / 高度なworkloadでは設計・測定の導入支援が必要 |
| 調達総額 | $8,500,000 | $113,000,000 | $78,000,000 | — | — | — | — | $800,000,000 | — | $115,000,000 | — | — | — | — | — | $16,200,000 | — | — | $640,000,000 | $57,800,000 | $480,000,000 | $20,000,000 | $1,990,000,000 |
| Bootstrapped | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ | いいえ |
| ARR | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | $120,000,000 (2026-01-20) | — |
| 従業員 | — | — | — | — | — | — | — | — | — | — | 50 | — | — | 40 | — | — | — | 16 | — | — | — | — | — |
| OSS | いいえ | はい | はい | はい | はい | はい | いいえ | はい | はい | いいえ | いいえ | いいえ | はい | いいえ | はい | はい | いいえ | いいえ | はい | はい | いいえ | いいえ | はい |
| ホスティング | Multi-cloud GPU and CPU infrastructure | Self-hosted or Zilliz Cloud | Qdrant Cloud + self-hosted | Self-hosted or LiteLLM hosted gateway | Self-hosted, Docker, Kubernetes, bare metal, or cloud VM | BentoCloud、public cloud、on-prem、Kubernetes、BYOC | Cloud platform | Together AI Native Cloud | Fermyon Cloud | Cloud API | Cloud | Fireworks AI Cloud | Cloud and self-hosted options | Customer cloud / Qovery managed cloud | LiveKit Cloud | Firecrawl Cloud | Pinecone Cloud | Koyeb bare metal infrastructure | La Plateforme API、self-hosted open-weight models、クラウド連携 | — | NVIDIA GPU instances, 1-Click Clusters, and single-tenant Superclusters with InfiniBand | Runpod globally distributed GPU cloud | Baseten Cloud、Self-hosted VPC、Hybrid。multi-cloud capacity managementで複数cloud・regionへ配置。 |
| 出典数 | 15 | 12 | 20 | 11 | 18 | 17 | 13 | 12 | 13 | 12 | 10 | 16 | 12 | 12 | 12 | 11 | 14 | 15 | 15 | 20 | 15 | 15 | 18 |
この比較の使い方
AI inference / GPU cloudを選ぶなら、モデル運用形態とworkloadから比較する。
- 既存モデルをAPIで呼びたいなら、Replicateを起点に比較する。
- 自前モデルを本番推論へ載せたいなら、Baseten、Runpod、Lambdaを比較する。
- trainingとinferenceを同じ基盤で扱いたいなら、GPU cluster、autoscaling、multi-cloud対応をsourceで確認する。
- 料金はGPU時間・ストレージ・egressの従量単位から見積もり、無料枠だけで判断しない。
「要補完」は未公開情報をゼロとして扱わないための表示。採用判断では各sourceの取得日と契約条件を再確認する。
公式source
Cerebrium
- Cerebrium developer documentation index (official)
- Cerebrium CLI GitHub (github)
- Cerebrium公式サイト (official)
- Cerebrium About (official)
Qdrant
- Qdrant GitHub repository (github)
- Qdrant official homepage (official)
- Qdrant About Us (official)
- Qdrant 2025 recap (official)
LiteLLM
- LiteLLM official documentation (official)
- LiteLLM GitHub repository (github)
- LiteLLM GitHub discussions (github)
- LiteLLM GitHub releases (github)
Open WebUI
- Open WebUI repository API metadata (github)
- Latest release API metadata (github)
- Open WebUI documentation (official)
- Features (official)
BentoML
- BentoML GitHub repository API metadata (github)
- Bento: Run Inference at Scale (official)
- BentoML documentation (official)
- BentoML documentation index (official)
Humanloop
- Humanloop GitHub organization (github)
- Humanloop Cookbook (github)
- Humanloop TypeScript SDK (github)
- Humanloop Python SDK (github)
Together AI
- Together AI documentation (official)
- Together Computer GitHub organization (github)
- Together AI official homepage (official)
- Together AI About Us and leadership (official)
Fermyon
- Fermyon developer docs (official)
- Cloud documentation (official)
- Spin documentation (official)
- Spin quickstart (official)
AssemblyAI
- AssemblyAI official homepage (official)
- AssemblyAI blog (official)
- AssemblyAI Series C announcement (official)
- Conformer-2 announcement (official)
Vapi
- Vapi documentation (official)
- Vapi Python SDK (official)
- Vapi TypeScript SDK (official)
- Vapi official homepage (official)
Fireworks AI
- Fireworks AI official homepage (official)
- Fireworks AI customers (official)
- Fireworks AI documentation (official)
- Fireworks AI Enterprise (official)
Weights & Biases
- Weights & Biases official source (official)
- Weights & Biases official source (official)
- Weights & Biases official source (official)
- Weights & Biases official source (official)
Qovery
- Qovery GitHub organization (github)
- Qovery official homepage (official)
- Qovery about and founder letter (official)
- Qovery blog (blog)
LiveKit
- 5 Qs for Russell D'Sa (news)
- LiveKit Documentation (official)
- LiveKit Agents GitHub repository (github)
- LiveKit JavaScript SDK (github)
Firecrawl
- Crawl feature documentation (official)
- Map feature documentation (official)
- Scrape feature documentation (official)
- Search feature documentation (official)
Pinecone
- Pinecone official (official)
- query-data (official)
- upsert-data (official)
- quickstart (official)
Koyeb
- Koyeb homepage (official)
- Koyeb about and history (official)
- Koyeb blog (official)
- Koyeb startup program (official)
Mistral AI
- Mistral AI La Plateforme (official)
- Mistral docs (official)
- Models overview (official)
- Mistral AI GitHub (github)
Replicate
- Story (blog)
- Cog OSS (github)
- Startup credits (blog)
- Official (official)
Lambda
- Lambda Docs (official)
- 1-Click Clusters | Lambda (official)
- Lambda Stack AI Software for Deep Learning & Machine Learning (official)
- Leadership | Lambda (official)
Runpod
- Welcome to Runpod - Runpod Documentation (official)
- Instant Clusters - Runpod Documentation (official)
- Pods Overview - Runpod Documentation (official)
- Pods pricing - Runpod Documentation (official)
Baseten
- Baseten documentation overview (official)
- How Baseten works documentation (official)
- Why Baseten documentation (official)
- Baseten autoscaling documentation (official)
データセットの利用について
構造化データの CSV や API を検討中です。どのような用途で使いたいか教えてください。