← News
By Leila Haddad
— Solutions Architect, Enterprise
·
· INSIGHT
Anthropic का Usage-Based Pricing Shift: Enterprises को क्या renegotiate करना है
Anthropic ने Claude Enterprise को एक low base seat fee plus API rates पर uncapped usage पर restructure किया है। 150 seats से ऊपर customers renewal पर transition करते हैं। Sign करने से पहले Procurement को tokens, caching, और batch discounts के around spend models फिर से build करने चाहिए।
Anthropic की 2026 enterprise pricing में क्या बदला
अप्रैल 2026 में, Anthropic ने confirm किया जो The Information ने पहले report किया था: Claude Enterprise bundled-token, flat per-seat pricing से एक model पर move हो गया है जो एक low base seat fee को standard API rates पर billed usage के साथ combine करता है। Anthropic की published structure के अनुसार, नया Enterprise seat roughly $20 per user per month पर बैठता है, और "आप usage के लिए billing disable नहीं कर सकते।" Older premium और standard seat tiers (पहले लगभग $200 और $40 per user per month bundled token allowances के साथ) renewal पर new model पर transition होते हैं।
Change initially 150 या ज़्यादा seats वाले deployments को target करता है। Smaller customers अभी legacy pricing पर रहते हैं, लेकिन direction of travel clear है। Anthropic Claude Code adoption और agentic workloads द्वारा driven एक compute crunch cite करता है। PYMNTS, The Register, और ITBrief से industry reporting इसे frontier-model enterprise contracts के लिए flat-fee era के end के रूप में frame करती है। Procurement teams के लिए, practical effect यह है कि line item predictable subscription से variable consumption पर shift हो जाता है, seat fee total cost का एक small fraction बनने के साथ।
Per-seat बनाम per-token: नया math कैसे काम करता है
Legacy Premium plan के तहत, $200 per user पर एक 200-seat deployment bundled token quotas के साथ एक predictable $480,000 annual commitment yield करता था। New model के तहत, same 200 seats केवल $48,000 base fees produce करती हैं; बाकी जो भी वो users actually API list prices पर consume करते हैं: Sonnet 4.6 लगभग $3 input और $15 output per million tokens पर, Opus 4.7 materially higher।
implicator.ai और The Register द्वारा cited licensing analysts estimate करते हैं कि heavy users total spend double या triple देखेंगे। एक single Claude Code engineer agentic loops run करते हुए per day कई million tokens burn कर सकता है; एक platform team में multiply करें और variable line base को dwarf कर देती है। Math अब तीन numbers पर hinges करता है जो procurement शायद ही पहले track करता था: per active user per day average tokens, output-to-input ratio (output 5x ज़्यादा expensive है), और Haiku, Sonnet, और Opus के बीच model mix। Finance teams जो पहले AI spend को एक per-seat SaaS line forecast करती थीं उन्हें consumption telemetry के around rebuild करना चाहिए, headcount के नहीं।
Cache discounts और Message Batches: use करने योग्य cost levers
दो Anthropic-native discounts new contract shape के तहत "nice to have" से load-bearing पर move होते हैं। Message Batches API asynchronous requests को 24 hours के अंदर standard token prices के exactly 50% पर process करता है, सभी Claude models पर input और output दोनों पर applied। जो कुछ भी एक day की latency tolerate कर सके — overnight document classification, evaluation runs, bulk summarization, embeddings refresh — इसके through routed होना चाहिए।
Prompt caching दूसरा lever है। Anthropic के published multipliers 5-minute cache writes को base input price के 1.25x पर, 1-hour writes को 2x पर, और crucially, cache reads को 0.1x — cached input tokens पर एक 90% discount पर रखते हैं। दो discounts stack करते हैं। एक long system prompt या document context जो Batches के माध्यम से repeatedly hit होता है, cached portion के लिए list price के लगभग 5% पर land कर सकता है। RAG pipelines और long-context agent workflows के लिए, prompts को deliberately structure करना ताकि stable prefix cacheable हो, अब एक procurement-visible decision है, सिर्फ एक engineering optimization नहीं।
New tiers के तहत annual spend को modeling
एक defensible 2026 spend model को per user cohort तीन inputs चाहिए: per month active days, per active day consumed tokens, और उस traffic का share जो caching या batch discounts के लिए eligible है। एक useful baseline: एक typical Claude-in-IDE developer reportedly per active day 1-3 million tokens consume करता है; Deep-Research-style flows running करने वाला एक research analyst 200-500k के closer है; Haiku पर एक customer-support agent 100k के नीचे बैठ सकता है।
तीन scenarios बनाएं — base, stretch, और adversarial — और हर एक को new seat plus usage formula के against pressure-test करें। Adversarial case zero caching adoption, zero batch routing, Opus default selection, और 30% month-over-month token growth (reported demand curves के साथ consistent) assume करना चाहिए। वो ceiling renewal conversation में लाने वाला number है। Anthropic का Compliance API और audit log access, Enterprise seat में included, वो हैं जो आपको monthly instrument करने चाहिए; sign करने से पहले contract में कम से कम 13 months token-level usage data retain करने का requirement bake करें।
OpenAI का counter-positioning और migration incentives
OpenAI ने अब तक ChatGPT Enterprise के लिए एक per-seat posture hold की है — published rates लगभग $45-$75 per user per month range में land करते हैं, एक 150-seat minimum और annual prepay के साथ। 2 अप्रैल 2026 तक, OpenAI ने standard seats और Codex-only seats के साथ flexible pricing introduce किया, plus credit pools जो Deep Research, Thinking models, image generation, और per-seat caps के ऊपर Codex को additional access unlock करते हैं। Structure entry tier पर अभी भी SaaS-shaped feel करती है, हालांकि credit-pool mechanic advanced features के लिए quietly usage-based dynamics import करता है।
यह एक near-term arbitrage create करता है। Q3 और Q4 2026 में renewal cycles run करने वाली Procurement teams OpenAI sales motions को Anthropic के variable model के against predictability को explicitly pitch करते देखेंगी। उन quotes को discipline के साथ treat करें: confirm करें कि कौन से features per-seat envelope के अंदर fall करते हैं, कौन से credit draw-down require करते हैं, और pool dry होने के बाद overage rate कैसा दिखता है। Realistic outcome single-vendor migration नहीं है, बल्कि किसी भी side पर hard caps के साथ एक dual-sourced posture है।
Hedge: एक single-vendor contract पर BYOK gateways layering
किसी भी single vendor के pricing shift के against सबसे clear structural hedge है model contract को application से decouple करना। एक gateway layer जो आपकी API keys hold करता है और providers में requests route करता है, model choice को एक procurement commitment के बजाय एक runtime decision बनाता है। जब Anthropic rates raise करे, एक discount tier expire करे, या contract के बीच एक model SKU deprecate करे, आप application code touch किए बिना Sonnet-class traffic को एक comparable competitor पर reroute करते हैं।
Operational requirements unglamorous लेकिन specific हैं: per-request provider selection, prompt-cache awareness ताकि reroute करते समय आप 90% discount न खोएं, per workspace या cost center token-level usage attribution, और pure pass-through billing ताकि gateway खुद vendor list prices के ऊपर एक margin layer add न करे। इस category के Platforms — osFoundry BYOK pass-through और no per-seat fees के around built एक example है — Anthropic, OpenAI, और open-weights providers के सामने simultaneously बैठते हैं। Point vendor abandonment नहीं है; यह एक position से renegotiate करने के option को preserve करना है जहां switching cost hours में measured है, quarters में नहीं।
अपने next renewal के लिए Action items
2026 renewal sign करने से पहले table पर चार artifacts लाएं। पहला, user cohort, model, और cache hit rate द्वारा broken down एक 13-month token consumption baseline — Anthropic का Compliance API आपको raw data देता है; अब instrument करें भले ही renewal months दूर हो। दूसरा, price-change notice windows पर Anthropic से एक written commitment: legacy seat-to-usage transition ने 2026 में customers को surprise किया, और आपका contract किसी भी rate या tier modification के लिए 90-day notice require करना चाहिए।
तीसरा, explicit Batch और cache discount survivability negotiate करें — language secure करें जो confirm करे कि 50% Batch discount और cache-read multiplier contractual हैं, unilaterally revocable terms of service नहीं। चौथा, एक gateway clause secure करें जो third-party routing layers और BYOK posture permit करे; कुछ enterprise agreements quietly reselling या proxying restrict करते हैं, और आप उस ambiguity को writing में resolved चाहते हैं। यदि Anthropic इनमें से किसी भी का resist करे, तो वो खुद एक useful data point है कि वे आपके contract term पर कितनी pricing flexibility exercise करने की expect करते हैं।
Frequently asked questions
- क्या नया Anthropic Enterprise pricing हर customer को affect करता है?
- Immediately नहीं। Anthropic ने confirm किया है कि new structure 150 या ज़्यादा seats वाले deployments को target करती है, existing premium और standard seat tiers contract renewal पर combined base-fee-plus-usage model पर transitioning के साथ। Smaller deployments अभी legacy pricing पर रहते हैं, हालांकि अधिकांश industry analysts expect करते हैं कि model समय के साथ downward extend होगा। Practical implication: यदि आपका renewal next दो से चार quarters में है और आप seat threshold से ऊपर हैं, तो sales team के quote भेजने के बाद नहीं, अब एक usage-based spend forecast बनाएं।
- क्या 50% Batch discount और 90% cache-read discount actually permanent हैं?
- Anthropic दोनों को current API documentation में publish करता है: Message Batches API 24 hours के अंदर returned asynchronous requests के लिए standard token rates का 50% charge करता है, और prompt caching cached input reads पर एक 0.1x multiplier use करता है, एक 90% discount के बराबर। Discounts stack करते हैं। वे documented features हैं, promotional pricing नहीं, लेकिन वे अभी भी terms-of-service-level commitments हैं जिन्हें Anthropic revise कर सकता है। Multi-year enterprise contracts negotiate करने वाली Procurement teams को contract term के लिए दोनों mechanisms को preserve करने वाली explicit contractual language request करनी चाहिए।
- मैं historical data के बिना per user token consumption कैसे forecast करूं?
- Anthropic के Compliance API या आपके gateway के logs का उपयोग करते हुए, role द्वारा segmented, एक दो से चार week instrumentation window से शुरू करें। Industry reporting से useful rough benchmarks: agentic IDE workflows में Claude use करने वाले developers अक्सर per active day one से three million tokens consume करते हैं, research और analyst roles 200,000 से 500,000 के closer, और Haiku पर high-volume support automation frequently 100,000 के नीचे। Working days से multiply करें, एक growth assumption (15 से 30 percent quarter over quarter typical है) apply करें, और expected cache और batch coverage के लिए adjust करें।
- क्या providers switch करना एक realistic hedge है या सिर्फ एक negotiating threat?
- दोनों, यदि architecture support करे। Real switching को एक gateway या routing layer चाहिए जो provider-specific SDKs abstract करे, reroutes पर prompt-cache hit rates preserve करे, और per workspace token level पर usage meter करे। उस layer के बिना, multi-provider portability के claims शायद ही एक real migration survive करते हैं। उसके साथ, आप model class द्वारा traffic route कर सकते हैं, incidents के दौरान एक secondary provider पर fall back कर सकते हैं, और renewal conversations में एक verbal threat के बजाय measurable elasticity use कर सकते हैं। Deploy करने की ज़रूरत होने से पहले capability build करें।
Sources