McKinsey has implemented an email alert system to notify employees when their artificial intelligence (AI) token usage exceeds predefined thresholds, according to Debasish Patnaik, global leader of QuantumBlack, the firm’s AI, data, and analytics division. The system, introduced firmwide in summer 2024, tracks consumption on a user-by-user basis and provides cost-saving recommendations via email, similar to mobile data usage alerts.
The move reflects broader industry shifts as companies transition from flat-rate AI access models to consumption-based pricing, where costs are tied to the number of tokens processed. This pricing model has intensified scrutiny over AI spending efficiency, particularly as usage scales. OpenAI reported in September that its most active users of AI coding agents were consuming over $7,000 worth of tokens daily, underscoring the financial stakes.
By May 2026, McKinsey was processing approximately five trillion AI tokens per month, with consumption heavily concentrated among a small group of users. Internal data showed that 10% of employees accounted for roughly 65% of total token usage, with consultants and software engineers among the highest consumers. The firm’s approach emphasizes transparency and education rather than outright caps on AI usage, aiming to balance innovation with cost control.
Industry Context and Cost Pressures
The shift to token-based pricing has forced enterprises to rethink AI adoption strategies. Unlike traditional subscription models, where costs remain fixed regardless of usage, consumption-based pricing ties expenses directly to activity levels. This has created new financial pressures for firms that initially encouraged widespread AI experimentation without fully accounting for long-term operational costs.
McKinsey’s solution aligns with its broader advisory role, where it guides clients on managing AI expenditures. The firm’s internal controls reflect a growing trend among consulting and technology companies to implement granular tracking systems that provide real-time feedback on resource consumption. Patnaik described the approach as a proactive measure to ensure cost efficiency while maintaining productivity.
Usage Patterns and Concentration Risks
The concentration of token usage among a minority of employees highlights potential inefficiencies in AI deployment. Heavy reliance on AI tools by consultants and engineers suggests that certain roles may be more prone to high-volume interactions, such as generating large datasets, running iterative queries, or processing extensive documentation. McKinsey’s data indicates that targeted interventions—such as alerts and optimization tips—could significantly reduce overall spending without restricting access.
The firm’s strategy also underscores the broader challenge of scaling AI tools in enterprise environments. As organizations integrate AI into workflows, they face the dual challenge of maximizing utility while mitigating financial risks associated with unpredictable usage patterns. McKinsey’s system represents one of the first publicly documented efforts to address these challenges through automated monitoring and user education.