Overview
- IndQA evaluates culturally grounded, reasoning-heavy questions drawn from everyday Indian contexts across 2,278 items, 10 domains, and a dozen languages written natively, including Hindi and Hinglish.
- It was built with 261 India-based domain experts using rubric-based grading with ideal answers and an AI model grader, with English translations provided for auditability.
- Questions were adversarially filtered by testing against OpenAI’s strongest models at creation, including GPT-4o, o3, GPT-4.5, and partially GPT-5 after its public launch, to retain items where models previously failed.
- Initial scores show GPT-5 Thinking High leading at 34.9% with Google’s Gemini 2.5 Pro Thinking close behind at 34.3%, with higher performance in Hindi and Hinglish and lower in Bengali and Telugu, and a domain low in History.
- The release is part of a broader India push that includes a Bengaluru developer event, a limited-time free year of ChatGPT Go for new users in India, and plans for a first office in Delhi with local hiring.