Particle.news

OpenAI Launches IndQA Benchmark for Indian Languages, Publishes First Results

OpenAI frames the India-focused test as a tool to track progress within model families, not a cross-language leaderboard.

Overview

  • IndQA evaluates culturally grounded, reasoning-heavy questions drawn from everyday Indian contexts across 2,278 items, 10 domains, and a dozen languages written natively, including Hindi and Hinglish.
  • It was built with 261 India-based domain experts using rubric-based grading with ideal answers and an AI model grader, with English translations provided for auditability.
  • Questions were adversarially filtered by testing against OpenAI’s strongest models at creation, including GPT-4o, o3, GPT-4.5, and partially GPT-5 after its public launch, to retain items where models previously failed.
  • Initial scores show GPT-5 Thinking High leading at 34.9% with Google’s Gemini 2.5 Pro Thinking close behind at 34.3%, with higher performance in Hindi and Hinglish and lower in Bengali and Telugu, and a domain low in History.
  • The release is part of a broader India push that includes a Bengaluru developer event, a limited-time free year of ChatGPT Go for new users in India, and plans for a first office in Delhi with local hiring.