← FeedCopy sourceOpenAI· GPT / ChatGPT / APISeparating signal from noise in coding evaluationsA new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.Visit source ↗openai.comAnnounced08 Jul 2026 13:00 UTCOther storiesOpenAI · 18 AugIntroducing ChatGPT for Teens: Built for learning, backed by protectionsProduct launchOpenAI · 18 AugIntroducing ChatGPT for Teens: Built for learning, backed by protectionsProduct launchOpenAI · 18 AugPartnering with CodeAI to prepare the first AI generationOpenAI · 17 AugThe Defender’s WindowProduct launchOpenAI · 17 AugOpenAI joins PORTS-Pike projectPartnershipOpenAI · 17 AugNew policy ideas for the Intelligence AgeProduct launch