LLM Evaluation & Quality Engineering Playbook 2026
A practical 2026 playbook for evaluating and monitoring LLM systems with task-specific tests, adversarial checks, and production telemetry grounded in standards.
A practical 2026 playbook for evaluating and monitoring LLM systems with task-specific tests, adversarial checks, and production telemetry grounded in standards.
Start With the Local Codex Threat Model, Not a Convenience Setting Local Codex hardening begins with a simple operating assumption: a coding assistant that can inspect a repository, run shell commands, write files, call network resources, consult web content, or use connected tools is crossing trust boundaries on every step….
Start with the operating surface, not with a “best” scanner Codex Security is not a single button with one deployment model. OpenAI documents four operating surfaces: the desktop plugin or workbench, the command-line interface, the TypeScript SDK, and a connected-GitHub cloud workflow. Each surface changes who initiates the scan, where…
Benchmark GPT-5.5 Reasoning Effort Before You Route Real Work GPT-5.5 is documented by OpenAI as an official flagship API model for complex professional work, and its API model ID is gpt-5.5. In the official GPT-5.5 model documentation, OpenAI lists text and image inputs with text output, a 1,050,000-token context window,…
What changed in September, and why Health permissions now need a closer look OpenAI’s September 14, 2026 update changed the starting permission posture for new Health plugin connections: when a user connects Health after the update, ChatGPT defaults that Health connection to the user’s existing global Plugins permission setting. If…
Start by separating the experience from the entitlement OpenAI’s current ChatGPT Work and Codex administration model separates three concepts that are easy to blur during rollout: the product surface a user starts in, the default configuration applied when that surface opens, and the permissions that decide whether the user can…
OpenAI has changed an important default in ChatGPT for Plus and Pro users: Instant no longer automatically escalates to a higher thinking level when a request appears complex. Reasoning has not been removed. Instead, Plus and Pro users must explicitly choose an available thinking level in the model picker when…
Why this playbook starts with disrupted misuse cases, not threat prevalence Anthropic’s September 2026 threat-intelligence report is best read as a set of selected notable disrupted cases observed between December 2025 and August 2026, not as a statistical prevalence study of AI misuse across the internet, cloud platforms, or all…
What OpenAI OneGov 2.0 changes for public-sector AI adoption OpenAI’s OneGov 2.0 announcement is best understood as a procurement and adoption framework for U.S. government organizations, not as a universal grant of unlimited free AI. OpenAI and the U.S. General Services Administration announced a 27-month agreement that begins on October…
How to use ChatGPT-5.5 as an administration planning partner without giving it authority it does not have ChatGPT-5.5 can help workspace administrators draft access reviews, critique group-provisioning plans, normalize audit evidence, prepare key-rotation runbooks, and identify missing approvals before a change reaches production. It is not, by itself, an administrator…
ChatGPT AI Hub Tools
© 2026 ChatGPT AI Hub