CHATGPT
30 days · UTC
Synchronizing with global intelligence nodes...
Cut LLM spend by moving embeddings local and squeezing inference — no new hardware required
Teams are quietly slashing LLM costs by running embeddings on CPU and optimizing inference before buying more GPUs. A dev building a small RAG app re...
Real-time AI DLP moves to the browser and the MCP layer
Nightfall AI now blocks sensitive data at paste-time and monitors MCP agent traffic, shifting DLP from after-the-fact alerts to real-time prevention. ...
OpenAI ships GPT-6 Astra: an agent-class model rolling out to ChatGPT and the API
OpenAI released GPT-6 Astra, a faster, more aligned agent-class model now rolling out across ChatGPT and major clouds. OpenAI says [GPT‑6 Astra](http...
Hidden-text prompt injections are gaming AI hiring screeners
AI hiring screeners are being gamed by hidden-text prompt injections in resumes and profiles. Candidates are slipping instructions into resumes and p...
OpenAI ships GPT-6 Astra across ChatGPT, API, Azure, and Bedrock — 1M context and breaking API knobs
OpenAI launched GPT-6 Astra across ChatGPT and developer channels, adding a 1M-token context window and changing key API controls. Rollout hits ChatG...
Anthropic pushes an AI‑native SDLC: code isn’t the bottleneck anymore
Anthropic is reframing software delivery around evidence, governance, and production feedback, not faster code generation. Anthropic’s AI‑Native SDLC...
WebMCP turns agent automation into first‑class site actions; make your agent docs testable
WebMCP is moving real agent automation from scraping to site-declared tools, and teams should treat agent-facing docs as testable interfaces. [WebMCP...
OpenAI trims GPT 5.6 Sol pricing by 20% — re-baseline your token budgets
OpenAI cut GPT 5.6 Sol API pricing by 20%, but teams should tighten token budgets instead of loosening them. OpenAI’s forum announcement points to a ...
Slack Code brings AI coding into shared Slack channels with built‑in review and PR handoff
Slack launched Slack Code, a shared Slack channel workspace where AI coding agents work under team oversight and ship pull requests with traceable con...
Hitting Claude’s cap? A free extension moves the whole chat to ChatGPT in seconds
A free browser extension makes it trivial to move an AI chat from Claude to ChatGPT when you hit usage limits. A TechRadar writer hit Claude’s usage ...
Codex Linux preview lands amid rate‑limit resets, phantom credit drain, and a severe bug report
OpenAI’s Codex app reached Linux preview while forum reports highlight rate‑limit resets, unexpected credit usage, freezes, and one severe file‑deleti...
LLM security meets architecture: defend against poisoning and design for model choice
Organized data poisoning and leakage concerns around LLMs are pushing teams toward model-agnostic orchestration and stronger data governance. TechRad...
Agent security shifts from alerts to runtime enforcement
Agent security is moving from alerts to autonomous enforcement, with OpenAI’s Codex Security CLI and new runtime blockers reacting to real prompt-worm...
OpenAI rolls out GPT-5.6; early tool gaps, agent bugs, and cost shifts surface
OpenAI shipped GPT-5.6 broadly, but developers report IDE gaps, agent quirks, and higher run costs. OpenAI signaled GPT-5.6 availability across ChatG...
ChatGPT 5.5 turns prompts into state: persistent memory and project context arrive
ChatGPT 5.5 now keeps user and project context across chats, shifting LLM app design from stateless prompts to stateful systems. A deep dive shows pe...
Agents Need a Governance Layer Before They Scale
Agentic AI is stalling on governance, not models or UI, and that changes where backend teams need to invest. An industry brief argues the real bottle...
From chat to delegation: Codex data shows agents are becoming workflows, not answers
OpenAI’s Codex data shows engineers are delegating multi-step work to agents, not chatting for answers. In [The Shift to Agentic AI: Evidence from Co...
OpenAI previews GPT-5.6 (Sol/Terra/Luna) with new pricing and cache semantics under limited rollout
OpenAI previewed the GPT-5.6 model family (Sol, Terra, Luna) with new pricing and stricter prompt-caching rules in a limited U.S.-only rollout. Per O...
RAG reliability is a context engineering problem, not a prompt problem
RAG reliability hinges on how you structure and retrieve context, not on prompt tweaks or chunk-size folklore. A teardown of failing production pipel...
RAG Reality Check: HNSW Everywhere, Filters Decide; Read Fewer Images
Most vector stores use HNSW, so your filtering and scale decide whether pgvector is enough or you need Qdrant, Pinecone, or Weaviate. A hands-on comp...
OpenAI sunsets GPT-5.2/5.3-Codex in Codex; expect migrations and sporadic capacity pain
OpenAI removed GPT-5.2 and GPT-5.3-Codex from Codex for ChatGPT subscribers, pushing teams to swap models while capacity errors surface. Users report...
ChatGPT 5.5 mode shift triggers real-world regressions; OpenAI SDK adds spend alerts
OpenAI shifted ChatGPT to 5.5 modes, retired older Codex models, and teams are seeing operational side effects. A practical explainer walks through h...
OpenAI Responses API adds conversation state, simplifying multi-turn chat backends
OpenAI's Responses API now includes built-in conversation state that replaces thread-like handling for multi-turn chats. OpenAI’s new [conversation s...