AWS says its Amazon Bedrock AgentCore platform supports MCP Apps, a Model Context Protocol extension that renders interactive HTML widgets inside AI hosts like ChatGPT and Claude. AWS demonstrated the capability with a sample app called Unicorn Rentals, showing browsing, booking, viewing, and returning rentals through AgentCore's runtime and gateway.
AWS published a blog post describing how AgentCore Evaluations and AWS DevOps Agent work together to monitor AI agents in production. AWS says AgentCore Evaluations scores live agent interactions for quality issues while DevOps Agent traces infrastructure failures across service boundaries. AWS demonstrated the approach using a four-agent airline reservation system.
Skild AI launched S1, a robot foundation model that learns new, multistep tasks from a single video demonstration without retraining. Built on NVIDIA's Isaac Lab, Cosmos, Omniverse and TensorRT technologies, Skild says S1 succeeded at 66% of steps on new tasks versus 9% for a comparable system, and the company has reached a $100 million annual revenue run rate.
Mistral described a project with a European energy operator that migrated 40,000 lines of Fortran 77 from a physics-intensive reservoir simulator to C++. The company said it built a numerical parity harness, used agents to document the codebase, and settled on workflows combining coder, tester and reviewer agents with human checkpoints.
AWS published a case study describing how Heurist Finance built an AI-powered investment workbench on Amazon Bedrock AgentCore. AWS says the system combines market and alternative data, filings, news, portfolio construction, scenario analysis and position monitoring in a chat experience, with AgentCore services handling identity, memory, code execution, payments and observability.
Databricks says Zepto built a dual-loop evaluation framework for customer support agents using Databricks and MLflow. The system combines development testing, production monitoring, execution traces, golden datasets and deployment quality gates. According to Databricks, the agents fully manage more than 80% of support tickets with human oversight.
AWS published a technical guide for deploying a restaurant-ordering assistant on WhatsApp using Amazon Bedrock AgentCore and Amazon Nova 2 models. The design supports text messages, voice notes and calls through one business number, with shared cross-channel memory and a common backend for menus, carts, orders and locations.
Anthropic's guidance reframes prompt engineering as context engineering: curate the smallest set of high-signal tokens for each turn instead of accumulating everything. In its own evals, automatically clearing stale tool results plus an external memory file improved a search task by 39% and cut token use by 84% over 100 turns.
A Skill lives or dies on one field: the description, written in third person, stating what it does and when to use it. Keep SKILL.md under 500 lines and reference files one level deep, test on every model you plan to use it with, and never install a Skill from a source you don't trust.
Developers interested in building AI agents can work through Google and Kaggle’s five-day course as a self-paced program. Study the codelabs, technical whitepapers and notebooks, then build a capstone that covers agent design, security and cloud deployment. Use Kaggle’s Discord for debugging help and study groups.
Anthropic announced life-sciences partnerships with the Allen Institute and Howard Hughes Medical Institute. HHMI will work with Anthropic on specialized laboratory agents, while the Allen Institute will collaborate on coordinated multi-agent systems for scientific analysis, experimental design and other research tasks. Both partnerships will also inform Claude’s broader life-science capabilities.
A harness is the code around the model: it assembles context, runs the tool-call loop, streams events, compacts long sessions, and holds irreversible actions behind human approval. OpenAI released Codex's harness — codex exec, the app-server and the SDK — under Apache-2.0, so any company can embed that same agent loop in its own software while still paying for the model behind it.
An AI agent is a language model that has been given tools it can call, a goal to pursue, and permission to take several steps without asking a human between each one. A chatbot answers a question and stops; an agent keeps acting until it decides the goal is met or it runs out of budget.
Demos run a short happy path once with a human watching. Production runs thousands of variations unattended, where per-step error rates compound and an unbounded permission scope turns a wrong decision into an incident. The deployments that work narrow the scope, verify each step cheaply, and gate every irreversible action.