Agents
AI news.
Current agents releases and research, summarized from live sources and linked to the original publisher.
Introducing Kimi K3 on Amazon Bedrock
Open-weight models are changing the economics of building and deploying AI at scale. Rapid gains in intelligence and efficiency mean companies can match each workload with the right balance of capability, speed, and cost. AWS is building for a future in which organizations can adopt open-weight innovation with the reliability and security required for production. Today, Kimi K3 from Moonshot AI is available on Amazon Bedrock, giving you a powerful new option for coding and knowledge work. According to Moonshot AI,
READ BRIEFING ↗Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime
Organizations building multi-model agentic AI applications face growing infrastructure complexity. Managing container orchestration, scaling policies, identity, and observability for multiple model types adds operational overhead. Teams often spend more time on infrastructure than on agent logic development. Developers running agentic frameworks on self-managed infrastructure such as Amazon Elastic Container Service (Amazon ECS) with AWS Fargate have full control over their deployment configuration. As agentic work
READ BRIEFING ↗The new AgentCore runtime: Elastic, optimized, and consistently fast starts
Agents are no longer experiments. They process claims, write and review code, coordinate across systems, and run for hours without supervision. As agents take on more complex, longer-running work, the infrastructure underneath them must evolve just as fast. We built Amazon Bedrock AgentCore to help developers build, connect, and optimize agents securely at scale. AgentCore runtime, a capability of Amazon Bedrock AgentCore, is the managed compute layer that gives developers a fully managed environment to deploy and
READ BRIEFING ↗Deploy Hugging Face models on Amazon SageMaker AI with coding agents
Deploying a Hugging Face model to production means making a dozen decisions: choosing the right serving container for the model’s architecture, confirming the current image tag for your AWS Region, and matching an instance type to the model’s memory footprint. Beyond infrastructure, you must wire autoscaling so you don’t burn GPU hours on an idle endpoint. You also set Amazon CloudWatch alarms that catch silent failures before your users do. After you’ve made those decisions, Amazon SageMaker AI collapses that work
READ BRIEFING ↗Selecting a vector store for Amazon Bedrock Knowledge Bases
When building a Retrieval Augmented Generation (RAG) solution with Amazon Bedrock Knowledge Bases , selecting the right vector store impacts performance and cost. Amazon Bedrock Knowledge Bases offers a fully managed option and a customer-managed option where you choose your own vector store. This post focuses on the customer-managed path, comparing the three supported backends: Amazon OpenSearch Service , Amazon Aurora PostgreSQL with pgvector , and Amazon S3 Vectors , a capability of Amazon Simple Storage Service
READ BRIEFING ↗A serverless, data-driven Git metrics dashboard using Amazon Quick Sight
Git activity is one of the richest signals engineering teams produce that can provide continuous observability into development analytics. The challenge is extracting these Git metrics at scale, which has traditionally required hand-rolled extract, transform, and load (ETL) jobs, dedicated infrastructure, and ongoing maintenance. Further, with modern development tools becoming more prevalent, teams need a clear way to measure whether these tools are making developers faster, or if the investment is not paying off.
READ BRIEFING ↗A shared agentic platform for Wood Mackenzie, on Amazon Bedrock AgentCore
Building a working agentic prototype takes an afternoon. Getting it to production is where the work explodes. The moment an agent has to serve more than one user, a new layer of engineering appears and it’s critical to tell whether the agent is doing the right thing on real traffic. Concurrency, session isolation, identity, persistent state, scaling, and guardrails are all layers that most teams rebuild from scratch every time, even though a standardized platform for each agent is a more repeatable approach. The re
READ BRIEFING ↗How MRH Trowe enabled secure self-service AI agents in financial services
Basic AI chat isn’t enough for financial services organizations that need secure, self-service AI agents. In financial services, employees need AI that can work with internal systems and sensitive client data, stay inside a governed environment, and remain auditable and cost-transparent. All of this must happen without every team standing up its own tools. This post shows how MRH Trowe , one of Germany’s leading commercial and industrial insurance brokers, gave approximately 400 employees secure access to self-serv
READ BRIEFING ↗How to Use AI Agents to Prepare 3D Scenes for Simulation
Agentic AI workflows can be used to prepare and validate digital twins for physical AI systems. Agents can inspect 3D scenes, author simulation-relevant data in...
READ BRIEFING ↗TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor
AI agents are moving from cloud data centers to vehicles, robots, and other edge devices. Unlike a chatbot that answers a single prompt, an agent works through...
READ BRIEFING ↗Improving HCLS AI reasoning with open-source agent skills
AI agents built on foundation models (FMs) often misapply healthcare and life sciences (HCLS) decision frameworks, even when they’ve seen the guidelines in training and in the system prompt. Ask an agent to classify a TP53 missense variant using ACMG/AMP criteria. It will cite the correct framework but misapply evidence categories, skip population frequency thresholds, or hallucinate computational predictor scores. The model knows facts but lacks the structured reasoning procedures that domain practitioners interna
READ BRIEFING ↗Fault tolerant distributed training on Amazon EKS using NVRx
Large-scale distributed training jobs run for hours or days across dozens of nodes. At that scale and duration, interruptions are statistically inevitable: network partitions, memory errors, software exceptions, or infrastructure events will eventually disrupt at least one worker. A single GPU fault triggers a cascade: NVIDIA Collective Communication Library (NCCL) timeouts propagate to healthy workers, pods crash and restart out of sync, and your cluster burns expensive GPU hours while making zero training progres
READ BRIEFING ↗Translating CUDA Tile Operations from Python to Rust Using Agentic AI
cuTile Rust (cutile-rs) is a tile-based system for safe, idiomatic GPU kernel authoring in the Rust programming language. Extending the Rust ownership model to...
READ BRIEFING ↗Optimizing agent system prompts with Amazon Bedrock AgentCore
In a previous launch post , we introduced AgentCore optimization, a capability of Amazon Bedrock AgentCore that can help you improve the quality of your agents. Improving a low-scoring agent has traditionally been a manual process. You review long traces to find where the agent goes wrong, tune individual components such as prompts, tool descriptions, and skills, and rerun evaluations to check for improvement. With AgentCore optimization, you can use production traces to propose configuration changes, validate them
READ BRIEFING ↗Build a serverless PII redaction pipeline with Amazon Bedrock Data Automation
Organizations that process thousands of scanned documents daily, including medical forms, insurance claims, and financial records, face a recurring compliance need: personally identifiable information (PII) redaction before documents are shared with third parties or processed downstream. Manual redaction doesn’t scale: It consumes staff hours, introduces human error, and creates compliance exposure. Redaction is also a precision problem, in addition to a detection problem. A single page can contain multiple names,
READ BRIEFING ↗Reimagining advertising with AI
Explore new AI-powered advertising experiences from OpenAI, including Sponsored Agents, tools for marketers, and integrations with HubSpot and Shopify.
READ BRIEFING ↗Hex turns complex analysis into visual reports with GPT‑6 Astra
GPT-6 Astra helps Hex’s data agents turn answers into interactive visualizations that employees are proud to share.
READ BRIEFING ↗Your Agent Aced the Task. Will It Do It Again?
Read Hugging Face’s complete announcement and technical details at the original link.
READ BRIEFING ↗Now everyone can put data to work
Meet the Data agent in ChatGPT Work. Connect company data, uncover insights, and build interactive dashboards with AI using natural language.
READ BRIEFING ↗Building a Memory-Driven Agent with NVIDIA NemoClaw
Enterprise work spans messages, decisions, projects, and obligations that change over time. An AI agent that starts without this context must reconstruct it...
READ BRIEFING ↗Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson
Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run...
READ BRIEFING ↗Give Your Coding Agents a Memory You Own
Read Hugging Face’s complete announcement and technical details at the original link.
READ BRIEFING ↗Training a coding model to paint watercolours with TRL and OpenEnv
Read Hugging Face’s complete announcement and technical details at the original link.
READ BRIEFING ↗AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR
Human-Computer Interaction and Visualization
READ BRIEFING ↗Orchard: An open framework for scalable agentic AI
At a glance Orchard is an open-source framework for scalable and cost-effective agentic AI research, built around Orchard Env, a reusable environment service for training and evaluating agents across task domains. The same Orchard infrastructure supports software-engineering, web-navigation, and personal-assistant agents, and can train them directly inside real deployment harnesses such as Codex, OpenClaw, and ZeroClaw—letting researchers reuse environments, data pipelines, and evaluation workflows across tasks. Or
READ BRIEFING ↗Echoverse: Deep, evolving environments for computer-use agents
Scaling fidelity over sheer count, targeting the capabilities agents actually lack, and evolving with the models they train. At a glance We built twelve training worlds for computer-use agents: ten deep domain worlds and two capability worlds, each drilling a single control rendered in many forms (date pickers and nested filters). Depth is what makes them worth training on: these worlds reproduce an application’s real behavior, come seeded with realistic data, and keep state coherent across screens and users. Train
READ BRIEFING ↗SymptomAI: Towards a conversational AI agent for everyday symptom assessment
General Science
READ BRIEFING ↗Verifying Rust cryptography in SymCrypt, from standards to code
How Rust, Lean, Aeneas, and AI agents are helping scale formal verification for production cryptographic algorithms At a glance SymCrypt develops new verified cryptography using Rust, Aeneas, and Lean to provide higher security assurance. We prove that their code safely and correctly implements standard algorithms, notably for post-quantum cryptography. We are releasing verified code, specs, properties, and proofs initially for SHA-3 and ML-KEM. Aeneas allows verifying a large subset of Rust code and provides effic
READ BRIEFING ↗Flint: A visualization language for the AI era
At a glance Polished charts from simple specs . Flint allows AI agents to reliably generate expressive, visually polished charts from simple, human-editable specifications. Semantic types guide design . Flint leverages semantic data types to express meanings of data. They help the compiler choose appropriate scales, baselines, formatting, and color schemes. Layouts adapt to the data . Flint automatically manages sizing, spacing, labels, and layout so charts remain readable as cardinality and density change, without
READ BRIEFING ↗