Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors
What happened
Learn how to use Amazon S3 Vectors as the persistent memory layer within the NVIDIA NeMo Agent Toolkit (NAT), deployed on Amazon Elastic Kubernetes (software that runs and manages applications across many servers) Service (Amazon EKS). NVIDIA NeMo Agent Toolkit (NAT) is an open source framework for building, profiling, and optimizing AI agents (AI that carries out multi-step tasks rather than answering one question). NVIDIA is a semiconductor company based in Santa Clara, and its products and services include software.
It’s framework-agnostic, working with Strands Agents, LangChain, LlamaIndex, CrewAI, and custom implementations. NAT provides four capabilities relevant to production agent systems: Agent orchestration – Define agents as composable workflows with configurable large language model (the kind of AI system trained on text to produce text)s (LLMs), tools, and prompts.
You run them locally with nat run or as persistent services with nat serve . Profiling – Track token usage, latency (the delay between asking for something and getting it), throughput (how much work a system gets through in a given time), and run times across agents and individual tools to identify bottlenecks in multi-agent workflows.
Key facts
- NAT — provides: four capabilities relevant to production agent systems: Agent orchestration – Define agents as composable workflows with configurable large language models (LLMs), tools, and prompts
Sources & evidence
- AWS Machine Learning Blog Primary / official
Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors ↗
https://aws.amazon.com/blogs/machine-learning/build-agent-memory-with-nvidia-nemo-agent-toolkit-and-amazon-s3-vectors/