World of Tavily is coming to SF Tech Week on Oct 9. Get on the list

How NVIDIA Uses Tavily to Post-Train Deep Research Agents

/Post-training3 min read

How NVIDIA Uses Tavily to Post-Train Deep Research Agents

See how NVIDIA used Tavily across 80,000 research trajectories, retained 67,000 for fine-tuning Nemotron 3 Super, and ranked first on DeepResearch Bench I and II

Noah Nefsky

Every unnecessary search in a research agent costs more than an API call. The model has to process the results, decide what matters, and choose its next action. Across a long investigation, weak research behavior becomes expensive.

Post-training gives AI labs a way to improve that behavior and specialize models for valuable research tasks. The aim is frontier-level quality on a defined task, using a smaller model and less inference compute. NVIDIA researchers make the case for specialized models in agentic systems. Getting there starts with training data that demonstrates effective research.

NVIDIA used Tavily to supply live web results for the trajectories used to fine-tune Nemotron 3 Super. In March 2026, its AI-Q system reached first place on DeepResearch Bench I and II, showing what specialized models can achieve when paired with a capable research workflow and tools.

How NVIDIA turned research runs into training data

NVIDIA started with open-source research questions and ran them through its AI-Q deep researcher, powered by GPT-OSS-120B. Each run produced a trajectory capturing the agent’s plans, searches, retrieved results, and final response. When the agent needed the live web, it queried Tavily and incorporated the returned results into the trajectory.

NVIDIA generated approximately 80,000 trajectories, then filtered them for quality using Qwen3-Nemotron-32B-GenRM-Principle. Approximately 67,000 passed the filter and became training data for supervised fine-tuning.

A flow chart showing how to generate fine-tuning data with Tavily

NVIDIA then fine-tuned Nemotron 3 Super on those trajectories for one epoch over 5,615 steps, taking approximately 25 hours on 128 H100 GPUs.

Where the specialized model fits

AI-Q divides the research process across a planner, researcher, and orchestrator. The planner determines how to approach the task, the researcher gathers and synthesizes evidence, and the orchestrator coordinates the work and determines when more research is needed.

For NVIDIA’s benchmark submission, the fine-tuned Nemotron 3 Super powered the researcher and its subagents. The researcher processed four times as many tokens as the planner and orchestrator combined, placing the specialized model in the most inference-intensive part of the system.

Why the search tool matters

The search tool helps determine what happens inside each trajectory. When retrieval returns noisy or off-target content, the agent may need to search again, spend more of its context processing irrelevant text, or reach its tool-call or time budget before finding enough evidence.

A flow chart showing noisy versus quality retrieval

With high-quality retrieval, relevant content reaches the agent with less noise. That can mean fewer, more purposeful turns and more context available for reasoning and synthesis. This gives the agent a better chance of completing the task within budget and producing a trajectory that passes the quality filter.

At NVIDIA's scale, this effect compounds. Across 80,000 research runs, every trajectory that fails the quality filter is wasted compute that never reaches the SFT dataset. Retrieval quality, therefore, shapes more than individual trajectories: it determines how much useful training data is produced, and ultimately, how capable the resulting model is.

There’s another reason the tool matters: the model is learning how to use it. Using the same retrieval tool during training and deployment keeps trajectories representative of the environment the model will encounter at inference time. The model learns research behavior against the search capabilities and responses it will actually use.

What specialization makes possible

The model can specialize in orchestrating retrieval and synthesizing findings, while external tools supply current information. A smaller model that becomes highly capable at those behaviors can deliver strong research performance with less inference compute. NVIDIA’s first-place results on DeepResearch Bench I and II show what’s possible when a specialized model is paired with a capable research workflow and tools.

NVIDIA isn't an isolated example. At Tavily, we’re working with multiple AI labs on web retrieval for post-training, helping teams tune search and agent harnesses for trajectory generation and long-horizon research.

Training a research agent? Talk to Tavily about retrieval for your post-training workload.