AI Prosumer
EN

Models

3 models

Prices per 1M tokens

nvidia/nemotron

Llama-3.1-Nemotron-70B-Instruct is a large language model customized by NVIDIA to improve the helpfulness of LLM generated responses to user queries.

17 model tags131K contextToken exchange

nvidia/llama3-chatqa

A model from NVIDIA based on Llama 3 that excels at conversational question answering (QA) and retrieval-augmented generation (RAG).

35 model tags8K contextToken exchange

nvidia/nemotron-mini

A commercial-friendly small language model by NVIDIA optimized for roleplay, RAG QA, and function calling.

17 model tags4K contextToken exchange

Filters

Providers
Creators
Access
Serverless