Spotlight

Case Study Microsoft

How Microsoft scaled global content delivery

Find out how Microsoft used Gcore to strengthen delivery across regions.

case study ProSieben GNTM app TOPSHOT

How ProSieben scaled GNTM's app TOPSHOT

Explore how ProSieben brought real-time AI portraits to GNTM's audience.

case study Higgsfield

How Higgsfield scaled AI video generation

See how Gcore helped Higgsfield scale with GPUs and Managed Kubernetes.

case study Fawkes Games

How Fawkes Games stopped DDoS attacks

See how Gcore protected gaming servers from massive DDoS threats without disrupting gameplay.

We're hiring

Help build the next chapter of the web

We're not just filling seats. We're building a team that will write the next chapter of the internet.

  1. Home
  2. Blog
  3. New AI inference models available now on Gcore
News
AI

New AI inference models available now on Gcore

  • November 17, 2025
  • 2 min read
New AI inference models available now on Gcore

We’ve expanded our Application Catalog with a new set of high-performance models across embeddings, text-to-speech, multimodal LLMs, and safety. All models are live today via Everywhere Inference and Everywhere AI, and are ready to deploy in just 3 clicks with no infrastructure management and no setup overhead.

This update brings stronger retrieval accuracy, more expressive voice generation, real-time audio-native LLMs, and enterprise-grade safety controls. Whether you’re building search pipelines, conversational agents, IVR systems, or production-scale AI applications, these additions give you more flexibility to optimize for quality, latency, and cost.

Text embeddings (5 new models)

High-quality embeddings are the backbone of any AI that needs to find, rank, or understand information, including RAG, semantic search, personalization, recommendations, and clustering. This new set of embedding models dramatically improves retrieval precision, cross-lingual reach, and overall RAG quality.

  • Alibaba-NLP/gte-Qwen2-7B-instruct: High-quality instruction-tuned embeddings for retrieval, reranking, and semantic search across broad domains. Ideal for RAG pipelines that need strong generalization.
  • BAAI/bge-m3: Multilingual, multi-function embeddings built for search, clustering, and cross-lingual retrieval. A great fit for global applications and multi-language knowledge bases.
  • intfloat/e5-mistral-7b-instruct: E5-style instruction-following embeddings optimized for retrieval tasks, question-answer matching, and ranking. Strong performance on RAG evaluation benchmarks.
  • Qwen/Qwen3-Embedding-4B: A cost-efficient, versatile embedding model delivering balanced performance for large-scale retrieval workloads.
  • Qwen/Qwen3-Embedding-8B: A higher-capacity sibling offering premium embedding quality for challenging retrieval, reranking, and high-accuracy semantic search.

Text-to-speech (2 new models)

Voice is becoming a first-class interface. These new TTS models make agents feel more natural, reduce robotic cadence, and improve clarity, especially in high-volume workflows like support, IVR, media generation, and automation.

  • microsoft/VibeVoice-1.5B: Neural TTS with natural prosody, expressive cadence, and fast synthesis, built for interactive applications where latency matters.
  • ResembleAI/chatterbox: Production-ready TTS capable of expressive, characterful speech. Ideal for agents, IVR, content workflows, and automated voice experiences.

Text + audio LLMs (2 new models)

These new multimodal LLMs accept both text and audio, enabling real-time voice agents, transcription intelligence, and interactive multimodal applications. They eliminate the need to stitch together separate ASR → LLM → TTS pipelines.

  • mistralai/Voxtral-Mini-3B-2507: A lightweight speech-and-text LLM for real-time voice agents. Handles both text and audio inputs/outputs and is optimized for low-latency scenarios.
  • mistralai/Voxtral-Small-24B: A mid-size Voxtral variant offering higher-quality multimodal reasoning and richer conversational speech. Suitable for advanced voice assistants, transcription workflows, and audio-aware applications.

Safety models (3 new models)

As enterprises deploy AI into production, safety is non-negotiable. These models offer high-quality classification, risk detection, and output transformation to help organizations stay compliant.

  • openai/gpt-oss-safeguard-120b: A high-capacity safety model supporting policy classification, risk detection, and output guidance. Built for enterprise-grade moderation systems.
  • openai/gpt-oss-safeguard-20b: A lighter, faster safeguard variant designed to power low-latency moderation pipelines without sacrificing accuracy.
  • Qwen/Qwen3Guard-Gen-8B: A guardrail model specialized in detecting unsafe content and transforming or steering outputs toward compliant responses.

Deploy the latest models in 3 clicks and 10 seconds

All models are available today via Gcore Everywhere AI and Gcore Everywhere Inference. Deploy publicly or privately, whichever fits your architecture.

You get:

  • Global low-latency routing
  • Predictable cost and usage visibility
  • Zero infrastructure management
  • Instant scaling to production workloads

Open the Gcore Customer Portal, choose a model, and deploy in just three clicks.

Deploy these new AI models today

Try Gcore AI

Gcore all-in-one platform: cloud, AI, CDN, security, and other infrastructure services.

Related articles

Sifted 100 France & Benelux 2026 announcement with falling confetti and spotlights.
Gcore named in Sifted Top 100 France & Benelux 2026

Gcore has been recognized as one of the top 100 fastest-growing technology startups in France and Benelux by Sifted — one of Europe's leading tech publications. Our inclusion in the B2B SaaS & Cloud Infrastructure category points to ris

GCORE and NVIDIA's Global Inference Routing, accelerated by NVIDIA Dynamo, features a glowing green network sphere.
Gcore introduces Global Inference Routing accelerated by NVIDIA Dynamo

Earlier this year we brought NVIDIA Dynamo to Gcore — one-click disaggregated inference that delivered up to 6× higher GPU throughput and 2× lower latency inside a deployment, by separating prefill and decode and routing each request to the

Two founders discuss Melious AI moving its CDN and DNS to Gcore.
Why Melious AI moved its CDN and DNS to Gcore: a founder conversation about sovereign AI in Europe

For many startups, infrastructure decisions are mostly about performance, pricing, and developer experience. For Melious AI, they are also about trust.Melious AI is a German startup building a European AI platform around privacy, transparen

GCORE and Graphiant logos connected by an 'X', signifying a partnership.
Gcore and Graphiant: Accelerating sovereign AI infrastructure with secure neo-cloud connectivity

As enterprises move AI from experimentation into production, they face a new infrastructure challenge. AI applications, models, and data are no longer confined to a single cloud or data center. Instead, they are distributed across multiple

5 insights on AI infrastructure from Nexus Luxembourg 2026

Nexus Luxembourg is Europe's premier AI and technology summit, and this year's edition brought together more than 10,000 visitors, 150+ speakers, and 250 startups from over 50 countries. Gcore CEO Andre Reitenbach joined LuxProvide's Arnaud

An isometric illustration of a secure server rack with a shield icon and glowing data activity.
AI sovereignty isn’t politics: it’s a sales requirement

Across Europe, I keep seeing the same pattern in public sector deals, regulated industries, and anything that smells like critical infrastructure: "AI sovereignty" has moved from a nice-to-have to the first real checkpoint in the deal. Not

Subscribe to our newsletter

Get the latest industry trends, exclusive insights, and Gcore updates delivered straight to your inbox.