# groq.com > AI-optimized mirror of groq.com containing 49 pages totalling 29,503 words of clean markdown content, structured data, and semantic HTML. Original source: https://groq.com/. Last updated: 2026-05-12T00:30:04.983Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [Groq delivers fast, low cost inference that doesn’t flake when things get real.](/site-root.html): The Groq LPU delivers inference with the speed and cost developers need. (476 words) ## Articles & Blog Posts - [Fast, Intelligent AI for Every Industry](/industry-solutions/index.html): The Groq LPU delivers inference with the speed and cost developers need. (899 words) - [Customer Stories](/customer-stories/index.html): The Groq LPU delivers inference with the speed and cost developers need. (783 words) - [Simplifying the Complexity of AI Agents with Server-Side Tool Use](/blog/how-to-build-your-own-ai-research-agent-with-one-groq-api-call.html): Learn to build a real-time AI research agent using the Groq API. Explore agentic workflows and achieve lightning-fast results—start today. (1,841 words) - [Our agentic AI system, rolling out in general availability](/blog/introducing-the-next-generation-of-compound-on-groqcloud.html): Compound Agentic AI is now GA on GroqCloud—run research, code, and web actions at low cost and top speed. Try the fastest agentic AI for free today. (1,007 words) - [**Lack of Standardization**](/blog/openbench-open-reproducible-evals/index.html): Discover OpenBench—Groq’s open, reproducible LLM evals suite. Standardized implementation, quick setup, and faster results. Try OpenBench for free today (636 words) - [blog/advancingamericanai/index.html](/blog/advancingamericanai/index.html) (1 words) - [Why Remote MCP Matters](/blog/introducing-remote-mcp-support-in-beta-on-groqcloud.html): The Groq LPU delivers inference with the speed and cost developers need. (583 words) - [privacy-policy/index.html](/privacy-policy/index.html) (1 words) - [Groq and Nvidia Enter Non-Exclusive Inference Technology Licensing Agreement to Accelerate AI Inference at Global Scale](/newsroom/index.html): The Groq LPU delivers inference with the speed and cost developers need. (322 words) - [_Blog_](/blog/index.html): The Groq LPU delivers inference with the speed and cost developers need. (275 words) - [**Why Orpheus on GroqCloud**](/blog/canopy-labs-orpheus-tts-is-live-on-groqcloud/index.html): Canopy Labs’ Orpheus TTS is now live on GroqCloud, offering low-latency, expressive English and authentic Saudi Arabic text-to-speech for real-time voice apps. (515 words) - [**How It Works — The Four‑Step Loop**](/blog/llms-inside-the-product-a-practical-field-guide/index.html): Discover practical methods for fast, reliable LLM integration in your product. See proven patterns, copy-ready code, and best practices. Create smarter AI today. (1,742 words) - [Terms of Use](/terms-of-use/index.html): The Groq LPU delivers inference with the speed and cost developers need. (2,909 words) - [What Is It?](/blog/the-official-llama-api-accelerated-by-groq/index.html): Access Meta’s official Llama API at record speed, powered by Groq. Enjoy secure, reliable, and scalable inference—start building with Llama models today. (488 words) - [Up to 100MB File Size & Faster Speed for Whisper Large v3 Powered by Groq](/blog/largest-most-capable-asr-model-now-faster-on-groqcloud.html): Experience ultra-fast Whisper Large v3 ASR on GroqCloud—now with 100MB file support and unmatched accuracy. Try today and power your voice-enabled apps. (658 words) - [What Can This Do That an LLM Can’t?](/blog/now-in-preview-groqs-first-compound-ai-system/index.html): Preview Groq's Compound AI System—AI with live web search and code tools for real-time, accurate results. Try it now and build next-generation assistants. (661 words) - [Throughput](/blog/12-hours-later-groq-is-running-llama-3-instruct-8-70b-by-meta-ai-on-its-lpu-inference-enginge.html): Llama 3 by Meta AI runs on Groq’s LPU™—leading token speed and top benchmarks. Build high-performance AI apps now. Test Groq for free. (316 words) - [Performance](/blog/build-fast-with-text-to-speech/index.html): Discover Dialog, the leading TTS model, now live on GroqCloud. Build real-time, expressive voice applications—try powerful text-to-speech AI today. (671 words) - [50% Discount Through April – Double the Savings!](/blog/batch-processing-with-groqcloud-for-ai-inference-workloads.html): Scale beyond speed with GroqCloud™ Batch Processing—efficiently handle massive AI workloads at enterprise scale. (458 words) - [**Scaling GroqCloud for Production Workloads**](/blog/groqcloud-expanding-to-meet-demand/index.html): The Groq LPU delivers inference with the speed and cost developers need. (425 words) - [**New, Lower Prices for GPT‑OSS Models**](/blog/gpt-oss-improvements-prompt-caching-and-lower-pricing.html): Discover how GroqCloud’s GPT-OSS prompt caching slashes costs and boosts speed. Get 50% discounts and instant integration—start building fast AI today. (595 words) - [Working at Groq](/careers-at-groq/index.html): The Groq LPU delivers inference with the speed and cost developers need. (598 words) - [Delivering Fast Inference with the Full 131k Context Window](/blog/groqcloud-tm-now-supports-qwen3-32b/index.html): Qwen3 32B now runs on GroqCloud—build fast, multilingual AI apps with 131k context and agentic reasoning. Enjoy low-cost, real-time inference—get started today. (342 words) - [Scaling Reinforcement Learning is All You Need for the Rise of Smaller, Smarter Models](/blog/a-guide-to-reasoning-with-qwen-qwq-32b/index.html): Discover fast reasoning with Qwen QwQ-32B. Explore RL scaling, benchmarks, and GroqCloud API tips for smarter, low-cost agentic models—try Qwen today. (1,221 words) - [Artificial Analysis Benchmarks Groq LPU™ Fast AI Inference Same Day as Meta Release](/blog/new-ai-inference-speed-benchmark-for-llama-3-3-70b-powered-by-groq.html): Benchmarked: Llama 3.3 70B runs fastest on Groq hardware. Achieve real-time, cost-efficient AI—start building for free today (429 words) - [Groq Performance & Pricing](/blog/llama-4-now-live-on-groq-build-fast-at-the-lowest-cost-without-compromise.html): Experience Llama 4’s blazing-fast Scout and Maverick models, low-cost AI inference on GroqCloud. Launch smarter apps instantly—try Llama 4 now. (526 words) - [Full Model Capabilities](/blog/day-zero-support-for-openai-open-models/index.html): Groq powers OpenAI open models—reach global scale, slash AI costs, and leverage advanced tools from day one. Try GroqCloud for free today! (392 words) - [The challenge: Lag at scale makes for a poor experience](/customer-stories/mem0-redefines-ai-memory-with-real-time-performance-on-groqcloud.html): Mem0 is redefining how AI remembers. It runs a memory system that learns from every user interaction, extracting, updating, and recalling information in real time. To make that possible, it needed fast, reliable inference, and that’s where Groq comes in. By running key parts of its memory pipeline on GroqCloud, Mem0 cut latency by roughly 5x, and end-to-end response times from hundreds of milliseconds to under 100, transforming slow interactions into seamless, real-time experiences that feel instant to users. Together, Mem0 and Groq are making AI capable of remembering instantly and responding intelligently. (816 words) - [What Are Word-Level Timestamps?](/blog/build-fast-with-word-level-timestamping/index.html): Experience precise, word-by-word audio timing—now on GroqCloud STT. Ideal for developers building video, caption, and social media solutions! (526 words) - [Introducing the LPU](/lpu-architecture/index.html): The Groq LPU delivers inference with the speed and cost developers need. (225 words) - [The Challenge](/customer-stories/groq-customer-use-case-vetted/index.html): Groq speed ensures that Vetted’s AI can efficiently process diverse and large datasets, enabling users to receive the most relevant product recommendations, the best pricing, and relevant product specs and insights in real-time. This speed, along with accuracy, is crucial to a superior user experience. (486 words) - [Welcome to Groq’s Galaxy, Elon | Groq is fast, low cost inference.](/blog/welcome-to-groqs-galaxy-elon/index.html): The Groq LPU delivers inference with the speed and cost developers need. (659 words) - [Who Should Care?](/blog/groq-recognized-gartner-cool-vendor/index.html): Groq has been recognized as a 2025 Gartner Cool Vendor in AI Infrastructure, which we believe demonstrates the uniqueness LPUs deliver for real-time AI systems compared to GPU-based alternatives. (579 words) - [Increased Rate Limits](/blog/developer-tier-now-available-on-groqcloud/index.html): Get GroqCloud Developer Tier now—easy self-serve signup, up to 10x rate limits, 25% cost discounts. Build faster AI apps with pay-as-you-go API—sign up today. (379 words) - [Accuracy Without Tradeoffs: TruePoint Numerics](/blog/inside-the-lpu-deconstructing-groq-speed/index.html): Discover how Groq’s LPUs set new AI inference speed records with SRAM design, static scheduling, tensor parallelism, and TruePoint numerics. Learn more today! (1,001 words) - [About Groq](/newsroom/saudi-arabia-announces-1-5-billion-expansion-to-fuel-ai-powered-economy-with-ai-tech-leader-groq.html): Mountain View, California & Riyadh, Saudi Arabia – February 10, 2025 –  Silicon Valley AI pioneer Groq has secured a $1.5 billion commitment from the (325 words) - [Fast, Low Cost, and Seamless AI Inference for Repetitive Workloads](/blog/introducing-prompt-caching-on-groqcloud/index.html): Cut LLM costs by up to 50% with prompt caching on GroqCloud. Enjoy instant speed-ups and no code change—perfect for chatbots and code assistants. Try it free today. (373 words) - [The Evolution of Advanced Openly-Available LLMs](/blog/from-speed-to-scale-how-groq-is-optimized-for-moe-other-large-models.html): Groq’s LPU is a significant advancement in AI hardware, offering the scalability and efficiency needed to support both small and large models. (659 words) - [**What is GPT‑OSS‑Safeguard‑20B?**](/blog/day-zero-support-for-openai-open-safety-model/index.html): Get day-zero, blazing-fast access to OpenAI’s open safety model on GroqCloud. Deploy policy-driven AI moderation with explainable reasoning—try it today. (463 words) - [Delivering Operational Excellence at Scale](/blog/saudi-arabia-announces-1-5-billion-expansion-to-fuel-ai-powered-economy-with-groq.html): Saudi Arabia's $1.5B investment boosts Groq AI infrastructure. Find out how GroqCloud transforms enterprise AI. Explore the latest innovations today. (318 words) - [Unmatched Price Performance](/pricing/index.html): Groq powers leading openly-available AI models. View the pricing of our core models including GPT-OSS, Kimi K2, Qwen3 32B, and more. (577 words) - [Build Fast](/blog/hey-elon-its-time-to-cease-de-grok/index.html): The Groq LPU delivers inference with the speed and cost developers need. (281 words) - [Simplicity of Hugging Face + Efficiency of Groq](/blog/build-faster-with-groq-hugging-face/index.html): Build and deploy LLMs at scale: Groq now on Hugging Face. Speed up your AI workflow, enjoy cost savings, and launch projects instantly. (289 words) - [Project Front Page](/demos/project-front-page/index.html): Perigon’s new web application collects 20 million data inputs daily from over 150,000 global sources, clusters the information into common themes and events, and relationally connects it across people, companies, and locations. Perigon integrates these data points with LLMs powered by Groq, providing an ultra-low latency solution with real-time contextual knowledge. (262 words) - [Why MCP Connectors Matter](/blog/introducing-mcp-connectors-in-beta-on-groqcloud/index.html): GroqCloud’s new MCP Connectors enable zero-setup tool use for Google Workspace. Cut costs and boost speed—explore seamless integration. Get started today. (517 words) - [Announced at GITEX Global 2025](/newsroom/groq-partners-with-aljammaz-technologies-to-power-ai-inference-across-mena.html): Groq, the inference-first AI company, is proud to partner with Aljammaz Technologies, the region's leading value-added distributor, to bring Groq's high-performance LPU technology to enterprises, governments, and developers across the region. (425 words) - [Key Features of Kimi K2‑0905 on GroqCloud](/blog/introducing-kimi-k2-0905-on-groqcloud/index.html): Experience Kimi K2‑0905 with 256k context and prompt caching on GroqCloud. Build agentic apps faster at lower cost. Try GroqCloud’s new model free today. (244 words) - [GroqCloud](/groqcloud/index.html): The Groq LPU delivers inference with the speed and cost developers need. (329 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/content/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/content/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/content/robots.txt): Crawler directives