WatchThatSite
Bring Your Own Keys

Cerebras

Configure Cerebras API key for fast open-source inference.

WatchThatSite integrates with Cerebras' API, offering fast inference on open-source models using their optimized hardware.

Get Your API Key

  1. Go to Cerebras Dashboard or Console
  2. Sign up or log in with your account
  3. Navigate to API Keys section
  4. Click Create New API Key or Generate Key
  5. Name it (e.g., "WatchThatSite")
  6. Click Create
  7. Copy the API key immediately

Add to WatchThatSite

  1. Go to Settings > AI Providers
  2. Click Add Provider Key
  3. Select Cerebras from the provider dropdown
  4. Paste your API key
  5. Optionally select specific models
  6. Click Save

Latest Cerebras models with tool calling:

  • Llama 3.1 70B — Large, powerful open model
  • Mixtral 8x7B — Fast, capable MoE model
  • Qwen 72B — Alibaba's powerful model

Plus additional open-source models optimized for Cerebras.

All Available Models

Your account has access to all Cerebras models supporting tool calling:

  1. Go to Settings > AI Providers
  2. Click Add Provider Key or Edit existing
  3. Models dropdown shows all available models
  4. Select specific models to restrict, or leave empty for all

Key Features

  • Optimized hardware: Specialized inference infrastructure
  • Fast inference: Faster than standard GPU/CPU serving
  • Cost-effective: Good pricing on inference
  • Open-source models: Community-focused model selection
  • Streaming: Real-time token generation
  • Reliable: High availability and SLA

Pricing

Cerebras pricing model:

  • Per 1M tokens: Standard token-based pricing
  • Input vs output: Different rates for input/output tokens
  • Varies by model: Different models have different costs
  • Free trial: May be available for new accounts

Check Cerebras pricing page for current rates.

Use Cases

Cerebras works well for:

  • Fast inference: Optimized hardware means lower latency
  • Cost control: Efficient inference reduces costs
  • Open models: Preference for open-source and community models
  • Predictable performance: Consistent, reliable infrastructure
  • Scalability: Handle high-volume requests

Advantages

  • Performance: Specialized hardware beats standard solutions
  • Open-source focus: Support for community models
  • Transparency: Open about infrastructure and pricing
  • No lock-in: Models available elsewhere
  • Growing capability: Expanding model and feature offerings

Account Setup

Configure Billing

  1. Go to Account Settings
  2. Add payment method
  3. Set spending limits if desired
  4. Review billing preferences

Monitor Usage

  1. Go to Usage Dashboard
  2. View token consumption
  3. Check costs
  4. Monitor rate limits

Model Selection

  • Speed-sensitive: Choose smaller models
  • Quality-sensitive: Choose larger models like LLaMA 70B
  • Balanced: Mixtral 8x7B offers good middle ground
  • Cost-conscious: Smaller models are cheaper
  • Latest research: Cerebras adds new models regularly

Troubleshooting

"Invalid API Key"

  • Verify complete key copied without spaces
  • Ensure key is active in Cerebras account
  • Check key hasn't been revoked
  • Regenerate key if needed

"Model Not Found"

  • Verify model ID is exactly correct
  • Check Cerebras dashboard for available models
  • Model names are case-sensitive
  • Some models may have limited availability

Rate Limits

Limits depend on account tier:

  • Implement exponential backoff
  • Space requests appropriately
  • Contact Cerebras support for quota increase
  • Upgrade account tier if needed

High Latency

While Cerebras optimizes latency:

  • Use smaller models for fastest response
  • Enable streaming for progressive results
  • Check network conditions
  • Batch similar requests

Authentication Issues

  • Regenerate key in Cerebras console
  • Verify no extra spaces in key
  • Check correct account is selected
  • Wait for key activation (usually instant)

Integration Tips

  • Streaming: Enable for real-time token streaming
  • Temperature: Lower values (0.2-0.5) for consistent output
  • Max tokens: Set reasonable limits
  • Batch processing: Group similar requests
  • Monitoring: Track usage and costs regularly

Comparing Providers

Cerebras vs similar services:

FeatureCerebrasAlternative
SpeedOptimized hardwareStandard
PriceCompetitiveVariable
ModelsOpen-source focusBroader selection
SupportProfessionalCommunity
SLAAvailableDepends

When to Use Cerebras

  • Latency is critical
  • Open-source preference
  • Moderate to high volume
  • Cost optimization important

On this page