Bring Your Own Keys
Cerebras
Configure Cerebras API key for fast open-source inference.
WatchThatSite integrates with Cerebras' API, offering fast inference on open-source models using their optimized hardware.
Get Your API Key
- Go to Cerebras Dashboard or Console
- Sign up or log in with your account
- Navigate to API Keys section
- Click Create New API Key or Generate Key
- Name it (e.g., "WatchThatSite")
- Click Create
- Copy the API key immediately
Add to WatchThatSite
- Go to Settings > AI Providers
- Click Add Provider Key
- Select Cerebras from the provider dropdown
- Paste your API key
- Optionally select specific models
- Click Save
Recommended Models
Latest Cerebras models with tool calling:
- Llama 3.1 70B — Large, powerful open model
- Mixtral 8x7B — Fast, capable MoE model
- Qwen 72B — Alibaba's powerful model
Plus additional open-source models optimized for Cerebras.
All Available Models
Your account has access to all Cerebras models supporting tool calling:
- Go to Settings > AI Providers
- Click Add Provider Key or Edit existing
- Models dropdown shows all available models
- Select specific models to restrict, or leave empty for all
Key Features
- Optimized hardware: Specialized inference infrastructure
- Fast inference: Faster than standard GPU/CPU serving
- Cost-effective: Good pricing on inference
- Open-source models: Community-focused model selection
- Streaming: Real-time token generation
- Reliable: High availability and SLA
Pricing
Cerebras pricing model:
- Per 1M tokens: Standard token-based pricing
- Input vs output: Different rates for input/output tokens
- Varies by model: Different models have different costs
- Free trial: May be available for new accounts
Check Cerebras pricing page for current rates.
Use Cases
Cerebras works well for:
- Fast inference: Optimized hardware means lower latency
- Cost control: Efficient inference reduces costs
- Open models: Preference for open-source and community models
- Predictable performance: Consistent, reliable infrastructure
- Scalability: Handle high-volume requests
Advantages
- Performance: Specialized hardware beats standard solutions
- Open-source focus: Support for community models
- Transparency: Open about infrastructure and pricing
- No lock-in: Models available elsewhere
- Growing capability: Expanding model and feature offerings
Account Setup
Configure Billing
- Go to Account Settings
- Add payment method
- Set spending limits if desired
- Review billing preferences
Monitor Usage
- Go to Usage Dashboard
- View token consumption
- Check costs
- Monitor rate limits
Model Selection
- Speed-sensitive: Choose smaller models
- Quality-sensitive: Choose larger models like LLaMA 70B
- Balanced: Mixtral 8x7B offers good middle ground
- Cost-conscious: Smaller models are cheaper
- Latest research: Cerebras adds new models regularly
Troubleshooting
"Invalid API Key"
- Verify complete key copied without spaces
- Ensure key is active in Cerebras account
- Check key hasn't been revoked
- Regenerate key if needed
"Model Not Found"
- Verify model ID is exactly correct
- Check Cerebras dashboard for available models
- Model names are case-sensitive
- Some models may have limited availability
Rate Limits
Limits depend on account tier:
- Implement exponential backoff
- Space requests appropriately
- Contact Cerebras support for quota increase
- Upgrade account tier if needed
High Latency
While Cerebras optimizes latency:
- Use smaller models for fastest response
- Enable streaming for progressive results
- Check network conditions
- Batch similar requests
Authentication Issues
- Regenerate key in Cerebras console
- Verify no extra spaces in key
- Check correct account is selected
- Wait for key activation (usually instant)
Integration Tips
- Streaming: Enable for real-time token streaming
- Temperature: Lower values (0.2-0.5) for consistent output
- Max tokens: Set reasonable limits
- Batch processing: Group similar requests
- Monitoring: Track usage and costs regularly
Comparing Providers
Cerebras vs similar services:
| Feature | Cerebras | Alternative |
|---|---|---|
| Speed | Optimized hardware | Standard |
| Price | Competitive | Variable |
| Models | Open-source focus | Broader selection |
| Support | Professional | Community |
| SLA | Available | Depends |
When to Use Cerebras
- Latency is critical
- Open-source preference
- Moderate to high volume
- Cost optimization important