Duck Brain AI
Duck Brain is an AI assistant that helps you write SQL queries using natural language. It understands your database schema and generates SQL tailored to your data.
Overview
Section titled “Overview”Duck Brain supports four providers. Pick one in Settings > AI:
- Local server (OpenAI compatible) - Ollama, LM Studio, or any endpoint that speaks the OpenAI API (vLLM, DeepSeek, and so on). Real models at native speed; nothing but the endpoint you set sees your prompts. This is the recommended private option.
- OpenAI - with your own API key.
- Anthropic - with your own API key.
- In-browser (experimental) - WebLLM runs a small model inside the tab via WebGPU. Nothing is installed and nothing leaves the browser, at the cost of a large download, a short context and weaker SQL than a local server.
Privacy: Whatever the provider, Duck Brain sends your question and a summary of your schema, never your data. The one exception is the optional “Explain results” action, which sends a small sample of rows and asks for your consent every time (unless the model runs in the browser).
Browser Requirements
Section titled “Browser Requirements”For the in-browser provider (WebLLM)
Section titled “For the in-browser provider (WebLLM)”The in-browser provider requires WebGPU support:
| Browser | Version | Support |
|---|---|---|
| Chrome | 113+ | Full support |
| Edge | 113+ | Full support |
| Firefox | - | Not supported |
| Safari | - | Not supported |
WebGPU Required: WebGPU is different from WebGL. Chrome/Edge 113+ have WebGPU enabled by default. Check
chrome://gpufor WebGPU status.
For every other provider
Section titled “For every other provider”A local server, OpenAI and Anthropic work in any modern browser.
Getting Started
Section titled “Getting Started”Opening Duck Brain
Section titled “Opening Duck Brain”Duck Brain opens as a panel on the right side of the app, from any of these places:
- the Ask Duck Brain button in the SQL editor toolbar
- the Duck Brain card on the Home tab
Until a provider is configured, the panel offers an Open AI settings button.
Using a local server
Section titled “Using a local server”- Start Ollama or LM Studio on your machine
- Go to Settings > AI and choose Local server (Ollama)
- Pick a preset (Ollama
http://localhost:11434/v1, LM Studiohttp://localhost:1234/v1) or type a base URL - Click Find models and pick one, or type the model name
- Click Test & save
Ollama and deployed Duck-UI: Ollama blocks unknown browser origins by default.
localhostworks out of the box; from a deployed Duck-UI, start Ollama withOLLAMA_ORIGINS=https://your-origin.
Using OpenAI or Anthropic
Section titled “Using OpenAI or Anthropic”- Go to Settings > AI
- Choose OpenAI or Anthropic
- Paste your API key and click Save. The key is verified with a test request and stored encrypted in this browser
- Pick a model
- Return to Duck Brain and start chatting
Using the in-browser provider
Section titled “Using the in-browser provider”- Go to Settings > AI, open In-browser models (experimental) and click Load next to a model
- Wait for the download (about 1 to 2.3 GB depending on the model)
- Once loaded, the model is cached in browser storage for future sessions. Clear cache removes it
- Type your question and press Enter
Available Models
Section titled “Available Models”In-browser models (WebLLM)
Section titled “In-browser models (WebLLM)”| Model | Size | Best For |
|---|---|---|
| Phi-3.5 Mini | ~2.3GB | Best balance of quality and performance, 4k context |
| Llama 3.2 1B | ~1.1GB | Fastest, good for quick queries |
| Qwen 2.5 1.5B | ~1GB | Good balance of size and capability |
Cloud Models
Section titled “Cloud Models”OpenAI: GPT-5.1, GPT-5, GPT-5 Mini (default), GPT-4o
Anthropic: Claude Opus 5, Claude Sonnet 5 (default), Claude Haiku 4.5
Local server: whatever your server offers. Find models lists them.
Using Duck Brain
Section titled “Using Duck Brain”Asking Questions
Section titled “Asking Questions”Duck Brain understands natural language. Just describe what you want:
Show me the top 10 customers by total salesFind all orders from last month where the amount is over $1000What's the average order value by product category?Referencing Tables with @
Section titled “Referencing Tables with @”Use @ to reference specific tables or columns:
Show me all rows from @customers where status is activeJoin @orders with @products and show the top sellersWhen you type @, an autocomplete menu shows available tables and columns.
Schema Awareness
Section titled “Schema Awareness”Duck Brain automatically knows your database schema:
- Table names and their columns
- Column types (VARCHAR, INTEGER, etc.) and nullability
- Approximate row counts for each table
This context helps generate accurate SQL for your specific data. The input shows a rough estimate of the tokens that will be sent (schema context, chat history, system prompt and your message).
Running Generated SQL
Section titled “Running Generated SQL”Every SQL block in a reply has a row of actions:
- Copy the SQL
- Insert replaces the query in the current SQL tab
- New tab opens the SQL in a new tab
- Run executes it right away, with the result shown inline in the chat
Fixing a failed query
Section titled “Fixing a failed query”When a query fails in the SQL editor, the error bar offers Fix with Duck Brain. The suggested fix replaces the query; review it and run again.
Actions on a result
Section titled “Actions on a result”The Duck Brain button in the result panel toolbar offers three tasks:
- Explain results: sends a small sample of rows and asks for consent first
- Optimize query: proposes a faster version of the SQL
- Suggest chart: proposes a chart configuration for the result
Switching Providers
Section titled “Switching Providers”If you have more than one provider configured, a provider selector appears in the Duck Brain header. For OpenAI and Anthropic a model selector sits next to it.
How It Works
Section titled “How It Works”In-browser flow
Section titled “In-browser flow”User Query -> Duck Brain -> WebLLM Engine -> WebGPU -> GPU | Schema Context | SQL Generation | Response- Your question is combined with database schema context
- The local LLM (running via WebLLM) generates SQL
- All processing happens on your GPU via WebGPU
- No data ever leaves your browser
Server flow (local server, OpenAI, Anthropic)
Section titled “Server flow (local server, OpenAI, Anthropic)”User Query + Schema -> API Request -> Provider -> Response- Your question and schema summary are sent to the endpoint
- The model generates SQL
- The response is streamed back to your browser
Privacy: The query and a summary of your schema are sent to the provider. Actual data values are never sent, except by the “Explain results” action, which asks first.
Best Practices
Section titled “Best Practices”Writing Good Prompts
Section titled “Writing Good Prompts”- Be specific: “Show sales by month for 2024” is better than “show me sales”
- Reference tables: Use
@table_nameto be explicit - Describe the output: “as a percentage” or “ordered by date descending”
For Complex Queries
Section titled “For Complex Queries”- Start simple, then refine
- Ask for explanations: “Explain this query”
- Request modifications: “Now add a filter for status = ‘active’”
Performance Tips
Section titled “Performance Tips”-
In-browser models:
- First load downloads 1 to 2.3 GB
- Subsequent loads use the cached model
- GPU memory affects performance
-
Local server:
- Runs full size models at native speed
- Nothing leaves your machine
-
Cloud AI:
- Faster initial response
- No local GPU required
- Requires an internet connection
Troubleshooting
Section titled “Troubleshooting”“WebGPU Not Supported”
Section titled ““WebGPU Not Supported””Problem: Can’t use the in-browser provider
Solutions:
- Update to Chrome/Edge 113+
- Check
chrome://gpufor WebGPU status - Try enabling
#enable-unsafe-webgpuflag - Use a local server or a cloud provider instead; they do not need WebGPU
“Model download failed”
Section titled ““Model download failed””Problem: Can’t download the in-browser model
Solutions:
- Check internet connection
- Clear browser cache and retry
- Try a smaller model (Llama 3.2 1B)
- Check available disk space
“Generation is slow”
Section titled ““Generation is slow””Problem: AI responses take too long
Solutions:
- Try a smaller model
- Close other GPU-intensive applications
- Use a local server or cloud AI for faster responses
- Reduce query complexity
“Connection failed” with a local server
Section titled ““Connection failed” with a local server”Solutions:
- Check the server is running and the base URL ends in
/v1 - From a deployed Duck-UI, set
OLLAMA_ORIGINSto your origin - Click Find models to confirm the server answers
“API key invalid”
Section titled ““API key invalid””Problem: Cloud AI not working
Solutions:
- Verify API key is correct
- Check API key permissions
- Ensure you have API credits
- Try generating a new key
Technical Details
Section titled “Technical Details”WebLLM Integration
Section titled “WebLLM Integration”Duck Brain uses WebLLM for in-browser inference:
- Runs optimized LLMs in browser via WebGPU
- Models are quantized (4-bit) for efficiency
- Cached in browser storage
- Loaded on first use, not at startup, so the app stays small for everyone else
Schema Context
Section titled “Schema Context”Before each generation, Duck Brain builds a schema context in CREATE TABLE form:
CREATE TABLE customers ( id INTEGER NOT NULL, name VARCHAR, email VARCHAR);-- Approximately 1,234 rows
CREATE TABLE orders ( id INTEGER NOT NULL, customer_id INTEGER, amount DECIMAL);-- Approximately 5,678 rowsThis context is prepended to your question so the model understands your data structure. Very large schemas are truncated to fit the context limit.
