Read-only by default · open source

Ask your database anything.

Connect PostgreSQL, MySQL or SQL Server and query in plain English. DBWhisper writes read-only SQL, checks it before it runs, and shows the answer with a chart — no SQL required.

dbwhisper.vercel.app/app
Top 5 products by total revenue

Summary

Aeron Chair leads with $128,400 in revenue, followed by the Standing Desk at $96,220.

policy: allow
SELECT p.name, SUM(oi.qty * oi.price) AS revenue
FROM order_items oi
JOIN products p ON p.id = oi.product_id
GROUP BY p.name ORDER BY revenue DESC LIMIT 5;
productrevenue
Aeron Chair$128,400
Standing Desk$96,220
Monitor Arm$54,900
Desk Lamp$31,050
Cable Kit$18,700

Measured, not claimed

Evaluated on an adversarial policy corpus

Adversarial deny cases denied
343 / 343Adversarial deny cases deniedEvery statement in the deny corpus — 343 case×dialect combinations drawn from 255 cases — was denied by policy sql_policy@2.0.0.
Benign queries still admitted
210 / 210Benign queries still admittedThe corpus checks the other direction too: 210 legitimate read-only expansions were admitted, so the deny figure is not bought with false positives.
Write paths to your data
0Write paths to your dataGenerated and hand-edited SQL take one route: an AST policy check, then execution inside a transaction that is rolled back. On PostgreSQL, MySQL and SQLite that transaction is opened read-only at the database.

Provenance. Corpus app/evaluation/datasets/adversarial/sql_policy_cases.yaml v2.0.0 — 255 distinct cases expanded across PostgreSQL, MySQL, SQL Server and SQLite to 577 case×dialect runs (343 deny, 210 allow, 24 needs-approval). System under test: sql_policy@2.0.0 on sqlglot 30.17.0. No model is involved — this measures a deterministic decision function, so there is no prompt version and no temperature. Reproduce it with uv run pytest tests/sqlpolicy (616 tests, verified 2026-08-24).

Limitations, stated plainly. This measures the policy decision only — nothing in the corpus is sent to a real database. It is a structural filter over a parsed AST, not a proof: the guarantee is bounded by sqlglot’s parse fidelity per dialect. And a corpus measures the attacks someone thought to write down, so 343 of 343 is evidence of no known bypass, not of no bypass.

Why there is no accuracy percentage here. Query-accuracy figures are tracked separately and are re-published only when a run can be reproduced from a clean checkout. Unsafe prompts are handled in two independent places — the model declines, and if it does not, the policy engine denies — and the one recorded end-to-end run only ever exercised the first of those, so it is not evidence about the second. The policy layer is measured on its own corpus above. Full audit: docs/v2/CLAIM_AUDIT.md.

How it works

From question to answer in three steps

1

Enroll a database

Point DBWhisper at PostgreSQL, MySQL and SQL Server (SQLite for local runs) using a read-only role. Enrollment is an API call (POST /schemas/enroll) that introspects the schema through SQLAlchemy reflection; the hosted demo ships with a sample database already enrolled.

2

Ask in plain English

“Top 5 products by revenue this quarter.” No SQL, no schema-hunting — just the question.

3

Get SQL + results you can check

It writes read-only SQL, runs it through the policy engine, shows you the statement, and returns a table, a chart, and a plain-English summary.

Features

Everything you need to check the answer

SQL you can read & edit

The generated SQL is shown with every answer — and you can edit it and re-run it, still through the same read-only policy engine.

Auto-charts & summaries

Results come back as a sortable table, an auto-selected chart, and a one-line plain-English summary.

Conversational follow-ups

It suggests follow-up questions and keeps context across a session so you can drill in.

Multi-engine & exportable

PostgreSQL · MySQL · SQL Server. Download any result as CSV or JSON, or copy it as a Markdown table.

Architecture

What happens when you hit run

  1. You ask in plain English

    “Top 5 products by revenue this quarter.” No SQL, no schema-hunting — just the question.

  2. Relevant tables are retrieved

    pgvector

    Embedding similarity pulls the tables a question needs, so the model sees a focused schema instead of the whole catalog — smaller prompts and fewer invented columns. Retrieval is top-k over per-table summaries; no benchmark on a large schema has been published yet.

  3. SQL is generated

    tool-calling agent

    An LLM agent writes the query with the retrieved tables in context, retrying across whichever of six configured providers have credentials (OpenAI → OpenRouter → DeepSeek → Groq → Anthropic → Gemini).

  4. Checked read-only before it runs

    sql_policy@2.0.0

    A SQLGlot AST policy engine parses the statement and rejects anything that is not a single read-only SELECT — writes, DDL, multi-statements, system-catalog access and blocked functions. This is code between generation and execution, not an instruction in the prompt.

  5. Run & explained

    The statement runs through one audited execution path — least-privilege role, statement timeout, row cap, rolled-back transaction — and comes back as a sortable table, an auto-selected chart, and a one-line summary.

Why it’s built this way

Schema-grounded, not schema-dumped

Instead of pasting a whole schema into every prompt, DBWhisper retrieves only the relevant tables via pgvector embeddings. Smaller prompts, fewer hallucinated columns — and the tables the statement actually touched come back with the answer, so you can check the grounding yourself.

Fail-closed, not prompt-please

“Read-only” is not an instruction the model might ignore — it is a policy engine between generation and execution. A statement the engine will not classify as read-only is not executed. It is a structural filter over a parsed AST, not a proof, so pair it with a least-privilege role.

Provider fallback, not a single point of failure

Six providers are wired in priority order (OpenAI → OpenRouter → DeepSeek → Groq → Anthropic → Gemini); a generation call that fails or is rate-limited moves to the next provider that has credentials, so how many are live depends on which keys you set. The v2 router goes further — it picks a model by the capability a step needs and opens a circuit breaker on one that keeps failing — and its local Ollama profile needs no API key at all.

A real graph, with real pauses

The /v2 API runs the workflow as a LangGraph StateGraph — retrieve → understand → generate → validate → execute → verify → summarize — where a repair loop re-enters validation instead of trusting its own fix. Clarification and approval use LangGraph interrupts against a durable checkpointer, so a run genuinely pauses and resumes after a restart. The console above still calls the v1 tool-calling endpoint.

Security

Read-only by default. You can see exactly what runs.

Read-only enforced in code

Generated and hand-edited SQL take the same path through the policy engine; writes, drops and DDL are rejected there. Execution runs through a least-privilege database role inside a transaction that is rolled back.

The exact SQL, every time

Every answer ships with the statement that produced it, so you can read it, edit it and re-run it through the same policy check.

Least-privilege connections

Connect with a read-only database role; DBWhisper refuses to enroll a connection that probes as writable. The probe is a backstop for the role you configure, not a substitute for it.

No black-box code execution

The model emits SQL and structured JSON, and calls a fixed set of named retrieval tools. There is no path that executes model-authored code on the server.

Scope, because it varies by engine: PostgreSQL, MySQL and SQLite additionally open each query inside a transaction the database itself marks read-only. SQL Server has no session-level read-only mode — there the controls are the policy engine, a least-privilege login and the driver query timeout. Every connection rolls back in a finally block.

Query your data in plain English.

Try it on the built-in sample database — no signup required on the hosted demo.

Open the console →