# AI Tokens
Source: https://docs.cube.dev/admin/account-billing/ai-tokens
Understand how AI token usage works in Cube, including per-seat grants, token packages, and usage tracking.
AI features in Cube use a token-based system to measure and manage consumption.
## Overview
Cube's [AI-powered features][ref-ai-overview] consume tokens based on the
resources required to complete each request. Token allocation differs by
customer type:
* **On-demand customers** receive [per-seat token grants](#per-seat-token-grants)
equal to half of their seat price, with optional
[on-demand consumption](#on-demand-consumption) beyond that
* **Contract customers** can purchase [pooled token packages](#token-packages)
## Token consumption
Token usage depends on several factors:
* **Task complexity** — More complex questions, multi-step analysis, or larger
datasets consume more tokens than simple lookups. Each message in a session
carries prior context, so longer sessions compound usage.
* **Data model context** — Cube sends context from your data model to the LLM
to improve answer accuracy. Larger models with more fields and descriptions
use more tokens per request.
* **AI model** — More capable models consume more tokens per request than
lighter models.
Not all AI features consume tokens. The list of features that consume tokens
is subject to change as the product evolves.
## Per-seat token grants
On-demand customers on paid plans receive **per-seat token grants** equal to
**half of the seat price**. Each user is awarded an individual monthly token
allocation based on their role.
Per-seat grants:
* Are assigned to the individual user
* Reset each billing cycle
* Cannot be transferred, shared, or rolled over
### On-demand consumption
When a user on an on-demand plan exceeds their monthly per-seat grant, usage
automatically continues as **on-demand consumption**. On-demand usage is billed
through the credit card on file.
Administrators can set a **monthly on-demand spending limit** to control
additional costs. This limit caps the total on-demand spend across the account
for each billing cycle.
## Token packages
Contract customers can purchase **pooled add-on token packages**. Token
packages are added to a shared pool accessible by all users in the account.
* Each package is valid for the duration of the contract or until fully
consumed, whichever comes first
* Multiple active packages can be combined
* Packages do not auto-renew
Contact your account executive for details on purchasing token packages.
## Free tier
Each user on a free plan receives an individual monthly token allowance. This
allowance resets at the start of each calendar month.
## Tracking usage
Administrators can monitor token consumption through the **AI Tokens Usage**
tab in the billing settings page. The dashboard shows:
* Total token usage over time
* Remaining allocation from per-seat grants and token packages
* Breakdown by usage dimension
## When limits are reached
When a user exhausts all available token sources (per-seat grant and token
packages), AI requests will return an error indicating the token limit has been
exceeded.
* **Administrators** are directed to the billing page to purchase additional
token packages
* **Non-admin users** are prompted to contact their account administrator to
increase token quotas
## FAQ
### Do tokens roll over?
* **Per-seat grants** reset each billing cycle and do not roll over
* **Token packages** remain active for the duration of the contract
### Can I bring my own AI model?
Yes. When using a [Bring Your Own Model (BYOM)][ref-byom] configuration, AI
requests bypass the token quota system entirely — no tokens are consumed or
tracked for those requests. You are billed directly by your model provider.
### Why does a single prompt sometimes use more tokens than expected?
Cube's AI features use an agentic architecture. A single prompt may trigger
multiple internal steps — such as searching the data model, building a query,
and summarizing results — each of which consumes tokens independently.
### Why does the same question use different amounts of tokens?
Token usage can vary between identical prompts due to differences in
conversation context (earlier messages in the session) or the AI choosing a
different reasoning path.
[ref-ai-overview]: /admin/ai
[ref-byom]: /admin/ai/bring-your-own-model
# API keys
Source: https://docs.cube.dev/admin/account-billing/api-keys
Admins can manage API keys through Admin → API Keys.
Create and revoke API keys for your Cube organization.
## Deployment scope
By default, an API key is **not scoped** to any deployment—it can be used with
every deployment in your organization (the **All deployments** option). When
creating or editing a key, you can instead scope it to one or more **specific
deployments**.
The scope determines which requests a key can authorize:
* **Deployment-scoped endpoints**—such as generating an embed session for a
deployment—require the key to be scoped either to that deployment or to all
deployments. A request targeting a deployment outside the key's scope is
rejected.
* **Endpoints that aren't tied to a specific deployment** continue to work with
any key, regardless of its scope.
Each key's scope—**All** or the specific deployments it's limited to—is shown in
the API keys table.
Scoping is optional and backward compatible. Existing keys remain unscoped (**All
deployments**) until you change them.
# Billing FAQ
Source: https://docs.cube.dev/admin/account-billing/billing-faq
Frequently asked questions about Cube billing, pricing, payments, and invoices.
## Billing basics
### How does the billing structure work?
#### On-demand customers
On-demand customers are billed monthly. The credit card on file is automatically
charged through Stripe upon invoice generation.
* Invoices bill active seats in advance for the upcoming period
* Any seat count changes made during a previously billed period are charged in
arrears on a prorated basis on the next invoice, as well as billed in advance
on all subsequent invoices until the customer makes further adjustments
#### Contract customers
**Contract customers** are billed annually and upfront according to the terms
in their contract.
* Committed seats are billed annually
* Overages for uncommitted seats and usage are billed in arrears separately on a
basis specified in the contract
* Standard payment terms are **Net 30**, unless otherwise specified in the agreement
### What payment methods does Cube accept?
#### On-demand customers
On-demand subscriptions must be paid by credit card via Stripe. Your card is
automatically charged on a recurring subscription basis at the start of each
billing cycle.
Cube does **not** support ACH or wire transfer payments for on-demand payment plans.
#### Contract customers
Contract customers can pay via:
* **Invoice billing** (e.g., Net 30 terms) — Payment terms are defined in your
contract and reflected on your invoice.
* **ACH / wire transfer** — Bank details (Stripe-managed virtual accounts) are
listed directly on the invoice. Please reach out to
[support@cube.dev](mailto:support@cube.dev) if you require validation.
* **Credit card** (via Stripe) — Contract customers may also pay by credit card.
### How can I update my payment information?
To update your payment information, log in to your Cube account and navigate to
the billing settings page.
If you need assistance confirming or updating your account email, please reach
out to [support@cube.dev](mailto:support@cube.dev).
### What happens if my credit card payment fails?
If your credit card payment fails:
1. **We automatically retry the charge** — Stripe will automatically retry the
payment using the card on file. Sometimes failures are temporary (e.g.,
insufficient funds, bank block, expired card).
2. **You may receive an email notification** — If the payment continues to fail,
you'll receive an email prompting you to update your payment method.
3. **Update your card** — You can update your card through the billing settings
page. Once updated, the system will automatically retry the charge.
4. **Continued failed payments** — If payment isn't resolved after multiple
attempts, your account may be restricted or suspended until the outstanding
balance is paid.
# Support in Cube Cloud
Source: https://docs.cube.dev/admin/account-billing/support
Summarizes support channels, coverage hours, and ticket allowances across Cube Cloud product tiers.
Cube Cloud [product tiers][ref-cloud-pricing] include different levels of support
in terms of speed and scope that you should review as you think about what is
right for your business.
## Support by product tiers
### Free
* [Free product tier][ref-free-tier] offers support via **online resources** such as
[documentation][ref-docs-intro], [webinars][cube-webinars], and [community
Slack][cube-slack].
### Starter
* [Starter product tier][ref-starter-tier] includes support via **online resources** such as
[documentation][ref-docs-intro], [webinars][cube-webinars], and [community
Slack][cube-slack].
* You can also connect with our support team during [support hours](#support-hours)
with **2 support tickets per month** through the in-product chat feature.
### Premium
* [Premium product tier][ref-premium-tier] includes support via **online resources** such as
[documentation][ref-docs-intro], [webinars][cube-webinars], and [community
Slack][cube-slack].
* It also includes **either 4 support tickets per month** (on-demand customers)
or **unlimited support tickets** (contract customers) for our support engineers during
[support hours](#support-hours) through our in-product chat.
| Priority | Response time during support hours |
| -------- | ---------------------------------- |
| P0 | 60 minutes |
| P1 | 4 hours |
| P2 | 8 business hours |
| P3 | 2 business days |
### Enterprise
* [Enterprise product tier][ref-enterprise-tier] includes support via **online resources** such as
[documentation][ref-docs-intro], [webinars][cube-webinars], and [community
Slack][cube-slack].
* It also includes **unlimited support tickets** for our support engineers during
[support hours](#support-hours) through our in-product chat with faster
response times as compared to the [Premium product tier](#premium).
| Priority | Response Time during Support Hours |
| -------- | ---------------------------------- |
| P0 | 30 minutes |
| P1 | 2 hours |
| P2 | 8 business hours |
| P3 | 2 business days |
* Enterprise product tier also includes a **dedicated customer success manager** (CSM)
who will provide quarterly reviews, sharing new features and training as well
as usage and optimization advice.
## Support hours
Official support hours are weekdays (Monday through Friday) from 8am ET to 8pm
ET. The above response times are only during support hours.
## Support priority definitions
We prioritize support requests based on their severity, as follows:
* **P0**: The platform is severely impacted or completely shut down. We will
assign specialists to work continuously to fix the issue, provide ongoing
updates, and start working on a temporary workaround or fix.
* **P1**: The platform is functioning with limited capabilities or facing
critical issues preventing a production deployment. We will assign specialists
to fix the issue, provide ongoing updates, and start working on a temporary
workaround or fix.
* **P2**: There are issues with workaround solutions or non-critical functions.
We will use resources during local business hours until the issue is resolved
or a workaround is in place.
* **P3**: There is a need for clarification in the documentation, or a
suggestion for product enhancement. We will triage the request, provide
clarification when possible, and may include a resolution in a future update.
[ref-docs-intro]: /docs/introduction
[cube-webinars]: https://cube.dev/events
[cube-slack]: https://slack.cube.dev
[ref-cloud-pricing]: /admin/account-billing/pricing
[ref-free-tier]: /admin/account-billing/pricing#free
[ref-starter-tier]: /admin/account-billing/pricing#starter
[ref-premium-tier]: /admin/account-billing/pricing#premium
[ref-enterprise-tier]: /admin/account-billing/pricing#enterprise
# Bring Your Own Model
Source: https://docs.cube.dev/admin/ai/bring-your-own-model
Configure custom LLM providers for AI agents in Cube, including supported providers, setup, and billing implications.
Available on the [Enterprise plan](https://cube.dev/pricing).
Bring Your Own Model (BYOM) lets you connect your own LLM provider to power
AI agents in Cube, instead of using the built-in models. This gives you full
control over which models your agents use, where your data is processed, and
how you manage AI costs.
## Supported providers
| Provider | Chat models | Embedding models |
| -------------------- | ----------- | ---------------- |
| **Anthropic** | Yes | No |
| **OpenAI** | Yes | Yes |
| **AWS Bedrock** | Yes | Yes |
| **GCP Vertex AI** | Yes | No |
| **Databricks** | Yes | No |
| **Snowflake Cortex** | Yes | No |
## Configuration
### Step 1: Add a model
Before assigning a BYOM model to an agent, you need to register it in the
admin panel:
1. Navigate to **Admin > Models**
2. Click **Add Model**
3. Provide a **name** for the model
4. Select the **model type** (LLM or Embedding)
5. Choose a **provider** and **model**
6. Enter the required credentials for the provider
### Step 2: Assign the model to an agent
Once a model is registered, reference it in the agents YAML configuration by
name or ID:
```yaml theme={"dark"}
agents:
- name: sales-analyst
llm:
byom:
name: "my-anthropic-model"
embedding_llm:
byom:
name: "my-bedrock-embeddings"
```
Each agent can use a different model. If no BYOM model is specified, the agent
uses the built-in default.
Switching embedding models for an agent means existing memories stored with
the previous embedding model will not be compatible. Memories are tied to the
embedding model that created them.
## Network configuration
When using BYOM, Cube connects to your model provider from its control plane.
If your provider requires IP allowlisting, ensure the Cube outbound IP
addresses are added to your allowlist.
For agents running in dedicated regions, additional per-region IP addresses
may also need to be allowlisted.
## Billing
When using a BYOM model, **Cube AI tokens are not consumed**. You are billed
directly by your model provider based on their pricing.
This means:
* No Cube token quota is deducted for BYOM chat requests
* No token usage is tracked in the AI Tokens Usage dashboard for BYOM requests
* Per-seat token grants and token packages do not apply
See [AI Tokens][ref-ai-tokens] for details on how token billing works with
built-in models.
## Provider-specific notes
### Anthropic
Supports extended thinking mode for compatible models. Configure this in the
model settings when creating the model.
### AWS Bedrock
* Credentials are optional — if left empty, the default AWS credential chain
is used (e.g., workload identity)
* Supports assume-role configuration for cross-account access
* Supports inference profiles
### GCP Vertex AI
Requires a service account JSON key for authentication.
### Databricks
Requires a workspace URL and access token.
### Snowflake Cortex
Supports two authentication methods:
* JWT authentication
* Key-pair authentication (requires an encrypted PKCS#8 PEM private key)
## Troubleshooting
### Rate limit errors
If you see rate limit errors, the limits are enforced by your model provider,
not by Cube. Check your provider's rate limits and usage quotas.
### Authentication errors
Verify that the API key or credentials configured for the model are valid and
have the necessary permissions.
### Model not found
Ensure the model ID configured in Cube matches a valid model offered by your
provider. Model availability may vary by region.
[ref-ai-tokens]: /admin/account-billing/ai-tokens
# Certified queries
Source: https://docs.cube.dev/admin/ai/certified-queries
Provide a library of trusted SQL queries that the agent uses as reference examples when answering user requests.
Certified queries are pre-approved SQL queries that the agent treats as a **library of trusted examples**. They guide the agent toward correct business logic and well-formed query patterns without restricting it to a fixed set of queries.
Certified queries are configured as code in your [data model repository](/admin/ai#agent-configuration), alongside your cubes and views.
The agent is not limited to running only certified queries. When answering a user request, the agent may:
* Use a certified query directly if it matches the request
* Use a certified query as a **starting point** and adapt it (e.g., add filters, dimensions, or measures) to fit the user's question
* Generate a new query independently if no certified query is relevant
Certified queries help the agent toward trusted patterns and correct business logic, but the agent retains full flexibility to construct the query that best answers each question.
## Defining certified queries
Certified queries are defined as Markdown files under `agents/certified_queries/`. Each query lives in its own file: the YAML frontmatter holds metadata, and the Markdown body is the SQL query.
```markdown theme={"dark"}
---
description: "Get revenue by fiscal quarter"
user_request: "What is the revenue by quarter?"
---
SELECT
DATE_TRUNC('quarter', order_date) AS quarter,
SUM(amount) AS revenue
FROM orders
WHERE status != 'cancelled'
GROUP BY 1
ORDER BY 1
```
Files placed under a `certified_queries/`, `certified-queries/`, or `queries/` directory are treated as certified queries automatically — no `kind` property is required. The `name` is inferred from the file name (e.g., `quarterly-revenue.md` → `quarterly-revenue`).
### Frontmatter properties
| Property | Type | Required | Description |
| -------------- | ------ | :------: | ---------------------------------------------------------- |
| `name` | string | No | Unique identifier. Inferred from the file name if omitted. |
| `description` | string | No | Human-readable description. |
| `user_request` | string | Yes | The user request pattern this query answers. |
| `sql_query` | string | No | The SQL query. Falls back to the Markdown body if omitted. |
### Inlining certified queries in YAML
You can also inline certified queries directly in `agents/config.yml` under a `certified_queries` key:
```yaml theme={"dark"}
# agents/config.yml
certified_queries:
- name: total-revenue
description: "Calculate total revenue"
user_request: "What is the total revenue?"
sql_query: "SELECT SUM(amount) AS total_revenue FROM orders WHERE status = 'completed'"
- name: monthly-sales
description: "Monthly sales breakdown"
user_request: "Show me sales by month"
sql_query: |
SELECT
DATE_TRUNC('month', order_date) AS month,
SUM(amount) AS total_sales
FROM orders
GROUP BY 1
ORDER BY 1
```
Inline certified queries accept the same properties as Markdown ones — `name`, `description`, `user_request` (required), and `sql_query` (required).
Certified queries inlined at the root of `agents/config.yml` are attached to the implicit `auto` space and applied to the default agent in a [single-agent setup](/admin/ai). In a [multi-agent setup](/admin/ai/multi-agent), attach certified queries to a specific space by inlining them under that space's `certified_queries` key (or by placing Markdown files under `agents/certified_queries//`).
## Writing effective certified queries
* **Phrase `user_request` like a real user question.** This is what the agent matches against incoming requests.
* **Keep queries focused.** A certified query should answer one well-defined question; the agent will adapt it for variations.
* **Use canonical business logic.** Certified queries are the reference for "the right way" to compute something — encode the definitions you want the agent to follow.
* **Cover common patterns first.** Start with the questions users ask most often, and grow the library based on usage.
# Evals
Source: https://docs.cube.dev/admin/ai/evals
Benchmark your agent's answers against a known-correct ground truth and track accuracy across data-model and agent changes.
Evals let you benchmark your agent's answers against a known-correct ground
truth, on any branch. You author a set of questions, each with the SQL or
[certified query](/admin/ai/certified-queries) that represents the right
answer, run your agent against them, and get a per-question pass/fail plus an
accuracy score for the run — so you can see, objectively, whether a data-model
or agent change made the agent better or worse.
You'll find evals in the model IDE under the **Evals** tab, with two
sub-tabs: **Evals** (runs) and **Questions** (the benchmark set).
## Concepts
| Term | What it is |
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Question** | A natural-language question plus its **ground truth** (the correct answer, as SQL or a certified-query reference). Authored as code in your data model. |
| **Eval (run)** | One execution of the agent against the whole question set, on a specific branch and agent. |
| **Result** | The agent's answer to a single question in a run, graded against that question's ground truth. |
| **Accuracy** | `passed / total` for a run, shown as `NN% (passed/total)`. |
## Authoring benchmark questions
Questions live in your [data model repository](/admin/ai#agent-configuration),
versioned and branched like the rest of it. You can keep them in a single
top-level `agents/eval_questions.yml` file — the simplest place to start — or
split them across any number of `agents/eval_questions/*.yml` files as your set
grows. The parser picks up both and merges every file's `eval_questions` list
into one set, so you can move from one file to many at any time without changing
anything else.
Each file has a top-level `eval_questions` list. A question needs a unique
`name`, a `question`, and exactly one ground truth: a `certifiedQuery`
reference **or** inline `sql`.
```yaml theme={"dark"}
# agents/eval_questions.yml
eval_questions:
- name: revenue_by_quarter
question: What was our revenue by quarter over the last two years?
certifiedQuery: revenue_by_quarter # reference an existing certified query by name
- name: arr_last_4_years
question: What was our ARR over the last 4 years?
sql: | # ...or inline SQL ground truth
SELECT date_trunc('year', created_at) AS year, SUM(arr) AS arr
FROM subscriptions GROUP BY 1 ORDER BY 1
```
* `certifiedQuery` references a [certified query](/admin/ai/certified-queries)
by name. Define it under `agents/certified_queries/` (or via **Certify this
query** in chat). A reference that doesn't resolve to an existing certified
query is flagged as a validation error.
* `sql` is inline ground-truth SQL, run through the same Cube SQL API the agent
uses (so `MEASURE(...)` and friends work).
* Omitting both — or setting both — is a validation error.
* An optional top-level `space` key scopes a file's questions to a named space
(defaults to `auto`). Question names are unique per space.
The **Questions** tab is a read-only view of these files. To add or edit
questions, edit the YAML in the IDE — there's no in-product question editor
yet.
## Running an eval
On the **Evals** tab, click **Run eval** and choose:
* **Branch** — which branch's data model and agent configuration to run
against. Defaults to the active branch.
* **Agent** — `auto` (the implicit auto-agent) or a configured agent name.
The run starts immediately and you can close the dialog — it executes in the
background. The run list shows live progress and then the outcome:
| Column | Meaning |
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Eval run** | When the run was created. |
| **Environment** | Where it ran — **dev** (your personal dev-mode branch, shown as "*Name* Dev Mode"), **staging**, or **prod** (the deploy branch, e.g. `master` or `main`). |
| **Agent** | The agent used. |
| **Execution status** | Running, Completed, or Failed. |
| **Accuracy** | `NN% (passed/total)`. |
| **Created by** | Who triggered the run. |
| **Last updated** | When it finished. |
## Reading the results
Open a run to see per-question results: the question list on the left, with a
pass/fail icon for each, and the selected question's detail on the right.
* **Assessment** — `pass`, `fail`, `review`, or `error`.
* **Score reason** — when a question doesn't pass, a tag categorizing why:
**Row count mismatch**, **Missing columns**, **Value mismatch**,
**Unexpected rows**, **Query error**, **Ground truth query failed**,
**Ground truth not found**, or **Agent error**.
* **Failure analysis** — a plain-English explanation, e.g. *"The agent
returned 3 rows, but the ground truth has 5 rows."*
* **Model output · SQL** vs. **Ground truth SQL answer** — the agent's query
side-by-side with the ground truth, so you can spot the difference.
* **Response** — the agent's full text answer, rendered as Markdown.
## How grading works
Grading is execution-based, not text-based — the same approach used by
industry text-to-SQL benchmarks such as BIRD and Spider 2.0. The agent's SQL
and the ground-truth SQL are both executed, and their result sets are
compared. So an answer that's worded or written differently but produces the
same data still passes.
The comparison is:
* **Sort-invariant** — row order never matters.
* **Numeric-tolerant** — values are compared to 4 significant figures, so
float/representation noise (`6646` vs. `6646.0`) doesn't fail.
* **Column-name-agnostic and lenient on extra columns** — each ground-truth
column must be reproduced by some agent column, matched by its values, so
`revenue` vs. `total` aliases don't matter. Extra columns the agent adds are
ignored.
* **No standalone row-count gate** — row count falls out of the comparison: a
"top 5" question is enforced because the golden result has exactly 5 rows.
Verdicts:
| Verdict | When |
| ---------- | ----------------------------------------------------------------------------------------------------------------------- |
| **pass** | The agent's result set matches the ground truth. |
| **fail** | It ran but the result set doesn't match (see the score reason). |
| **review** | Nothing to compare automatically — the question has no ground truth, or the agent didn't run a query. Compare manually. |
| **error** | The agent run failed, the ground-truth query failed, or a referenced certified query wasn't found. |
## Limitations
* Questions are authored as code only; the **Questions** tab is read-only.
* Very large question sets can be slow to run.
* Grading is execution-based on the result set; it does not semantically judge
prose answers.
# Overview
Source: https://docs.cube.dev/admin/ai/index
Configure the agent that ships with every Cube deployment — accessible views, LLM, runtime, and memory.
Every Cube deployment ships with an agent that powers AI features such as [Analytics Chat](/docs/explore-analyze/analytics-chat). You can customize the agent's behavior with rules, certified queries, accessible views, model selection, and memory configuration.
Configuring the agent in the UI is **deprecated** in favor of code-first configuration. Cube version 1.6.5 or above is required.
Deployments created after **April 30, 2026** have code-first agent configuration enabled by default. Existing deployments need to opt in by setting `CUBE_CLOUD_AGENTS_CONFIG_ENABLED=true`. The directory path defaults to `agents` and can be overridden with `CUBE_CLOUD_AGENTS_CONFIG_PATH`.
## Agent configuration
Agent configuration lives in your **Cube data model repository**, alongside your cubes and views. Configuration is defined as YAML and Markdown files under an `agents/` directory in your project:
```text theme={"dark"}
your-cube-project/
├── model/
│ └── cubes/
│ └── orders.yml
├── agents/
│ ├── config.yml # Agent configuration
│ ├── rules/ # Rules as Markdown
│ │ └── fiscal-year.md
│ └── certified_queries/ # Certified queries as Markdown
│ └── quarterly-revenue.md
└── cube.py
```
Storing configuration as code in your repository enables version control, code review, and consistent behavior across environments.
## Configure the agent
Place agent properties at the root of `agents/config.yml`:
```yaml theme={"dark"}
# agents/config.yml
llm: claude_4_sonnet
runtime: plain
accessible_views:
- orders_view
- customers_view
memory_mode: user
```
### Properties
| Property | Type | Default | Description |
| ------------------ | ---------------- | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `llm` | string or object | `auto` | LLM provider — `auto`, a [predefined model](#llm) name, or a [BYOM](/admin/ai/bring-your-own-model) reference. |
| `embedding_llm` | string or object | `text-embedding-3-large` | Embedding model — a predefined name or a BYOM reference. |
| `runtime` | string | `plain` | [Runtime mode](#runtime) — `plain` or `reasoning`. |
| `accessible_views` | array | *all views* | List of view names exposed to the agent as context guidance for which views to query. This is not a security control — if omitted or empty, all views are exposed. Use [access policies](/reference/data-modeling/data-access-policies) to enforce actual access restrictions. |
| `memory_mode` | string | `space` | [Memory isolation mode](#memory) — `space`, `user`, or `disabled`. |
Rules and certified queries have their own dedicated pages:
* [Rules](/admin/ai/rules) — instructions that guide the agent's behavior
* [Certified queries](/admin/ai/certified-queries) — a library of trusted SQL examples for the agent
* [Skills](/admin/ai/skills) — reusable, named agent workflows users can run on demand from chat
## LLM
The default value is `auto` — Cube picks a recommended model on your behalf and may change it as better models become available. Use `auto` unless you have a reason to pin a specific model.
To pin a specific model, set the `llm` property to one of the predefined models:
**Anthropic Claude:**
* `claude_3_5_sonnetv2`
* `claude_3_7_sonnet`
* `claude_3_7_sonnet_thinking`
* `claude_4_sonnet`
* `claude_4_5_sonnet`
* `claude_4_5_haiku`
* `claude_4_5_opus`
* `claude_4_6_sonnet`
* `claude_4_6_opus`
* `claude_4_7_opus`
* `claude_4_8_opus`
* `claude_5_sonnet`
* `claude_5_opus`
**OpenAI GPT:**
* `gpt_4o`
* `gpt_4_1`
* `gpt_4_1_mini`
* `gpt_5`
* `gpt_5_mini`
* `gpt_5_3`
* `gpt_5_4`
* `o3`
* `o4_mini`
To use your own model, see [Bring your own model](/admin/ai/bring-your-own-model).
### Embedding models
Predefined embedding models for the `embedding_llm` property:
* `text-embedding-3-large`
* `text-embedding-3-small`
BYOM is also supported for embedding models.
## Runtime
The `runtime` property controls how the agent processes requests:
| Mode | Description |
| ----------- | ---------------------------------------------------------------------- |
| `plain` | Default. Optimized for speed and cost. Recommended for most use cases. |
| `reasoning` | Enables extended thinking for complex analysis. |
`reasoning` is **experimental**. It may be unstable and is not yet feature-complete. Stick with `plain` unless you have a specific need for extended reasoning.
## Memory
The `memory_mode` property controls how the agent persists context across conversations:
| Mode | Description |
| ---------- | ------------------------------------------------------------------------------------------- |
| `space` | Default. Memories are shared across all users — useful when the agent serves a single team. |
| `user` | Memories are isolated per user — useful when each user has private context. |
| `disabled` | The agent does not persist memory between conversations. |
See [Memories](/admin/ai/memory-isolation) for details on how memories are stored and used.
## Customize the agent
Define instructions that guide how the agent responds and analyzes data.
Provide a library of trusted SQL queries for the agent to reference.
Package reusable, named agent workflows that users can run on demand from chat.
Configure the agent to use your own LLM provider or model.
Control how the agent persists context across conversations and users.
# MCP Connectors
Source: https://docs.cube.dev/admin/ai/mcp-connectors
Connect external MCP servers so the agent can use their tools — search Notion, file Linear issues, query Sentry, and more — from chat.
MCP Connectors are available on [Premium and Enterprise plans](https://cube.dev/pricing).
MCP Connectors let the [agent](/admin/ai) use tools from external services. An
administrator connects an external [MCP](https://modelcontextprotocol.io) server — such as
Notion, Linear, Sentry, or Attio — and its tools become available to the agent in
[Analytics Chat](/docs/explore-analyze/analytics-chat). The agent can then search a Notion
workspace, file a Linear issue, look up a Sentry error, or call any tool a connected server
exposes, alongside the data it queries from your semantic model.
**MCP Connectors are the inverse of the [MCP server](/docs/integrations/mcp-server).** A
connector lets the **Cube agent reach out** to an external MCP server and call its tools.
The MCP server lets **external MCP clients reach in** to Cube and query your data. One is
outbound, the other inbound — you can use either or both.
## Concepts
* A **connector** is a connection to one external MCP server. Each connector is configured
and authenticated once, at the organization level, by an administrator.
* Each connector exposes one or more **tools** — the individual actions the agent can call
(for example, "search pages" or "create issue"). A single connector typically exposes
many tools.
* Connectors are managed in the admin panel under **MCP Connectors** and apply across the
organization, so every agent can use the tools you enable.
## Adding a connector
Open the admin panel and go to **MCP Connectors**. You can add a connector from the
built-in directory or connect a custom MCP server.
### From the directory
The connector directory includes vetted, first-party integrations with streamlined setup.
Select **Browse directory** and choose a service (for example, Notion, Linear, Sentry,
or Attio).
Complete the connector's authentication flow (see [Authentication](#authentication)
below). For OAuth-based connectors you'll be redirected to the provider to authorize
access.
Review the tools the connector exposes and choose which ones are available to the
agent. See [Choosing which tools are available](#choosing-which-tools-are-available).
### Custom connector
Any service that exposes a remote MCP endpoint can be connected as a custom connector.
Select **Add custom connector** and provide a name and the server's HTTPS endpoint URL
(for example, `https://mcp.example.com/mcp`).
Choose the authentication method the server requires — OAuth or a user-provided
credential such as an API key or token.
Once connected, the server's tools are discovered automatically. Choose which ones the
agent may call.
## Authentication
Connectors authenticate to the external service in one of two ways, depending on what the
service supports:
| Method | How it works |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **OAuth** | You authorize Cube with the provider through a standard OAuth flow. The connector stores the resulting tokens and refreshes them as needed. Used by most directory connectors. |
| **User-provided credential** | You supply a credential — such as an API key or access token — that the connector uses to authenticate. Used when a service does not offer OAuth. |
A connector's status indicator in the connectors list shows whether it is connected and
authenticated.
## Choosing which tools are available
Each connected server reports the full set of tools it exposes (shown as a count, for
example `16 / 16`). You control which of those tools the agent is allowed to call.
Enabling only the tools you need keeps the agent focused and limits what it can do through
each connector.
## How the agent uses connector tools
Once a connector is configured and its tools are enabled, the agent can call them in
[Analytics Chat](/docs/explore-analyze/analytics-chat) as part of answering a request — the
same way it queries your semantic model. The agent decides when a tool is relevant based on
the user's request and the tool's description. If a tool requires the user to authenticate
to the external service, the agent prompts for authorization in chat before the tool runs.
## Permissions
Managing MCP Connectors requires administrator access in Cube Cloud — the same access
needed to manage other organization-level settings in the admin panel. Connectors apply
across the organization; using the tools they expose is available to anyone with chat
access, subject to the tools you enable.
## Related
Let external MCP clients connect to Cube and query your data over HTTPS.
Configure the agent that powers Cube's AI features.
Package reusable, named agent workflows users can run on demand from chat.
Ask questions of your data in natural language.
# Memories
Source: https://docs.cube.dev/admin/ai/memory-isolation
How the agent stores, retrieves, and isolates memories across tenants, spaces, and users.
Memory in Cube allows Agents to learn from and recall past interactions, user preferences, and contextual information. When users interact with Agents, relevant information is stored as memories that can be retrieved in future conversations, enabling more personalized and context-aware responses. Memories help Agents understand user preferences, remember previous decisions, and provide continuity across sessions.
Memories are scoped and enforced at the Tenant/Space boundary. Cube applies both application-layer and (optionally) infrastructure-layer isolation so one customer's end users cannot see another customer's memories.
## How It Works
### Space-Scoped Memories
Cube memories are stored at the Space level. Agents only learn from and retrieve memories within the current Space, ensuring no cross-Space exposure by design.
Because a Space can be [scoped globally or to a single deployment](/admin/ai/multi-agent#space-scope), that boundary also decides whether memories are shared across deployments: agents in different deployments that use the same global Space share its memories, while a Space scoped to one deployment keeps them to that deployment.
### Tenant-Aware Security Context
Every request carries a tenant-bound security context (JWT). Cube maps that context to an app/tenant ID and uses it across caching, orchestration, and query flows. This is the backbone of multi-tenancy isolation.
### RBAC and Policy Guardrails
Role-based access policies gate what entities and content are visible within a tenant. These same guardrails apply to what agents can read and write as memories.
### Data Model and API Isolation
Even when using the SQL API or custom views, hidden members and non-public entities remain inaccessible. Multitenancy configuration ensures queries and artifacts resolve only within the current tenant's scope.
### Optional Infrastructure Isolation
Many customers run in multi-tenant regions, but single-tenant infrastructure and BYOC (Bring Your Own Cloud) variants are available. These provide stronger blast-radius isolation at the cluster, storage, and key-management levels.
## Practical Implications
* **Tenant Separation**: An end user in Customer A can only create and retrieve memories in Customer A's Spaces
* **Cross-Tenant Protection**: Memories are not retrievable by Customer B's users or agents, because requests from B carry a different security context and resolve to different Space and tenant IDs
* **Intra-Tenant Controls**: Even within a customer, RBAC and policies can further restrict which users or agents can contribute to or benefit from memories
## Technical Implementation
Cube ensures memory isolation through multiple layers of security controls:
1. **Tenant Isolation**: Every request is scoped to a specific tenant via JWT and middleware
2. **User Isolation**: Additional user-level filtering for user-mode memories
3. **Automatic Filtering**: Database queries are automatically filtered by tenant using decorators
4. **Vector Store Security**: All vector searches include tenant and user filters
5. **Container Isolation**: Each tenant gets its own dependency injection container
6. **JWT Security**: All security context comes from cryptographically signed JWT tokens
7. **Memory Mode Controls**: Configurable memory isolation levels (user/space/disabled)
# Multi-agent
Source: https://docs.cube.dev/admin/ai/multi-agent
Configure multiple agents within a single deployment, each with its own model, accessible views, rules, and certified queries.
Every Cube deployment ships with a single agent by default — see the [Overview](/admin/ai) for the standard configuration. For more advanced setups, you can configure **multiple agents** within the same deployment, each with its own model, accessible views, rules, and certified queries.
Multi-agent is useful when:
* Different teams need agents tuned to their own domain (e.g., a Sales Assistant and a Marketing Analyst).
* You want specialized agents with distinct instructions or tool access in the same deployment.
* You need to isolate context (rules, certified queries, memories) between user groups.
`agents/config.yml` is the source of truth for spaces and agents: declaring them there and [reconciling](#reconciliation) is the only supported way to create either. Agents can't be created in the UI, and their settings are read-only there. Cube links every record it creates to the entry it came from, and matches them by that entry's `name`.
## Architecture
A multi-agent setup introduces one new concept on top of the single-agent model: **spaces**.
* A **space** belongs to the account and is used by one or more **deployments** — see [Space scope](#space-scope).
* A **space** is an isolated context. It owns its **rules**, **certified queries**, and **memories** — they are not shared across spaces.
* Each **agent** belongs to exactly one space and inherits everything that space owns. Multiple agents can live in the same space and share the same rules, certified queries, and memories.
This means you choose a space's boundary based on what context should be shared. Two agents serving the same team usually live in one space; agents serving different domains (Sales vs. Marketing) live in different spaces so their context stays separate.
```mermaid theme={"dark"}
flowchart LR
classDef deployment fill:#4f46e5,stroke:#3730a3,color:#fff
classDef space fill:#e0e7ff,stroke:#6366f1,color:#1e1b4b
classDef agent fill:#10b981,stroke:#047857,color:#fff
classDef resource fill:#fef3c7,stroke:#d97706,color:#78350f
D[Deployment]:::deployment
SA[Space sales-analytics]:::space
SM[Space marketing-analytics]:::space
SAR[Rules]:::resource
SAC[Certified queries]:::resource
SAM[Memories]:::resource
A1[Agent sales-assistant]:::agent
A2[Agent sales-reporter]:::agent
SMR[Rules]:::resource
SMC[Certified queries]:::resource
SMM[Memories]:::resource
A3[Agent marketing-analyst]:::agent
D -->|uses| SA
D -->|uses| SM
SA --- SAR
SA --- SAC
SA --- SAM
SA --> A1
SA --> A2
SM --- SMR
SM --- SMC
SM --- SMM
SM --> A3
```
In the [single-agent setup](/admin/ai), there is an implicit `auto` space that holds all rules, certified queries, and memories — you don't need to think about it. In a multi-agent setup, you define spaces explicitly and attach rules and certified queries to specific spaces.
## What changes from the single-agent setup
The [`agents/` file structure](/admin/ai#agent-configuration) is the same. What's different is how `agents/config.yml` is shaped:
1. **Agents are defined as an array.** Each agent gets a unique `name` and an optional `description`, in addition to the standard agent [properties](/admin/ai#properties) (`llm`, `runtime`, `accessible_views`, `memory_mode`, etc.).
2. **Spaces are introduced.** A `spaces` array defines the contexts agents operate in. Each space gets a unique `name`.
3. **Rules and certified queries attach to spaces.** Use the `space` property in the frontmatter of each rule or certified query Markdown file to attach it to a specific space.
You cannot mix flat root-level properties (the single-agent style) with `spaces` or `agents` arrays in the same file. Use one style per file.
## Agents
Replace the flat root-level agent properties with an `agents` array:
```yaml theme={"dark"}
# agents/config.yml
agents:
- name: sales-assistant # Required
description: "AI assistant for sales analytics"
space: sales-analytics # Required: reference to a space
llm: claude_4_6_sonnet
accessible_views:
- orders_view
- customers_view
memory_mode: user
- name: marketing-analyst # Required
description: "AI assistant for marketing analytics"
space: marketing-analytics # Required
llm: gpt_5
```
The properties available on each agent are the same as in the [single-agent setup](/admin/ai#properties), plus:
| Property | Type | Required | Description |
| ------------- | ------ | :------: | --------------------------------------------------- |
| `name` | string | Yes | Unique identifier for the agent. |
| `description` | string | No | Human-readable description. |
| `space` | string | Yes | Name of the [space](#spaces) this agent belongs to. |
## Spaces
A space is the context an agent operates in. Spaces own the rules, certified queries, and memories that the agents inside them share. Define spaces alongside agents:
```yaml theme={"dark"}
# agents/config.yml
spaces:
- name: sales-analytics # Required
description: "Space for sales team analytics and reporting"
- name: marketing-analytics # Required
description: "Space for marketing team analytics"
```
| Property | Type | Required | Description |
| ------------- | ------ | :------: | -------------------------------- |
| `name` | string | Yes | Unique identifier for the space. |
| `description` | string | No | Human-readable description. |
Each agent must reference exactly one space via its `space` property. Multiple agents can share the same space and inherit its rules, certified queries, and memories.
## Reconciliation
A space or agent declared in `agents/config.yml` needs a matching record in Cube before users can chat with it. Cube creates those records from your config — that step is **reconciliation**:
Add the `spaces:` and `agents:` entries to `agents/config.yml`. In dev mode the pending list reflects your branch, so you can reconcile before merging.
Cube compares the config with the records already available to that deployment, matching each entry by `name`. Entries with no record yet are listed under **Pending Configurations** on the **Agents** page and on **Agents** → **Spaces**, once you pick the deployment. The Semantic Model IDE also shows a **Reconcile agent configs** button with the pending count that links there.
Choose the [space scope](#space-scope) and press **Create All**. Spaces are created first, then each agent is linked to the space its `space` property names.
Reconciliation only creates missing spaces and agents, so you need it only when an entry has no record available to the deployment yet: after adding a `name` to the config, and on each deployment you reconcile for the first time — an agent belongs to one deployment, as does a space created per deployment, while a global space counts as created everywhere. Agent behavior — `llm`, `description`, `accessible_views`, `memory_mode`, rules, certified queries — is read from the `agents/` directory of the deployment's data model and takes effect without reconciling. The implicit `auto` space and agent of the [single-agent setup](/admin/ai) are never listed as pending.
An agent's `space` is the exception: the link is made when the agent is created, so changing it in YAML doesn't move an existing agent. The agent's page flags the mismatch between the space its config names and the space it is linked to. If the newly named space has no record yet, it appears under **Pending Configurations**, and **Create All** creates it and moves the agent onto it in the same action. If that space is already available to the deployment, nothing about the agent is pending and **Create All** won't move it — delete the agent so its config reads as pending again, then **Create All** recreates it in the space the config names. The recreated agent is a new agent, so chats from before the delete don't carry over to it.
Changing an entry's `name` reads as a new entry: reconciling creates a new space or agent, and the record created from the old name stays behind with everything tied to it — an agent's chats, a space's memories — flagged **(misconfigured)** because its config no longer exists. Delete it from the **Agents** or **Spaces** page once you no longer need it.
## Space scope
Spaces live at the account level, so one space can be used by agents in more than one deployment. When a space is created, you choose its scope:
* **Global** — one space that agents in every deployment can use.
* **Per deployment** — the space is available to a single deployment only.
You pick the scope when you [reconcile](#reconciliation): the **Create All** panel's **Create Spaces** control offers **Global** and **Per deployment**. You can change the scope of an existing space later on its page, under **Agents** → **Spaces**.
Spaces created per deployment are named after the deployment, for example `Product (production)` and `Product (staging)`. Only the displayed name changes — the space stays linked to the `spaces:` entry it was created from, so the YAML entry keeps matching.
Choose per-deployment scope when the same `spaces:` entry is declared in several deployments — typically development, staging, and production fed from branches of one data model — and you don't want them sharing what the space stores.
Creating the spaces again does not split a space that was created as global: it already counts as created for every deployment, so **Create All** creates nothing new. To split one, re-scope it to a deployment on its space page, then run **Create All** for the other deployments.
Spaces created before this option existed have no scope stored. Cube infers the deployment they belong to from the agents linked to them, and the Spaces list shows them as **Not scoped**. Set the scope explicitly on the space page to make it definite.
### Scope affects storage, not configuration
Scope never changes where configuration comes from. Agents always belong to a single deployment, and the `spaces:` and `agents:` entries, rules, and certified queries that shape an agent are read from the `agents/` directory of the data model of that deployment, on that deployment's branch. A global space does not merge the configuration of the deployments that use it.
What a global space shares is storage: the space itself and the data held against it — memories in particular. Two deployments using the same global space read and write the same memories, while each of them still applies its own `agents/` configuration.
## Attaching rules and certified queries to a space
In the single-agent setup, [rules](/admin/ai/rules) and [certified queries](/admin/ai/certified-queries) belong to the implicit `auto` space. In a multi-agent setup, you must attach each rule and certified query to a specific space using the `space` property in the Markdown frontmatter:
```markdown theme={"dark"}
---
space: sales-analytics
type: always
---
Always use fiscal year starting April 1st when analyzing dates.
```
```markdown theme={"dark"}
---
space: sales-analytics
description: "Apply when the user asks about quarterly revenue"
user_request: "What is the revenue by quarter?"
---
SELECT
DATE_TRUNC('quarter', order_date) AS quarter,
SUM(amount) AS revenue
FROM orders
WHERE status != 'cancelled'
GROUP BY 1
ORDER BY 1
```
You can also organize rules and certified queries into space-named subdirectories. Files placed under `agents/rules//` or `agents/certified_queries//` are attached to that space automatically — no `space` frontmatter required.
## Complete example
```yaml theme={"dark"}
# agents/config.yml
spaces:
- name: sales-analytics
description: "Space for sales team analytics and reporting"
- name: marketing-analytics
description: "Space for marketing team analytics"
agents:
- name: sales-assistant
description: "AI assistant for sales analytics and reporting"
space: sales-analytics
llm: claude_4_6_sonnet
accessible_views:
- orders_view
- customers_view
- products_view
memory_mode: user
- name: marketing-analyst
description: "AI assistant for marketing analytics"
space: marketing-analytics
llm: gpt_5
memory_mode: user
```
# Rules
Source: https://docs.cube.dev/admin/ai/rules
Define instructions in your data model repository that guide the agent's behavior.
Rules are instructions that guide the agent's behavior — encoding business definitions, calculation methods, domain terminology, and analytical approaches the agent should follow when answering questions.
Rules are configured as code in your [data model repository](/admin/ai#agent-configuration), alongside your cubes and views. This enables version control, code review, and consistent behavior across environments.
## Rule types
| Type | When applied | Best for |
| ----------------- | -------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- |
| `always` | Injected into every agent interaction. | Fundamental business definitions, default calculations, domain terms. |
| `agent_requested` | Conditionally applied when the agent determines the rule is relevant to the current request. | Scenario-specific guidance, specialized analysis methods. |
For `agent_requested` rules, the `description` field is what the agent matches against the user's request to decide whether the rule is relevant — write it as a short summary of when this rule applies. `always` rules don't need a `description` since they're injected into every interaction regardless.
## Defining rules
Rules are defined as Markdown files under `agents/rules/`. Each rule lives in its own file: the YAML frontmatter holds metadata, and the Markdown body is the rule prompt.
```markdown theme={"dark"}
---
type: always
---
Always use fiscal year starting April 1st when analyzing dates.
Q1 is April–June, Q2 is July–September, Q3 is October–December, Q4 is January–March.
```
For `agent_requested` rules, add a `description` so the agent can decide when the rule applies:
```markdown theme={"dark"}
---
description: "Apply when the user asks about cart abandonment"
type: agent_requested
---
For cart abandonment analysis, segment by device type and traffic source.
```
Files placed under a `rules/` directory are treated as rules automatically — no `kind` property is required. The `name` is inferred from the file name (e.g., `fiscal-year.md` → `fiscal-year`).
### Frontmatter properties
| Property | Type | Required | Description |
| ------------- | ------ | :------: | --------------------------------------------------------------------------------------------------------------------------------- |
| `name` | string | No | Unique identifier. Inferred from the file name if omitted. |
| `description` | string | No | Required for `agent_requested` rules — used by the agent to decide when the rule applies. Optional and unused for `always` rules. |
| `type` | string | Yes | Either `always` or `agent_requested`. |
| `prompt` | string | No | Rule prompt. Falls back to the Markdown body if omitted. |
### Inlining rules in YAML
You can also inline rules directly in `agents/config.yml` under a `rules` key:
```yaml theme={"dark"}
# agents/config.yml
rules:
- name: fiscal-year
prompt: "Always use fiscal year starting April 1st when analyzing dates."
type: always
- name: efficiency-analysis
description: "Apply when the user asks about sales efficiency"
prompt: "When analyzing sales efficiency, calculate as deal size divided by sales cycle length."
type: agent_requested
```
Inline rules accept the same properties as Markdown rules — `name`, `description`, `prompt` (required), and `type` (required).
Rules inlined at the root of `agents/config.yml` are attached to the implicit `auto` space and applied to the default agent in a [single-agent setup](/admin/ai). In a [multi-agent setup](/admin/ai/multi-agent), attach rules to a specific space by inlining them under that space's `rules` key (or by placing Markdown files under `agents/rules//`).
## Writing effective rules
Good rules are specific, actionable, and encode context the agent wouldn't otherwise know about your business.
**Do:**
* "Customer churn rate is customers lost ÷ total customers at the start of the period."
* "When analyzing quarterly performance, always compare against the same quarter of the previous year."
* "Our fiscal year starts in October."
**Don't:**
* "Be helpful." (too vague)
* "Always be accurate." (redundant)
* "Consider all factors." (too broad)
### Domain examples
**E-commerce:**
```markdown theme={"dark"}
---
type: always
---
Customer lifetime value equals average order value × purchase frequency × customer lifespan.
```
```markdown theme={"dark"}
---
description: "Apply when the user asks about cart abandonment"
type: agent_requested
---
For cart abandonment analysis, segment by device type and traffic source.
```
**SaaS:**
```markdown theme={"dark"}
---
type: always
---
MRR growth rate excludes one-time charges and setup fees.
```
```markdown theme={"dark"}
---
description: "Apply when the user asks about churn analysis"
type: agent_requested
---
When analyzing churn, distinguish between voluntary and involuntary churn.
```
## Resolving conflicts
When multiple rules could apply, follow these guidelines:
1. **Review existing rules** before adding new ones to avoid contradictions.
2. **Use specific triggers** in `agent_requested` rules so the agent knows when each rule applies.
3. **Prefer specificity over breadth** — narrowly scoped rules override broader defaults more cleanly.
4. **Test rule combinations** with sample queries before relying on them.
If two `always` rules directly contradict each other, the agent will surface the conflict in its response. Resolve such conflicts by editing the rules in your repository and redeploying.
# Skills
Source: https://docs.cube.dev/admin/ai/skills
Package reusable, named agent workflows in your data model repository that users can run on demand from chat.
Skills are reusable, named instruction packages for the agent — saved workflows a user
can run on demand wherever they work with the agent, including
[Analytics Chat](/docs/explore-analyze/analytics-chat), Workbooks, dashboards, and the
IDE. Instead of re-typing the same multi-step request ("produce a weekly revenue report,
broken down by region, with week-over-week trends…"), a user picks a skill and the agent
follows the workflow you've defined.
Skills are configured as code in your [data model repository](/admin/ai#agent-configuration),
alongside your cubes and views, so they're versioned, reviewed, and deployed like the rest
of your Cube project.
A skill guides what the agent does using the same data access the agent already has. The
first version of skills is instructions-only — `title`, `description`, and instructions.
There is no per-skill data scoping or external actions yet.
## Defining skills
Skills are defined as Markdown files under `agents/skills/`. Each skill lives in its own
file: the YAML frontmatter holds metadata, and the Markdown body is the instructions the
agent follows.
```markdown theme={"dark"}
---
title: "Weekly revenue report"
description: "Use when the user asks for a weekly revenue summary, a weekly revenue report, or week-over-week revenue trends."
---
Produce a weekly revenue report:
1. Report total revenue for the requested week, alongside the prior week and the
week-over-week percentage change.
2. Break revenue down by region, sorted from highest to lowest.
3. Highlight any region whose revenue changed by more than 10% week over week.
4. If the user names a region or a time range, scope the report accordingly.
```
Files placed under a `skills/` directory are treated as skills automatically — no `kind`
property is required. The `name` is inferred from the file name (e.g.,
`weekly-revenue-report.md` → `weekly-revenue-report`). Nested folders are allowed for
organization but do not namespace the skill — skill names must be unique across the
entire `skills/` directory.
### Frontmatter properties
| Property | Type | Required | Description |
| ------------- | ------ | :------: | --------------------------------------------------------------------------- |
| `title` | string | Yes | User-facing label shown on the skill button and in the `/` menu. |
| `description` | string | Yes | What the agent matches free-text requests against to auto-select the skill. |
| `name` | string | No | Unique identifier. Inferred from the file name if omitted. |
The Markdown body is the skill's instructions.
### Inlining skills in YAML
You can also inline skills directly in `agents/config.yml` under a `skills` key:
```yaml theme={"dark"}
# agents/config.yml
skills:
- name: weekly-revenue-report
title: "Weekly revenue report"
description: "Use when the user asks for a weekly revenue summary or week-over-week trends."
instructions: |
Produce a weekly revenue report:
1. Total revenue for the requested week, with the week-over-week change.
2. A breakdown by region, sorted by revenue descending.
3. Highlight any region that moved more than 10% week over week.
```
Inline skills use the same `name`, `title` (required), and `description` (required) as
Markdown skills. Since there is no Markdown body in YAML, provide the instructions inline
via the `instructions` key.
Skills inlined at the root of `agents/config.yml` are attached to the implicit `auto`
space and applied to the default agent in a [single-agent setup](/admin/ai). In a
[multi-agent setup](/admin/ai/multi-agent), attach skills to a specific space by inlining
them under that space's `skills` key (or by placing Markdown files under
`agents/skills//`).
## How skills are surfaced
Skills use **progressive disclosure**. The agent is given a compact catalog of every
available skill — just the `name`, `title`, and `description` — so it knows what each
skill is for without carrying the full instructions. When a skill is run, the agent loads
its complete instructions on demand.
* The chat UI only ever receives a skill's `name`, `title`, and `description`. The full
instruction body stays server-side and is never sent to the browser.
* Skills are surfaced through a viewer-accessible deployment endpoint, so they work for
any configured agent as well as the default `auto` agent.
* Because the agent matches against `description`, a well-written description makes
automatic matching more reliable.
For how skills appear and run in chat — buttons, the `/` menu, and automatic matching —
see [Agent skills](/docs/explore-analyze/skills).
## Deploying and testing
Skills ship through the normal Cube development flow: author the skill on a development
branch, commit, and merge to your production branch. Because skills live on the branch, a
skill on a development branch is testable before it reaches production — open chat against
the dev branch and run the skill to confirm it behaves as intended.
## Permissions
Authoring skills requires data-model edit access — the same access needed to define
[rules](/admin/ai/rules) and [certified queries](/admin/ai/certified-queries). In Cube
Cloud's built-in roles, that means the **Admin** or **Developer**
[roles](/admin/users-and-permissions/roles-and-permissions), or a custom role with
semantic-model edit access.
Running skills is available to anyone with chat access, including **Explorer** and
**Viewer** roles. In-product tips that promote authoring skills, rules, and certified
queries are role-aware and shown only to developers.
## Writing effective skills
* **Phrase `description` as a "use when…" statement.** This is what the agent matches
against incoming requests, so describe the situations the skill applies to.
* **Write instructions as an explicit workflow.** Number the steps and state the output
you expect, the same way you'd brief an analyst.
* **Account for added context.** Users can add specifics after selecting a skill (e.g.,
`/weekly-revenue-report for EMEA, last 6 weeks`), so instruct the agent to honor a
named region or time range when provided.
* **Keep each skill focused.** One skill should cover one well-defined workflow; create
separate skills for distinct tasks.
# Querying concurrency
Source: https://docs.cube.dev/admin/connect-to-data/concurrency
All queries to data APIs are processed asynchronously via a query queue. It allows to optimize the load and increase querying performance.
## Query queue
The query queue allows to deduplicate queries to API instances and insulate upstream
data sources from query spikes. It also allows to execute queries to data sources
concurrently for increased performance.
By default, Cube uses a *single* query queue for queries from all API instances and
the refresh worker to all configured data sources.
You can read more about the query queue in the [this blog post](https://cube.dev/blog/how-you-win-by-using-cube-store-part-1#query-queue-in-cube).
### Multiple query queues
You can use the [`context_to_orchestrator_id`][ref-context-to-orchestrator-id]
configuration option to route queries to multiple queues based on the security
context.
If you're configuring multiple connections to data sources via the [`driver_factory`
configuration option][ref-driver-factory], you **must** also configure
`context_to_orchestrator_id` to ensure that queries are routed to correct queues.
## Data sources
Cube supports various kinds of [data sources][ref-data-sources], ranging from cloud
data warehouses to embedded databases. Each data source scales differently,
therefore Cube provides sound defaults for each kind of data source out-of-the-box.
### Data source concurrency
By default, Cube uses the following concurrency settings for data sources:
| Data source | Default concurrency |
| ------------------------------- | ---------------------------------------------------------------------------- |
| [Amazon Athena][ref-athena] | 10 |
| [Amazon Redshift][ref-redshift] | 5 |
| [Apache Pinot][ref-pinot] | 10 |
| [ClickHouse][ref-clickhouse] | 10 |
| [Databricks][ref-databricks] | 10 |
| [Firebolt][ref-firebolt] | 10 |
| [Google BigQuery][ref-bigquery] | 10 |
| [Snowflake][ref-snowflake] | 8 |
| All other data sources | 5 or [less, if specified in the driver][link-github-data-source-concurrency] |
You can use the [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) environment variable to adjust the maximum
number of concurrent queries to a data source. It's recommended to use the default
configuration unless you're sure that your data source can handle more concurrent
queries.
### Connection pooling
For data sources that support connection pooling, the maximum number of concurrent
connections to the database can also be set by using the [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool)
environment variable. If changing this from the default, you must ensure that the
new value is greater than the number of concurrent connections used by Cube's query
queues and the refresh worker.
## Refresh worker
By default, the refresh worker uses the same concurrency settings as API instances.
However, you can override this behvaior in the refresh worker
[configuration][ref-preagg-refresh].
[ref-data-apis]: /reference
[ref-data-sources]: /admin/connect-to-data/data-sources
[ref-context-to-orchestrator-id]: /reference/configuration/config#context_to_orchestrator_id
[ref-driver-factory]: /reference/configuration/config#driver_factory
[ref-preagg-refresh]: /docs/pre-aggregations/refreshing-pre-aggregations#configuration
[ref-athena]: /admin/connect-to-data/data-sources/aws-athena
[ref-clickhouse]: /admin/connect-to-data/data-sources/clickhouse
[ref-databricks]: /admin/connect-to-data/data-sources/databricks-jdbc
[ref-firebolt]: /admin/connect-to-data/data-sources/firebolt
[ref-pinot]: /admin/connect-to-data/data-sources/pinot
[ref-redshift]: /admin/connect-to-data/data-sources/aws-redshift
[ref-snowflake]: /admin/connect-to-data/data-sources/snowflake
[ref-bigquery]: /admin/connect-to-data/data-sources/google-bigquery
[link-github-data-source-concurrency]: https://github.com/search?q=repo%3Acube-js%2Fcube+getDefaultConcurrency+path%3Apackages%2Fcubejs-&type=code
# Amazon Athena
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/aws-athena
Give Cube IAM access to Athena, an S3 query-results path, region and catalog settings, plus optional assumed-role credentials.
## Prerequisites
* [A set of IAM credentials][aws-docs-athena-access] which allow access to [AWS
Athena][aws-athena]
* [The AWS region][aws-docs-regions]
* [The S3 bucket][aws-s3] on AWS to [store query results][aws-docs-athena-query]
**Read-only access to your source data is sufficient.** Cube reads it with
`athena:*QueryExecution`/`GetQueryResults`/`GetWorkGroup`, `glue:Get*` metadata,
and `s3:GetObject`/`ListBucket`. Athena also **always** writes query results to
[`CUBEJS_AWS_S3_OUTPUT_LOCATION`](#environment-variables) — even for `SELECT`s —
so that location needs `s3:PutObject` (plus the multipart actions), but this is
Athena's own staging bucket, not your source data. Writes to source data are
only needed for [pre-aggregations](#pre-aggregation-build-strategies) built
inside Athena. For a ready-to-use IAM policy, see the
[AWS OIDC guide][ref-oidc-aws-athena].
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
#### Static Credentials
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=athena
CUBEJS_AWS_KEY=AKIA************
CUBEJS_AWS_SECRET=****************************************
CUBEJS_AWS_REGION=us-east-1
CUBEJS_AWS_S3_OUTPUT_LOCATION=s3://my-athena-output-bucket
CUBEJS_AWS_ATHENA_WORKGROUP=primary
CUBEJS_DB_NAME=my_non_default_athena_database
CUBEJS_AWS_ATHENA_CATALOG=AwsDataCatalog
```
#### IAM Role Assumption
For enhanced security, you can configure Cube to assume an IAM role to access Athena:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=athena
CUBEJS_AWS_ATHENA_ASSUME_ROLE_ARN=arn:aws:iam::123456789012:role/AthenaAccessRole
CUBEJS_AWS_REGION=us-east-1
CUBEJS_AWS_S3_OUTPUT_LOCATION=s3://my-athena-output-bucket
CUBEJS_AWS_ATHENA_WORKGROUP=primary
# Optional: if the role requires an external ID
CUBEJS_AWS_ATHENA_ASSUME_ROLE_EXTERNAL_ID=unique-external-id
```
When using role assumption:
* If running in AWS (EC2, ECS, EKS with IRSA), the driver will use the instance's IAM role or service account to assume the target role
* You can also provide [`CUBEJS_AWS_KEY`](/reference/configuration/environment-variables#cubejs_aws_key) and [`CUBEJS_AWS_SECRET`](/reference/configuration/environment-variables#cubejs_aws_secret) as master credentials for the role assumption
### Cube Cloud
In some cases you'll need to allow connections from your Cube Cloud deployment
IP address to your database. You can copy the IP address from either the
Database Setup step in deployment creation, or from **Settings →
Configuration** in your deployment.
In Cube Cloud, select **AWS Athena** when creating a new deployment and fill in
the required fields:
#### OIDC workload identity
Instead of static credentials, Cube Cloud deployments can authenticate to
Athena with [OIDC workload identity][ref-oidc-aws-athena]: an IAM role in
your account trusts Cube's OIDC issuer, and the driver assumes it through
the AWS SDK's default credential chain — no access keys to provision or
rotate.
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=athena
AWS_ROLE_ARN=arn:aws:iam::123456789012:role/cube-deployment-acme
CUBEJS_AWS_REGION=us-east-1
CUBEJS_AWS_S3_OUTPUT_LOCATION=s3://my-athena-output-bucket
```
See the [AWS OIDC guide][ref-oidc-aws-athena] for the IAM role, trust
policy, and permissions setup.
Cube Cloud also supports connecting to data sources within private VPCs
if [single-tenant infrastructure][ref-dedicated-infra] is used. Check out the
[VPC connectivity guide][ref-cloud-conf-vpc] for details.
[ref-dedicated-infra]: /admin/deployment/infrastructure#dedicated-infrastructure
[ref-cloud-conf-vpc]: /admin/deployment/dedicated
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| --------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | :-----------: |
| [`CUBEJS_AWS_KEY`](/reference/configuration/environment-variables#cubejs_aws_key) | The AWS Access Key ID to use for database connections | A valid AWS Access Key ID | ❌1 |
| [`CUBEJS_AWS_SECRET`](/reference/configuration/environment-variables#cubejs_aws_secret) | The AWS Secret Access Key to use for database connections | A valid AWS Secret Access Key | ❌1 |
| [`CUBEJS_AWS_REGION`](/reference/configuration/environment-variables#cubejs_aws_region) | The AWS region of the Cube deployment | [A valid AWS region][aws-docs-regions] | ✅ |
| `CUBEJS_AWS_S3_OUTPUT_LOCATION` | The S3 path to store query results made by the Cube deployment | A valid S3 path | ❌ |
| [`CUBEJS_AWS_ATHENA_WORKGROUP`](/reference/configuration/environment-variables#cubejs_aws_athena_workgroup) | The name of the workgroup in which the query is being started | [A valid Athena Workgroup][aws-athena-workgroup] | ❌ |
| [`CUBEJS_AWS_ATHENA_CATALOG`](/reference/configuration/environment-variables#cubejs_aws_athena_catalog) | The name of the catalog to use by default | [A valid Athena Catalog name][awsdatacatalog] | ❌ |
| [`CUBEJS_AWS_ATHENA_ASSUME_ROLE_ARN`](/reference/configuration/environment-variables#cubejs_aws_athena_assume_role_arn) | The ARN of the IAM role to assume for Athena access | A valid IAM role ARN | ❌ |
| [`CUBEJS_AWS_ATHENA_ASSUME_ROLE_EXTERNAL_ID`](/reference/configuration/environment-variables#cubejs_aws_athena_assume_role_external_id) | The external ID to use when assuming the IAM role (if required by the role's trust policy) | A string | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to use by default | A valid Athena Database name | ❌ |
| [`CUBEJS_DB_SCHEMA`](/reference/configuration/environment-variables#cubejs_db_schema) | The name of the schema to use as `information_schema` filter. Reduces count of tables loaded during schema generation. | A valid schema name | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
1 Either provide [`CUBEJS_AWS_KEY`](/reference/configuration/environment-variables#cubejs_aws_key) and [`CUBEJS_AWS_SECRET`](/reference/configuration/environment-variables#cubejs_aws_secret) for static credentials, or use [`CUBEJS_AWS_ATHENA_ASSUME_ROLE_ARN`](/reference/configuration/environment-variables#cubejs_aws_athena_assume_role_arn) for role-based authentication. When using role assumption without static credentials, the driver will use the AWS SDK's default credential chain (IAM instance profile, EKS IRSA, or [OIDC workload identity][ref-oidc-aws-athena] in Cube Cloud).
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
be used in pre-aggregations when using AWS Athena as a source database. To learn
more about AWS Athena's support for approximate aggregate functions, [click
here][aws-athena-docs-approx-agg-fns].
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ❌ | ✅ |
| Export Bucket | ❌ | ❌ |
By default, AWS Athena uses a [batching][self-preaggs-batching] strategy to
build pre-aggregations.
### Batching
No extra configuration is required to configure batching for AWS Athena.
Batching builds pre-aggregations inside Athena via `CREATE TABLE AS SELECT`, the
only case needing more than read-only access: `glue:CreateTable` on the
pre-aggregations schema plus `s3:PutObject` on its S3 location.
### Export Bucket
AWS Athena **only** supports using AWS S3 for export buckets.
#### AWS S3
For [improved pre-aggregation performance with large
datasets][ref-caching-large-preaggs], enable export bucket functionality by
configuring Cube with the following environment variables:
Ensure the AWS credentials are correctly configured in IAM to allow reads and
writes to the export bucket in S3.
The export bucket strategy uses Athena's `UNLOAD` to write directly to the
bucket — no temporary tables, no Glue Data Catalog writes. The only extra
permission beyond read-only querying is `s3:PutObject` on the export bucket.
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my.bucket.on.s3
CUBEJS_DB_EXPORT_BUCKET_AWS_KEY=
CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET=
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
## SSL
Cube does not require any additional configuration to enable SSL as AWS Athena
connections are made over HTTPS.
[aws-athena]: https://aws.amazon.com/athena
[aws-athena-workgroup]: https://docs.aws.amazon.com/athena/latest/ug/workgroups-benefits.html
[awsdatacatalog]: https://docs.aws.amazon.com/athena/latest/ug/understanding-tables-databases-and-the-data-catalog.html
[aws-s3]: https://aws.amazon.com/s3/
[aws-docs-athena-access]: https://docs.aws.amazon.com/athena/latest/ug/security-iam-athena.html
[aws-docs-athena-query]: https://docs.aws.amazon.com/athena/latest/ug/querying.html
[aws-docs-regions]: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-regions-availability-zones.html#concepts-available-regions
[aws-athena-docs-approx-agg-fns]: https://prestodb.io/docs/current/functions/aggregate.html#approximate-aggregate-functions
[ref-caching-large-preaggs]: /docs/pre-aggregations/using-pre-aggregations#export-bucket
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-oidc-aws-athena]: /admin/deployment/oidc/aws#athena
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-batching]: #batching
# AWS Redshift
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/aws-redshift
Authenticate Cube to Amazon Redshift with database passwords or IAM roles and ensure network reachability from your deployment.
## Prerequisites
* The [hostname][aws-redshift-docs-connection-string] for the [AWS
Redshift][aws-redshift] cluster
* The [username/password][aws-redshift-docs-users] for the [AWS
Redshift][aws-redshift] cluster **or** IAM credentials with
`redshift:GetClusterCredentialsWithIAM` and `redshift:DescribeClusters`
permissions
* The name of the database to use within the [AWS Redshift][aws-redshift]
cluster
If the cluster is configured within a [VPC][aws-vpc], then Cube **must** have a
network route to the cluster.
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
#### Password Authentication
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=redshift
CUBEJS_DB_HOST=my-redshift-cluster.cfbs3dkw1io8.eu-west-1.redshift.amazonaws.com
CUBEJS_DB_NAME=my_redshift_database
CUBEJS_DB_USER=
CUBEJS_DB_PASS=
```
#### IAM Authentication
For enhanced security, you can configure Cube to use IAM authentication
instead of username and password. When running in AWS (EC2, ECS, EKS with
IRSA), the driver can use the instance's IAM role to obtain temporary
database credentials automatically.
Omit [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) and [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) to enable IAM authentication:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=redshift
CUBEJS_DB_HOST=my-redshift-cluster.xxx.eu-west-1.redshift.amazonaws.com
CUBEJS_DB_NAME=my_redshift_database
CUBEJS_DB_SSL=true
CUBEJS_DB_REDSHIFT_AWS_REGION=eu-west-1
CUBEJS_DB_REDSHIFT_CLUSTER_IDENTIFIER=my-redshift-cluster
```
The driver uses the AWS SDK's default credential chain (IAM instance profile,
EKS IRSA, etc.) to obtain temporary database credentials via the
`redshift:GetClusterCredentialsWithIAM` API.
#### IAM Role Assumption
For cross-account access or enhanced security, you can configure Cube to assume
an IAM role:
```dotenv theme={"dark"}
CUBEJS_DB_REDSHIFT_AWS_REGION=eu-west-1
CUBEJS_DB_REDSHIFT_CLUSTER_IDENTIFIER=my-redshift-cluster
CUBEJS_DB_REDSHIFT_ASSUME_ROLE_ARN=arn:aws:iam::123456789012:role/RedshiftAccessRole
CUBEJS_DB_REDSHIFT_ASSUME_ROLE_EXTERNAL_ID=unique-external-id
```
### Cube Cloud
In some cases you'll need to allow connections from your Cube Cloud deployment
IP address to your database. You can copy the IP address from either the
Database Setup step in deployment creation, or from **Settings →
Configuration** in your deployment.
The following fields are required when creating an AWS Redshift connection:
#### OIDC workload identity
Instead of a database password, Cube Cloud deployments can authenticate to
Redshift with [OIDC workload identity][ref-oidc-aws-redshift]: an IAM role
in your account trusts Cube's OIDC issuer, and the driver uses it to obtain
temporary database credentials via [IAM authentication](#iam-authentication).
Set `AWS_ROLE_ARN` alongside the IAM authentication variables and omit
`CUBEJS_DB_USER` / `CUBEJS_DB_PASS`:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=redshift
AWS_ROLE_ARN=arn:aws:iam::123456789012:role/cube-deployment-acme
CUBEJS_DB_HOST=my-redshift-cluster.xxx.eu-west-1.redshift.amazonaws.com
CUBEJS_DB_NAME=my_redshift_database
CUBEJS_DB_SSL=true
CUBEJS_DB_REDSHIFT_AWS_REGION=eu-west-1
CUBEJS_DB_REDSHIFT_CLUSTER_IDENTIFIER=my-redshift-cluster
```
See the [AWS OIDC guide][ref-oidc-aws-redshift] for the IAM role, trust
policy, and permissions setup.
Cube Cloud also supports connecting to data sources within private VPCs
if [single-tenant infrastructure][ref-dedicated-infra] is used. Check out the
[VPC connectivity guide][ref-cloud-conf-vpc] for details.
[ref-dedicated-infra]: /admin/deployment/infrastructure#dedicated-infrastructure
[ref-cloud-conf-vpc]: /admin/deployment/dedicated
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | -------------------------------------- | :-----------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅1 |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅1 |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `16` | A valid number | ❌ |
| [`CUBEJS_DB_REDSHIFT_CLUSTER_IDENTIFIER`](/reference/configuration/environment-variables#cubejs_db_redshift_cluster_identifier) | The Redshift cluster identifier. Required for IAM authentication | A valid cluster identifier | ❌ |
| [`CUBEJS_DB_REDSHIFT_AWS_REGION`](/reference/configuration/environment-variables#cubejs_db_redshift_aws_region) | The AWS region of the Redshift cluster. Required for IAM authentication | [A valid AWS region][aws-docs-regions] | ❌ |
| [`CUBEJS_DB_REDSHIFT_ASSUME_ROLE_ARN`](/reference/configuration/environment-variables#cubejs_db_redshift_assume_role_arn) | The ARN of the IAM role to assume for cross-account access | A valid IAM role ARN | ❌ |
| [`CUBEJS_DB_REDSHIFT_ASSUME_ROLE_EXTERNAL_ID`](/reference/configuration/environment-variables#cubejs_db_redshift_assume_role_external_id) | The external ID for the assumed role's trust policy | A string | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_REDSHIFT_ARN`](/reference/configuration/environment-variables#cubejs_db_export_bucket_redshift_arn) | | | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
1 Required when using password-based authentication. When using IAM authentication, omit these and set [`CUBEJS_DB_REDSHIFT_CLUSTER_IDENTIFIER`](/reference/configuration/environment-variables#cubejs_db_redshift_cluster_identifier) and [`CUBEJS_DB_REDSHIFT_AWS_REGION`](/reference/configuration/environment-variables#cubejs_db_redshift_aws_region) instead. The driver uses the AWS SDK's default credential chain (IAM instance profile, EKS IRSA, or [OIDC workload identity][ref-oidc-aws-redshift] in Cube Cloud) to obtain temporary database credentials.
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
not be used in pre-aggregations when using AWS Redshift as a source database.
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ❌ | ✅ |
| Export Bucket | ❌ | ❌ |
By default, AWS Redshift uses [batching][self-preaggs-batching] to build
pre-aggregations.
### Batching
Cube requires the Redshift user to have ownership of a schema in Redshift to
support pre-aggregations. By default, the schema name is `prod_pre_aggregations`.
It can be set using the [`pre_aggregations_schema` configration
option][ref-conf-preaggs-schema].
No extra configuration is required to configure batching for AWS Redshift.
### Export bucket
AWS Redshift **only** supports using AWS S3 for export buckets.
#### AWS S3
For [improved pre-aggregation performance with large
datasets][ref-caching-large-preaggs], enable export bucket functionality by
configuring Cube with the following environment variables:
Ensure the AWS credentials are correctly configured in IAM to allow reads and
writes to the export bucket in S3.
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my.bucket.on.s3
CUBEJS_DB_EXPORT_BUCKET_AWS_KEY=
CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET=
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
## SSL
To enable SSL-encrypted connections between Cube and AWS Redshift, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
[aws-redshift-docs-connection-string]: https://docs.aws.amazon.com/redshift/latest/mgmt/configuring-connections.html#connecting-drivers
[aws-redshift-docs-users]: https://docs.aws.amazon.com/redshift/latest/dg/r_Users.html
[aws-redshift]: https://aws.amazon.com/redshift/
[aws-vpc]: https://aws.amazon.com/vpc/
[aws-docs-regions]: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-regions-availability-zones.html#concepts-available-regions
[ref-caching-large-preaggs]: /docs/pre-aggregations/using-pre-aggregations#export-bucket
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-oidc-aws-redshift]: /admin/deployment/oidc/aws#redshift
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-batching]: #batching
[ref-conf-preaggs-schema]: /reference/configuration/config#pre_aggregations_schema
# ClickHouse
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/clickhouse
Connect Cube to ClickHouse for low-latency analytics, including host credentials and driver-specific connection options.
[ClickHouse](https://clickhouse.com) is a fast and resource efficient
[open-source database](https://github.com/ClickHouse/ClickHouse) for real-time
applications and analytics.
## Prerequisites
* The hostname for the [ClickHouse][clickhouse] database server
* The [username/password][clickhouse-docs-users] for the
[ClickHouse][clickhouse] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=clickhouse
CUBEJS_DB_HOST=my.clickhouse.host
CUBEJS_DB_NAME=my_clickhouse_database
CUBEJS_DB_USER=clickhouse_user
CUBEJS_DB_PASS=**********
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_CLICKHOUSE_READONLY`](/reference/configuration/environment-variables#cubejs_db_clickhouse_readonly) | Whether the ClickHouse user has read-only access or not | `true`, `false` | ❌ |
| [`CUBEJS_DB_CLICKHOUSE_COMPRESSION`](/reference/configuration/environment-variables#cubejs_db_clickhouse_compression) | Whether the ClickHouse client has compression enabled or not | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `20` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
When using [pre-aggregations][ref-preaggs] with ClickHouse, you have to define
[indexes][ref-preaggs-indexes] in pre-aggregations. Otherwise, you might get
the following error: `ClickHouse doesn't support pre-aggregations without indexes`.
### `count_distinct_approx`
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
not be used in pre-aggregations when using ClickHouse as a source database.
### `rollup_join`
You can use [`rollup_join` pre-aggregations][ref-preaggs-rollup-join] to join
data from ClickHouse and other data sources inside Cube Store.
Alternatively, you can leverage ClickHouse support for [integration table
engines](https://clickhouse.com/docs/en/engines/table-engines#integration-engines)
to join data from ClickHouse and other data sources inside ClickHouse.
To do so, define table engines in ClickHouse and connect your ClickHouse as the
only data source to Cube.
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ✅ | ✅ |
| Export Bucket | ✅ | - |
By default, ClickHouse uses [batching][self-preaggs-batching] to build
pre-aggregations.
### Batching
No extra configuration is required to configure batching for ClickHouse.
### Export Bucket
Clickhouse driver **only** supports using AWS S3 for export buckets.
#### AWS S3
For [improved pre-aggregation performance with large
datasets][ref-caching-large-preaggs], enable export bucket functionality by
configuring Cube with the following environment variables:
Ensure the AWS credentials are correctly configured in IAM to allow reads and
writes to the export bucket in S3.
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my.bucket.on.s3
CUBEJS_DB_EXPORT_BUCKET_AWS_KEY=
CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET=
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
## SSL
To enable SSL-encrypted connections between Cube and ClickHouse, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
## Custom headers
The ClickHouse driver supports forwarding custom HTTP headers on every request to
the ClickHouse server. This is useful when requests pass through a proxy or gateway
that expects additional headers (e.g., for routing or tracing). See the [ClickHouse
JavaScript client configuration][clickhouse-docs-js-config] for more details.
Custom headers can't be configured via environment variables. Instead, use the
[`driver_factory`](/reference/configuration/config#driver_factory) configuration
option to pass a `headers` object to the driver:
```python title="Python" theme={"dark"}
from cube import config
@config('driver_factory')
def driver_factory(ctx: dict) -> dict:
return {
'type': 'clickhouse',
'headers': {
'X-Custom-Header': 'value',
'X-Routing-Group': 'analytics'
}
}
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
driverFactory: ({ dataSource }) => ({
type: "clickhouse",
headers: {
"X-Custom-Header": "value",
"X-Routing-Group": "analytics"
}
})
};
```
In multitenant deployments, you can use the [security
context](/embedding/authentication/security-context) to pass per-tenant headers, for example
to forward a user token from the API request down to ClickHouse:
```python title="Python" theme={"dark"}
from cube import config
@config('driver_factory')
def driver_factory(ctx: dict) -> dict:
security_context = ctx['securityContext']
return {
'type': 'clickhouse',
'headers': {
'X-Custom-User-Token': security_context['token']
}
}
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
driverFactory: ({ securityContext }) => ({
type: "clickhouse",
headers: {
"X-Custom-User-Token": securityContext.token
}
})
};
```
## Query attribution
Cube sets the ClickHouse [`query_id`](https://clickhouse.com/docs/operations/system-tables/query_log)
of every statement it runs to `-`, where `` is the
identifier shown for the query in Query History and `` is generated per statement, so
one Cube query matches as many rows as it ran statements. Use it to trace a Cube query to the
statements it produced in ClickHouse:
```sql theme={"dark"}
SELECT query_id, query, event_time
FROM system.query_log
WHERE query_id LIKE '%'
```
## Additional Configuration
You can connect to a ClickHouse database when your user's permissions are
[restricted][clickhouse-readonly] to read-only, by setting
[`CUBEJS_DB_CLICKHOUSE_READONLY`](/reference/configuration/environment-variables#cubejs_db_clickhouse_readonly) to `true`.
You can connect to a ClickHouse database with compression enabled, by setting
[`CUBEJS_DB_CLICKHOUSE_COMPRESSION`](/reference/configuration/environment-variables#cubejs_db_clickhouse_compression) to `true`.
[clickhouse]: https://clickhouse.tech/
[clickhouse-docs-users]: https://clickhouse.tech/docs/en/operations/settings/settings-users/
[clickhouse-docs-js-config]: https://clickhouse.com/docs/integrations/javascript#configuration
[clickhouse-readonly]: https://clickhouse.com/docs/en/operations/settings/permissions-for-queries#readonly
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
[ref-caching-large-preaggs]: /docs/pre-aggregations/using-pre-aggregations#export-bucket
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-batching]: #batching
[ref-preaggs]: /docs/pre-aggregations/using-pre-aggregations
[ref-preaggs-indexes]: /reference/data-modeling/pre-aggregations#indexes
[ref-preaggs-rollup-join]: /reference/data-modeling/pre-aggregations#rollup_join
# Databricks
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/databricks-jdbc
Databricks is a unified data intelligence platform.
[Databricks](https://www.databricks.com) is a unified data intelligence platform.
## Prerequisites
* [A JDK installation][gh-cubejs-jdbc-install]
* The [JDBC URL][databricks-docs-jdbc-url] for the [Databricks][databricks]
cluster
## Setup
### Environment Variables
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=databricks-jdbc
# CUBEJS_DB_NAME is optional
CUBEJS_DB_NAME=default
# You can find this inside the cluster's configuration
CUBEJS_DB_DATABRICKS_URL=jdbc:databricks://dbc-XXXXXXX-XXXX.cloud.databricks.com:443/default;transportMode=http;ssl=1;httpPath=sql/protocolv1/o/XXXXX/XXXXX;AuthMech=3;UID=token
# You can specify the personal access token separately from [`CUBEJS_DB_DATABRICKS_URL`](/reference/configuration/environment-variables#cubejs_db_databricks_url) by doing this:
CUBEJS_DB_DATABRICKS_TOKEN=XXXXX
# This accepts the Databricks usage policy and must be set to `true` to use the Databricks JDBC driver
CUBEJS_DB_DATABRICKS_ACCEPT_POLICY=true
```
### Docker
Create a `.env` file [as above](#environment-variables), then extend the
`cubejs/cube:jdk` Docker image tag to build a Cube image with the JDBC driver:
```dockerfile theme={"dark"}
FROM cubejs/cube:jdk
COPY . .
RUN npm install
```
You can then build and run the image using the following commands:
```bash theme={"dark"}
docker build -t cube-jdk .
docker run -it -p 4000:4000 --env-file=.env cube-jdk
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | --------------------- | :------: |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_DATABRICKS_URL`](/reference/configuration/environment-variables#cubejs_db_databricks_url) | The URL for a JDBC connection | A valid JDBC URL | ✅ |
| [`CUBEJS_DB_DATABRICKS_ACCEPT_POLICY`](/reference/configuration/environment-variables#cubejs_db_databricks_accept_policy) | Whether or not to accept the license terms for the Databricks JDBC driver | `true`, `false` | ✅ |
| [`CUBEJS_DB_DATABRICKS_OAUTH_CLIENT_ID`](/reference/configuration/environment-variables#cubejs_db_databricks_oauth_client_id) | The OAuth client ID for [service principal][ref-databricks-m2m-oauth] authentication | A valid client ID | ❌ |
| [`CUBEJS_DB_DATABRICKS_OAUTH_CLIENT_SECRET`](/reference/configuration/environment-variables#cubejs_db_databricks_oauth_client_secret) | The OAuth client secret for [service principal][ref-databricks-m2m-oauth] authentication | A valid client secret | ❌ |
| [`CUBEJS_DB_DATABRICKS_TOKEN`](/reference/configuration/environment-variables#cubejs_db_databricks_token) | The [personal access token][databricks-docs-pat] used to authenticate the Databricks connection | A valid token | ❌ |
| [`CUBEJS_DB_DATABRICKS_CATALOG`](/reference/configuration/environment-variables#cubejs_db_databricks_catalog) | The name of the [Databricks catalog][databricks-catalog] to connect to | A valid catalog name | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_MOUNT_DIR`](/reference/configuration/environment-variables#cubejs_db_export_bucket_mount_dir) | The path for the [Databricks DBFS mount][databricks-docs-dbfs] (Not needed if using Unity Catalog connection) | A valid mount path | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
be used in pre-aggregations when using Databricks as a source database. To learn
more about Databricks's support for approximate aggregate functions, [click
here][databricks-docs-approx-agg-fns].
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Simple | ✅ | ✅ |
| Export Bucket | ✅ | ❌ |
By default, Databricks JDBC uses a [simple][self-preaggs-simple] strategy to
build pre-aggregations.
### Simple
No extra configuration is required to configure simple pre-aggregation builds
for Databricks.
### Export Bucket
Databricks supports using both [AWS S3][aws-s3] and [Azure Blob
Storage][azure-bs] for export bucket functionality.
#### AWS S3
To use AWS S3 as an export bucket, first complete [the Databricks guide on
connecting to cloud object storage using Unity Catalog][databricks-docs-uc-s3].
Ensure the AWS credentials are correctly configured in IAM to allow reads and
writes to the export bucket in S3.
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=s3://my.bucket.on.s3
CUBEJS_DB_EXPORT_BUCKET_AWS_KEY=
CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET=
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
#### Google Cloud Storage
When using an export bucket, remember to assign the **Storage Object Admin**
role to your Google Cloud credentials ([`CUBEJS_DB_EXPORT_GCS_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_export_gcs_credentials)).
To use Google Cloud Storage as an export bucket, first complete [the Databricks guide on
connecting to cloud object storage using Unity Catalog][databricks-docs-uc-gcs].
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET=gs://databricks-export-bucket
CUBEJS_DB_EXPORT_BUCKET_TYPE=gcs
CUBEJS_DB_EXPORT_GCS_CREDENTIALS=
```
#### Azure Blob Storage
To use Azure Blob Storage as an export bucket, follow [the Databricks guide on
connecting to Azure Data Lake Storage Gen2 and Blob Storage][databricks-docs-azure].
[Retrieve the storage account access key][azure-bs-docs-get-key] from your Azure
account and use as follows:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=azure
CUBEJS_DB_EXPORT_BUCKET=wasbs://my-container@my-storage-account.blob.core.windows.net
CUBEJS_DB_EXPORT_BUCKET_AZURE_KEY=
```
Access key provides full access to the configuration and data,
to use a fine-grained control over access to storage resources, follow [the Databricks guide on authorize with Azure Active Directory][authorize-with-azure-active-directory].
[Create the service principal][azure-authentication-with-service-principal] and replace the access key as follows:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_AZURE_TENANT_ID=
CUBEJS_DB_EXPORT_BUCKET_AZURE_CLIENT_ID=
CUBEJS_DB_EXPORT_BUCKET_AZURE_CLIENT_SECRET=
```
## SSL/TLS
Cube does not require any additional configuration to enable SSL/TLS for
Databricks JDBC connections.
## Additional Configuration
### Cube Cloud
To accurately show partition sizes in the Cube Cloud APM, [an export
bucket][self-preaggs-export-bucket] **must be** configured.
[aws-s3]: https://aws.amazon.com/s3/
[azure-bs]: https://azure.microsoft.com/en-gb/services/storage/blobs/
[azure-bs-docs-get-key]: https://docs.microsoft.com/en-us/azure/storage/common/storage-account-keys-manage?toc=%2Fazure%2Fstorage%2Fblobs%2Ftoc.json&tabs=azure-portal#view-account-access-keys
[authorize-with-azure-active-directory]: https://learn.microsoft.com/en-us/rest/api/storageservices/authorize-with-azure-active-directory
[azure-authentication-with-service-principal]: https://learn.microsoft.com/en-us/azure/developer/java/sdk/identity-service-principal-auth
[databricks]: https://databricks.com/
[databricks-docs-dbfs]: https://docs.databricks.com/en/dbfs/mounts.html
[databricks-docs-azure]: https://docs.databricks.com/data/data-sources/azure/azure-storage.html
[databricks-docs-uc-s3]: https://docs.databricks.com/en/connect/unity-catalog/index.html
[databricks-docs-uc-gcs]: https://docs.databricks.com/gcp/en/connect/unity-catalog/cloud-storage.html
[databricks-docs-jdbc-url]: https://docs.databricks.com/integrations/bi/jdbc-odbc-bi.html#get-server-hostname-port-http-path-and-jdbc-url
[databricks-docs-pat]: https://docs.databricks.com/dev-tools/api/latest/authentication.html#token-management
[databricks-catalog]: https://docs.databricks.com/en/data-governance/unity-catalog/create-catalogs.html
[gh-cubejs-jdbc-install]: https://github.com/cube-js/cube/blob/master/packages/cubejs-jdbc-driver/README.md#java-installation
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[databricks-docs-approx-agg-fns]: https://docs.databricks.com/en/sql/language-manual/functions/approx_count_distinct.html
[self-preaggs-simple]: #simple
[self-preaggs-export-bucket]: #export-bucket
[ref-databricks-m2m-oauth]: https://docs.databricks.com/aws/en/dev-tools/auth/oauth-m2m
# Druid
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/druid
The driver for Druid is community-supported and is not maintained by Cube or the database vendor.
The driver for Druid is community-supported and is not maintained by Cube or the database vendor.
## Prerequisites
* The URL for the [Druid][druid] database
* The username/password for the [Druid][druid] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=druid
CUBEJS_DB_URL=https://my.druid.host:8082
CUBEJS_DB_USER=druid
CUBEJS_DB_PASS=**********
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------------------ | :------: |
| [`CUBEJS_DB_URL`](/reference/configuration/environment-variables#cubejs_db_url) | The URL for a database | A valid database URL for Druid | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## SSL
Cube does not require any additional configuration to enable SSL as Druid
connections are made over HTTPS.
[druid]: https://druid.apache.org/
# DuckDB
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/duckdb
Set up Cube with DuckDB for cloud object storage, local database files, or MotherDuck-backed hybrid query execution.
[DuckDB][duckdb] is an in-process SQL OLAP database management system, and has
support for querying data in CSV, JSON and Parquet formats from an AWS
S3-compatible blob storage. This means you can [query data stored in AWS S3,
Google Cloud Storage, or Cloudflare R2][duckdb-docs-s3-import].
You can also use the [`CUBEJS_DB_DUCKDB_DATABASE_PATH`](/reference/configuration/environment-variables#cubejs_db_duckdb_database_path) environment variable to
connect to a local DuckDB database.
Cube can also connect to [MotherDuck][motherduck], a cloud-based serverless
analytics platform built on DuckDB. When connected to MotherDuck, DuckDB uses
[hybrid execution][motherduck-docs-architecture] and routes queries to S3
through MotherDuck for better performance.
## Prerequisites
* A set of IAM credentials which allow access to the S3-compatible data source.
Credentials are only required for private S3 buckets.
* The region of the bucket
* The name of a bucket to query data from
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=duckdb
```
### Cube Cloud
In Cube Cloud, select **DuckDB** when creating a new deployment
and fill in the required fields:
If you are not using MotherDuck, leave the **MotherDuck Token**
field blank.
You can also explore how DuckDB works with Cube if you create a [demo
deployment][ref-demo-deployment] in Cube Cloud.
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- | :------: |
| [`CUBEJS_DB_DUCKDB_MEMORY_LIMIT`](/reference/configuration/environment-variables#cubejs_db_duckdb_memory_limit) | The maximum memory limit for DuckDB. Equivalent to `SET memory_limit=`. Default is 75% of available RAM | A valid memory limit | ❌ |
| [`CUBEJS_DB_DUCKDB_SCHEMA`](/reference/configuration/environment-variables#cubejs_db_duckdb_schema) | The [default search schema][link-duckdb-configuration-ref] | A valid schema name | ❌ |
| [`CUBEJS_DB_DUCKDB_MOTHERDUCK_TOKEN`](/reference/configuration/environment-variables#cubejs_db_duckdb_motherduck_token) | The service token to use for connections to MotherDuck | A valid [MotherDuck service token][motherduck-docs-svc-token] | ❌ |
| [`CUBEJS_DB_DUCKDB_DATABASE_PATH`](/reference/configuration/environment-variables#cubejs_db_duckdb_database_path) | The database filepath to use for connection to a local database. | A valid duckdb database file path | ❌ |
| `CUBEJS_DB_DUCKDB_S3_ACCESS_KEY_ID` | The Access Key ID to use for database connections | A valid Access Key ID | ❌ |
| `CUBEJS_DB_DUCKDB_S3_SECRET_ACCESS_KEY` | The Secret Access Key to use for database connections | A valid Secret Access Key | ❌ |
| `CUBEJS_DB_DUCKDB_S3_ENDPOINT` | The S3 endpoint | A valid [S3 endpoint][duckdb-docs-s3-import] | ❌ |
| `CUBEJS_DB_DUCKDB_S3_REGION` | The [region of the bucket][duckdb-docs-s3-import] | A valid AWS region | ❌ |
| `CUBEJS_DB_DUCKDB_S3_USE_SSL` | Use SSL for connection | A boolean | ❌ |
| `CUBEJS_DB_DUCKDB_S3_URL_STYLE` | To choose the S3 URL style(vhost or path) | `vhost` or `path` | ❌ |
| `CUBEJS_DB_DUCKDB_S3_SESSION_TOKEN` | The token for the S3 session | A valid Session Token | ❌ |
| [`CUBEJS_DB_DUCKDB_EXTENSIONS`](/reference/configuration/environment-variables#cubejs_db_duckdb_extensions) | A comma-separated list of DuckDB extensions to install and load | A comma-separated list of DuckDB extensions | ❌ |
| [`CUBEJS_DB_DUCKDB_COMMUNITY_EXTENSIONS`](/reference/configuration/environment-variables#cubejs_db_duckdb_community_extensions) | A comma-separated list of DuckDB community extensions to install and load | A comma-separated list of DuckDB community extensions | ❌ |
| `CUBEJS_DB_DUCKDB_S3_USE_CREDENTIAL_CHAIN` | A flag to use credentials chain for secrets for S3 connections | `true`, `false`. Defaults to `false` | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
be used in pre-aggregations when using DuckDB as a source database. To learn
more about DuckDB's support for approximate aggregate functions, [click
here][duckdb-docs-approx-agg-fns].
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ❌ | ✅ |
| Export Bucket | - | - |
By default, DuckDB uses a [batching][self-preaggs-batching] strategy to build
pre-aggregations.
### Batching
No extra configuration is required to configure batching for DuckDB.
### Export Bucket
DuckDB does not support export buckets.
## SSL
Cube does not require any additional configuration to enable SSL as DuckDB
connections are made over HTTPS.
[duckdb]: https://duckdb.org/
[duckdb-docs-approx-agg-fns]: https://duckdb.org/docs/sql/aggregates.html#approximate-aggregates
[duckdb-docs-s3-import]: https://duckdb.org/docs/guides/import/s3_import
[link-duckdb-configuration-ref]: https://duckdb.org/docs/sql/configuration.html#configuration-reference
[motherduck]: https://motherduck.com/
[motherduck-docs-architecture]: https://motherduck.com/docs/architecture-and-capabilities#hybrid-execution
[motherduck-docs-svc-token]: https://motherduck.com/docs/authenticating-to-motherduck/#authentication-using-a-service-token
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-batching]: #batching
[ref-demo-deployment]: /admin/deployment#demo-deployments
# Firebolt
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/firebolt
The driver for Firebolt is supported by its vendor. Please report any issues to their Help Center.
The driver for Firebolt is supported by its vendor. Please report any issues to
their [Help Center][firebolt-help].
## Prerequisites
* The id/secret (client id/client secret) for your [Firebolt][firebolt] [service account][firebolt-service-accounts]
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_NAME=firebolt_database
CUBEJS_DB_USER=aaaa-bbb-3244-wwssd
CUBEJS_DB_PASS=**********
CUBEJS_FIREBOLT_ACCOUNT=cube
CUBEJS_FIREBOLT_ENGINE_NAME=engine_name
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | :------: |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | A service account ID for accessing Firebolt programmatically | A valid service id | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | A service account secret for accessing Firebolt programmatically | A valid service secret | ✅ |
| [`CUBEJS_FIREBOLT_ACCOUNT`](/reference/configuration/environment-variables#cubejs_firebolt_account) | Account name | An account name | ✅ |
| [`CUBEJS_FIREBOLT_ENGINE_NAME`](/reference/configuration/environment-variables#cubejs_firebolt_engine_name) | Engine name to connect to | A valid engine name | ✅ |
| [`CUBEJS_FIREBOLT_API_ENDPOINT`](/reference/configuration/environment-variables#cubejs_firebolt_api_endpoint) | Firebolt API endpoint. Used for authentication | `api.dev.firebolt.io`, `api.staging.firebolt.io`, `api.app.firebolt.io` | - |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `20` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
[firebolt]: https://www.firebolt.io/
[firebolt-help]: https://help.firebolt.io/
[firebolt-service-accounts]: https://docs.firebolt.io/godocs/Guides/managing-your-organization/service-accounts.html
# Google BigQuery
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/google-bigquery
Grant a Google Cloud service account the right BigQuery IAM roles, then point Cube at your project, region, and key material.
## Prerequisites
In order to connect Google BigQuery to Cube, you need to provide service account
credentials. Cube requires the service account to have **BigQuery Data Viewer**
and **BigQuery Job User** roles enabled. If you plan to use pre-aggregations,
the account will need the **BigQuery Data Editor** role instead of **BigQuery Data Viewer**.
You can learn more about acquiring
Google BigQuery credentials [here][bq-docs-getting-started].
In Cube Cloud, you can authenticate with [OIDC workload identity
federation][ref-oidc-gcp-bigquery] instead of a key file — the same roles
apply to the impersonated service account.
* The [Google Cloud Project ID][google-cloud-docs-projects] for the
[BigQuery][bq] project
* A set of [Google Cloud service credentials][google-support-create-svc-account]
which [allow access][bq-docs-getting-started] to the [BigQuery][bq] project
* The [Google Cloud region][bq-docs-regional-locations] for the [BigQuery][bq]
project
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=bigquery
CUBEJS_DB_BQ_PROJECT_ID=my-bigquery-project-12345
CUBEJS_DB_BQ_KEY_FILE=/path/to/my/keyfile.json
```
You could also encode the key file using Base64 and set the result to
[`CUBEJS_DB_BQ_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_bq_credentials):
```dotenv theme={"dark"}
CUBEJS_DB_BQ_CREDENTIALS=$(cat /path/to/my/keyfile.json | base64)
```
### Cube Cloud
In some cases you'll need to allow connections from your Cube Cloud deployment
IP address to your database. You can copy the IP address from either the
Database Setup step in deployment creation, or from **Settings →
Configuration** in your deployment.
The following fields are required when creating a BigQuery connection:
#### OIDC workload identity federation
Instead of a service account key file, Cube Cloud deployments can
authenticate to BigQuery with [OIDC workload identity
federation][ref-oidc-gcp-bigquery]: a Workload Identity Federation provider
in your GCP project trusts Cube's OIDC issuer, and the driver authenticates
through the GCP default credential chain — no JSON key to provision or
rotate. Select **OIDC workload identity federation** in the connection
wizard, or set the equivalent environment variables:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=bigquery
CUBEJS_DB_BQ_PROJECT_ID=my-bigquery-project-12345
GCP_POOL_AUDIENCE=//iam.googleapis.com/projects/123456789/locations/global/workloadIdentityPools/cube-pool/providers/cube
GCP_SERVICE_ACCOUNT_EMAIL=cube-deployment@my-bigquery-project-12345.iam.gserviceaccount.com
```
`GCP_SERVICE_ACCOUNT_EMAIL` selects the service account Cube impersonates;
leave it unset to authenticate as the federated principal directly. See the
[GCP OIDC guide][ref-oidc-gcp-bigquery] for the Workload Identity Pool,
provider, and IAM setup.
Cube Cloud also supports connecting to data sources within private VPCs
if [single-tenant infrastructure][ref-dedicated-infra] is used. Check out the
[VPC connectivity guide][ref-cloud-conf-vpc] for details.
[ref-dedicated-infra]: /admin/deployment/infrastructure#dedicated-infrastructure
[ref-cloud-conf-vpc]: /admin/deployment/dedicated
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | :-----------: |
| [`CUBEJS_DB_BQ_PROJECT_ID`](/reference/configuration/environment-variables#cubejs_db_bq_project_id) | The Google BigQuery project ID to connect to | A valid Google BigQuery Project ID | ✅ |
| [`CUBEJS_DB_BQ_KEY_FILE`](/reference/configuration/environment-variables#cubejs_db_bq_key_file) | The path to a JSON key file for connecting to Google BigQuery | A valid Google BigQuery JSON key file | ✅1 |
| [`CUBEJS_DB_BQ_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_bq_credentials) | A Base64 encoded JSON key file for connecting to Google BigQuery | A valid Google BigQuery JSON key file encoded as a Base64 string | ❌ |
| [`CUBEJS_DB_BQ_LOCATION`](/reference/configuration/environment-variables#cubejs_db_bq_location) | The Google BigQuery dataset location to connect to. Required if used with pre-aggregations outside of US. If not set then BQ driver will fail with `Dataset was not found in location US` error | [A valid Google BigQuery regional location][bq-docs-regional-locations] | ⚠️ |
| [`CUBEJS_DB_EXPORT_BUCKET`](/reference/configuration/environment-variables#cubejs_db_export_bucket) | The name of a bucket in cloud storage | A valid bucket name from cloud storage | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_TYPE`](/reference/configuration/environment-variables#cubejs_db_export_bucket_type) | The cloud provider where the bucket is hosted | `gcp` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `40` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
1 Not required when authenticating with [OIDC workload identity
federation][ref-oidc-gcp-bigquery] in Cube Cloud, or when providing the key
material via [`CUBEJS_DB_BQ_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_bq_credentials).
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
be used in pre-aggregations when using Google BigQuery as a source database. To
learn more about Google BigQuery's support for approximate aggregate functions,
[click here][bq-docs-approx-agg-fns].
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ❌ | ✅ |
| Export Bucket | ❌ | ❌ |
By default, Google BigQuery uses [batching][self-preaggs-batching] to build
pre-aggregations.
### Batching
No extra configuration is required to configure batching for Google BigQuery.
### Export bucket
BigQuery only supports using Google Cloud Storage for export buckets.
#### Google Cloud Storage
For [improved pre-aggregation performance with large
datasets][ref-caching-large-preaggs], enable export bucket functionality by
configuring Cube with the following environment variables:
When using an export bucket, remember to assign the **BigQuery Data Editor** and
**Storage Object Admin** role to your BigQuery service account.
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET=export_data_58148478376
CUBEJS_DB_EXPORT_BUCKET_TYPE=gcp
```
## SSL
Cube does not require any additional configuration to enable SSL as Google
BigQuery connections are made over HTTPS.
[bq]: https://cloud.google.com/bigquery
[bq-docs-getting-started]: https://cloud.google.com/docs/authentication/getting-started
[bq-docs-regional-locations]: https://cloud.google.com/bigquery/docs/locations#regional-locations
[bq-docs-approx-agg-fns]: https://cloud.google.com/bigquery/docs/reference/standard-sql/approximate_aggregate_functions
[google-cloud-docs-projects]: https://cloud.google.com/resource-manager/docs/creating-managing-projects#before_you_begin
[google-support-create-svc-account]: https://support.google.com/a/answer/7378726?hl=en
[ref-caching-large-preaggs]: /docs/pre-aggregations/using-pre-aggregations#export-bucket
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-oidc-gcp-bigquery]: /admin/deployment/oidc/gcp#bigquery
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-batching]: #batching
# Hive / SparkSQL
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/hive
The driver for Hive / SparkSQL is deprecated and community-supported, and is not maintained by Cube or the database vendor.
The driver for Hive / SparkSQL is deprecated and will be removed in a future
release. It is community-supported and is not maintained by Cube or the database
vendor.
## Prerequisites
* The hostname for the [Hive][hive] database server
* The username/password for the [Hive][hive] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=hive
CUBEJS_DB_HOST=my.hive.host
CUBEJS_DB_NAME=my_hive_database
CUBEJS_DB_USER=hive_user
CUBEJS_DB_PASS=**********
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_HIVE_TYPE`](/reference/configuration/environment-variables#cubejs_db_hive_type) | | | ❌ |
| [`CUBEJS_DB_HIVE_VER`](/reference/configuration/environment-variables#cubejs_db_hive_ver) | | | ❌ |
| [`CUBEJS_DB_HIVE_THRIFT_VER`](/reference/configuration/environment-variables#cubejs_db_hive_thrift_ver) | | | ❌ |
| [`CUBEJS_DB_HIVE_CDH_VER`](/reference/configuration/environment-variables#cubejs_db_hive_cdh_ver) | | | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
[hive]: https://hive.apache.org/
# Data Sources
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/index
Choose and configure a data warehouse, query engine, or other data source to connect to Cube.
Cube connects directly to your data warehouse or database. Select your data
source below to get started with the setup process.
You can also connect [multiple data sources][ref-config-multi-data-src] at the same
time and adjust the [concurrency settings][ref-data-source-concurrency] for data
sources.
## Data warehouses
Connect to Amazon Redshift.
Connect to Google BigQuery.
Connect to Snowflake.
Connect to Databricks.
Connect to Microsoft Fabric.
Connect to ClickHouse.
Connect to SingleStore.
Connect to Apache Pinot.
Connect to Firebolt.
Connect to Vertica.
## Query engines
Connect to Amazon Athena.
Connect to DuckDB or MotherDuck.
Connect to Hive or SparkSQL.
Connect to Presto.
Connect to Trino.
## Transactional databases
Connect to Postgres.
Connect to Microsoft SQL Server.
Connect to MySQL.
Connect to Oracle.
Connect to SQLite.
## Time series & streaming
Connect to QuestDB.
Connect to ksqlDB.
Connect to Materialize.
Connect to RisingWave.
## Other data sources
Connect to MongoDB.
Connect to Apache Druid.
Query Parquet files via DuckDB.
Query CSV files via DuckDB.
Query JSON files via DuckDB.
### API endpoints
Cube is designed to work with data sources that allow querying them with SQL.
Cube is not designed to access data files directly or fetch data from REST,
or GraphQL, or any other API. To use Cube in that way, you either need to use
a supported data source (e.g., use [DuckDB][ref-duckdb] to query Parquet files
on Amazon S3) or create a [custom data source driver](#currently-unsupported-data-sources).
## Data source drivers
### Driver support
Most of the drivers for data sources are supported either directly by the Cube
team or by their vendors. The rest are community-supported and will be
highlighted as such in their respective pages.
You can find the [source code][link-github-packages] of the drivers that are part
of the Cube distribution in `cubejs-*-driver` folders on GitHub.
### Third-party drivers
The following drivers were contributed by the Cube community. They are not part
of the Cube distribution, however, they can still be used with Cube:
* [ArangoDB](https://www.npmjs.com/package/arangodb-cubejs-driver)
* [CosmosDB](https://www.npmjs.com/package/cosmosdb-cubejs-driver)
* [CrateDB](https://www.npmjs.com/package/cratedb-cubejs-driver)
* [Dremio](https://www.npmjs.com/package/mydremio-cubejs-driver)
* [Dremio ODBC](https://www.npmjs.com/package/dremio-odbc-cubejs-driver)
* [OpenDistro Elastic](https://www.npmjs.com/package/opendistro-cubejs-driver)
* [SAP Hana](https://www.npmjs.com/package/cubejs-hana-driver)
* [Trino](https://www.npmjs.com/package/trino-cubejs-driver)
* [Vertica](https://www.npmjs.com/package/@knowitall/vertica-driver)
You need to configure [`driver_factory`][ref-driver-factory] to use a third-party
driver.
### Currently unsupported data sources
If you'd like to connect to a data source which is not yet listed on this page,
please see the list of [requested drivers](https://github.com/cube-js/cube/issues/7076)
and [file an issue](https://github.com/cube-js/cube/issues) on GitHub.
You're more than welcome to contribute new drivers as well as new features and
patches to
[existing drivers](https://github.com/cube-js/cube/tree/master/packages). Please
check the
[contribution guidelines](https://github.com/cube-js/cube/blob/master/CONTRIBUTING.md#contributing-database-drivers)
and join the `#contributing-to-cube` channel in our
[Slack community](https://slack.cube.dev).
[ref-config-multi-data-src]: /admin/connect-to-data/multiple-data-sources
[ref-driver-factory]: /reference/configuration/config#driver_factory
[ref-duckdb]: /admin/connect-to-data/data-sources/duckdb
[link-github-packages]: https://github.com/cube-js/cube/tree/master/packages
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-sources
# ksqlDB
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/ksqldb
ksqlDB is a purpose-built database for stream processing applications, ingesting data from Apache Kafka.
[ksqlDB](https://ksqldb.io) is a purpose-built database for stream processing
applications, ingesting data from [Apache Kafka](https://kafka.apache.org).
Available on the [Enterprise plan](https://cube.dev/pricing).
[Contact us](https://cube.dev/contact) for details.
See how you can use ksqlDB and Cube Cloud to power real-time analytics in Power BI:
In this video, the SQL API is used to connect to Power BI.
Currently, it's recommended to use the DAX API.
## Prerequisites
* Hostname for the ksqlDB server
* Username and password (or an API key) to connect to ksqlDB server
### Confluent Cloud
If you are using [Confluent Cloud](https://www.confluent.io/confluent-cloud/),
you need to generate an API key and use the API key name as your username and
the API key secret as your password.
You can generate an API key by installing `confluent-cli` and running the
following commands in the command line:
```sh theme={"dark"}
brew install --cask confluent-cli
confluent login
confluent environment use
confluent ksql cluster list
confluent api-key create --resource
```
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=ksql
CUBEJS_DB_URL=https://xxxxxx-xxxxx.us-west4.gcp.confluent.cloud:443
CUBEJS_DB_USER=username
CUBEJS_DB_PASS=password
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| --------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | ------------------------- | :------: |
| [`CUBEJS_DB_URL`](/reference/configuration/environment-variables#cubejs_db_url) | The host URL for ksqlDB with port | A valid database host URL | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the ksqlDB. API key for Confluent Cloud. | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the ksqlDB. API secret for Confluent Cloud. | A valid database password | ✅ |
| [`CUBEJS_DB_KAFKA_HOST`](/reference/configuration/environment-variables#cubejs_db_kafka_host) | Kafka broker host(s) for [Kafka streams mode](#kafka-streams-mode). Multiple brokers can be comma-separated. | A valid Kafka broker URL | ❌ |
| [`CUBEJS_DB_KAFKA_USER`](/reference/configuration/environment-variables#cubejs_db_kafka_user) | Username for Kafka broker authentication (SASL PLAIN) | A valid Kafka username | ❌ |
| [`CUBEJS_DB_KAFKA_PASS`](/reference/configuration/environment-variables#cubejs_db_kafka_pass) | Password for Kafka broker authentication (SASL PLAIN) | A valid Kafka password | ❌ |
| [`CUBEJS_DB_KAFKA_USE_SSL`](/reference/configuration/environment-variables#cubejs_db_kafka_use_ssl) | If `true`, enables SASL\_SSL for the Kafka connection | `true`, `false` | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
## Pre-Aggregations Support
ksqlDB supports only
[streaming pre-aggregations][ref-streaming-pre-aggs].
## Kafka streams mode
By default, Cube connects to ksqlDB via its REST API. ksqlDB uses its REST
API both for metadata (discovering tables and streams) and for streaming
data into Cube Store during pre-aggregation builds.
In this default mode, Cube may create tables and streams in ksqlDB as part
of the pre-aggregation build process (e.g., `CREATE TABLE ... AS SELECT`
statements for non-read-only pre-aggregations).
When **Kafka streams mode** is enabled, Cube reads data directly from the
underlying Kafka topics instead of going through the ksqlDB REST API for
data streaming. ksqlDB is still used for metadata operations such as
discovering tables, streams, and their schemas, but Cube Store subscribes
to the backing Kafka topic directly.
In this mode, Cube does not create any tables or streams in ksqlDB. All
pre-aggregations use the read-only refresh path: Cube discovers the
existing ksqlDB objects and their backing Kafka topics, then streams data
directly from Kafka into Cube Store.
### When to use Kafka streams mode
Kafka streams mode is useful when:
* You want to prevent Cube from creating any objects in ksqlDB
* You need higher throughput for data ingestion by reading Kafka directly
* Your ksqlDB environment has restricted permissions that don't allow
creating tables or streams
* You prefer Cube Store to consume from Kafka topics without an
intermediary
### Enabling Kafka streams mode
Set the `CUBEJS_DB_KAFKA_HOST` environment variable to the address of your
Kafka broker(s). This activates Kafka streams mode automatically:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=ksql
CUBEJS_DB_URL=https://xxxxxx-xxxxx.us-west4.gcp.confluent.cloud:443
CUBEJS_DB_USER=ksql_username
CUBEJS_DB_PASS=ksql_password
CUBEJS_DB_KAFKA_HOST=pkc-xxxxx.us-west4.gcp.confluent.cloud:9092
CUBEJS_DB_KAFKA_USER=kafka_api_key
CUBEJS_DB_KAFKA_PASS=kafka_api_secret
CUBEJS_DB_KAFKA_USE_SSL=true
```
Multiple Kafka brokers can be specified as a comma-separated list:
```dotenv theme={"dark"}
CUBEJS_DB_KAFKA_HOST=broker1:9092,broker2:9092,broker3:9092
```
When using [Confluent Cloud](https://www.confluent.io/confluent-cloud/),
the Kafka credentials are separate from the ksqlDB credentials. Generate
an API key for the Kafka cluster (not the ksqlDB cluster) and use it as
`CUBEJS_DB_KAFKA_USER` and `CUBEJS_DB_KAFKA_PASS`.
### How it works
With Kafka streams mode enabled:
1. Cube uses the ksqlDB REST API to discover available tables and streams
and to retrieve their schemas via `DESCRIBE`.
2. For each table or stream, Cube resolves the backing Kafka topic name
from the ksqlDB metadata.
3. Instead of streaming data through ksqlDB, Cube Store connects directly
to the Kafka broker(s) and consumes from the resolved topic.
4. Pre-aggregation builds use the read-only refresh strategy. Cube does
not issue any `CREATE TABLE` or `CREATE STREAM` statements to ksqlDB.
### Data modeling
ksqlDB is typically used as an additional data source alongside a primary
data warehouse. To use Kafka streams mode, configure ksqlDB as a named
data source using [decorated environment variables][ref-decorated-env-vars]
and point your cubes to it with the
[`data_source`][ref-cube-data-source] property.
First, declare the data sources and configure the ksqlDB connection with
Kafka credentials:
```dotenv theme={"dark"}
CUBEJS_DATASOURCES=default,ksql
CUBEJS_DB_TYPE=postgres
CUBEJS_DB_HOST=my.postgres.host
CUBEJS_DB_NAME=my_database
CUBEJS_DB_USER=postgres_user
CUBEJS_DB_PASS=postgres_password
CUBEJS_DS_KSQL_DB_TYPE=ksql
CUBEJS_DS_KSQL_DB_URL=https://xxxxxx-xxxxx.us-west4.gcp.confluent.cloud:443
CUBEJS_DS_KSQL_DB_USER=ksql_api_key
CUBEJS_DS_KSQL_DB_PASS=ksql_api_secret
CUBEJS_DS_KSQL_DB_KAFKA_HOST=pkc-xxxxx.us-west4.gcp.confluent.cloud:9092
CUBEJS_DS_KSQL_DB_KAFKA_USER=kafka_api_key
CUBEJS_DS_KSQL_DB_KAFKA_PASS=kafka_api_secret
CUBEJS_DS_KSQL_DB_KAFKA_USE_SSL=true
```
Then, create cubes that reference your data. A common pattern is to
combine a **batch cube** (reading historical data from your warehouse)
with a **streaming cube** (reading real-time data from ksqlDB via Kafka)
using a [lambda pre-aggregation][ref-lambda-pre-aggs].
The batch cube queries the warehouse and builds daily partitions
incrementally. The streaming cube points at an existing ksqlDB stream
with `data_source: ksql` and uses a read-only streaming pre-aggregation
that consumes from the backing Kafka topic directly. The lambda
pre-aggregation in the batch cube merges both, serving historical data
from the warehouse rollup and real-time data from the streaming rollup:
```yaml title="YAML" theme={"dark"}
cubes:
- name: order_events
data_source: default
sql: >
SELECT
order_id,
user_id,
status,
amount,
created_at
FROM ecommerce.order_events
WHERE {FILTER_PARAMS.order_events.created_at.filter(
(from, to) =>
`created_at >= ${from} AND created_at < ${to}`
)}
measures:
- name: count
type: count
- name: total_amount
sql: amount
type: sum
- name: failed_count
sql: "CASE WHEN status = 'failed' THEN 1 ELSE 0 END"
type: sum
dimensions:
- name: order_id
sql: order_id
type: string
primary_key: true
- name: user_id
sql: user_id
type: string
- name: status
sql: status
type: string
- name: created_at
sql: created_at
type: time
pre_aggregations:
- name: lambda
type: rollup_lambda
rollups:
- order_events.batch
- order_events_stream.stream
- name: batch
type: rollup
measures:
- CUBE.count
- CUBE.total_amount
- CUBE.failed_count
dimensions:
- CUBE.order_id
- CUBE.user_id
- CUBE.status
time_dimension: CUBE.created_at
granularity: second
partition_granularity: day
build_range_start:
sql: SELECT NOW() - INTERVAL '90 days'
build_range_end:
sql: SELECT NOW()
refresh_key:
every: 8 hour
update_window: 1 day
incremental: true
indexes:
- name: user_status
columns:
- CUBE.user_id
- CUBE.status
- name: order_events_stream
data_source: ksql
sql: "SELECT * FROM ORDER_EVENTS_STREAM"
measures:
- name: count
type: count
- name: total_amount
sql: AMOUNT
type: sum
- name: failed_count
sql: "CASE WHEN STATUS = 'failed' THEN 1 ELSE 0 END"
type: sum
dimensions:
- name: order_id
sql: ORDER_ID
type: string
primary_key: true
- name: user_id
sql: USER_ID
type: string
- name: status
sql: STATUS
type: string
- name: created_at
sql: CREATED_AT
type: time
pre_aggregations:
- name: stream
type: rollup
read_only: true
measures:
- CUBE.count
- CUBE.total_amount
- CUBE.failed_count
dimensions:
- CUBE.order_id
- CUBE.user_id
- CUBE.status
unique_key_columns:
- order_id
- user_id
- status
- created_at_second
time_dimension: CUBE.created_at
granularity: second
partition_granularity: day
# ksqlDB does not support NOW(); use CURRENT_TIMESTAMP with INTERVAL arithmetic instead.
build_range_start:
sql: "SELECT date_trunc('day', CURRENT_TIMESTAMP - INTERVAL '5 hour')"
build_range_end:
sql: "SELECT CURRENT_TIMESTAMP + INTERVAL '15 minute'"
refresh_key:
every: 1 minute
update_window: 1 hour
incremental: true
indexes:
- name: user_status
columns:
- CUBE.user_id
- CUBE.status
- orders__created_at_second
stream_offset: latest
output_column_types:
- member: CUBE.order_id
type: text
- member: CUBE.user_id
type: text
- member: CUBE.status
type: text
- member: CUBE.created_at.second
type: timestamp
- member: CUBE.count
type: int
- member: CUBE.total_amount
type: decimal
- member: CUBE.failed_count
type: int
```
```javascript title="JavaScript" theme={"dark"}
cube("order_events", {
data_source: "default",
sql: `
SELECT
order_id,
user_id,
status,
amount,
created_at
FROM ecommerce.order_events
WHERE ${FILTER_PARAMS.order_events.created_at.filter(
(from, to) => `created_at >= ${from} AND created_at < ${to}`
)}
`,
measures: {
count: {
type: `count`,
},
total_amount: {
sql: `amount`,
type: `sum`,
},
failed_count: {
sql: `CASE WHEN status = 'failed' THEN 1 ELSE 0 END`,
type: `sum`,
},
},
dimensions: {
order_id: {
sql: `order_id`,
type: `string`,
primary_key: true,
},
user_id: {
sql: `user_id`,
type: `string`,
},
status: {
sql: `status`,
type: `string`,
},
created_at: {
sql: `created_at`,
type: `time`,
},
},
pre_aggregations: {
lambda: {
type: `rollup_lambda`,
rollups: [
order_events.batch,
order_events_stream.stream,
],
},
batch: {
type: `rollup`,
measures: [CUBE.count, CUBE.total_amount, CUBE.failed_count],
dimensions: [CUBE.order_id, CUBE.user_id, CUBE.status],
time_dimension: CUBE.created_at,
granularity: `second`,
partition_granularity: `day`,
build_range_start: {
sql: `SELECT NOW() - INTERVAL '90 days'`,
},
build_range_end: {
sql: `SELECT NOW()`,
},
refresh_key: {
every: `8 hour`,
update_window: `1 day`,
incremental: true,
},
indexes: {
user_status: {
columns: [CUBE.user_id, CUBE.status],
},
},
},
},
});
cube("order_events_stream", {
data_source: "ksql",
sql: `SELECT * FROM ORDER_EVENTS_STREAM`,
measures: {
count: {
type: `count`,
},
total_amount: {
sql: `AMOUNT`,
type: `sum`,
},
failed_count: {
sql: `CASE WHEN STATUS = 'failed' THEN 1 ELSE 0 END`,
type: `sum`,
},
},
dimensions: {
order_id: {
sql: `ORDER_ID`,
type: `string`,
primary_key: true,
},
user_id: {
sql: `USER_ID`,
type: `string`,
},
status: {
sql: `STATUS`,
type: `string`,
},
created_at: {
sql: `CREATED_AT`,
type: `time`,
},
},
pre_aggregations: {
stream: {
type: `rollup`,
read_only: true,
measures: [CUBE.count, CUBE.total_amount, CUBE.failed_count],
dimensions: [CUBE.order_id, CUBE.user_id, CUBE.status],
unique_key_columns: [
`order_id`,
`user_id`,
`status`,
`created_at_second`,
],
time_dimension: CUBE.created_at,
granularity: `second`,
partition_granularity: `day`,
// ksqlDB does not support NOW(); use CURRENT_TIMESTAMP with INTERVAL arithmetic instead.
build_range_start: {
sql: `SELECT date_trunc('day', CURRENT_TIMESTAMP - INTERVAL '5 hour')`,
},
build_range_end: {
sql: `SELECT CURRENT_TIMESTAMP + INTERVAL '15 minute'`,
},
refresh_key: {
every: `1 minute`,
update_window: `1 hour`,
incremental: true,
},
indexes: {
user_status: {
columns: [CUBE.user_id, CUBE.status, `orders__created_at_second`],
},
},
stream_offset: `latest`,
output_column_types: [
{ member: CUBE.order_id, type: `text` },
{ member: CUBE.user_id, type: `text` },
{ member: CUBE.status, type: `text` },
{ member: CUBE.created_at.second, type: `timestamp` },
{ member: CUBE.count, type: `int` },
{ member: CUBE.total_amount, type: `decimal` },
{ member: CUBE.failed_count, type: `int` },
],
},
},
});
```
Key properties for the streaming pre-aggregation:
* `read_only: true` — Cube will not create any objects in ksqlDB. The
data is consumed directly from the backing Kafka topic.
* `stream_offset` — controls where Cube Store starts consuming from in
the Kafka topic. Set to `"latest"` to only consume new messages
arriving after the pre-aggregation is created. Set to `"earliest"` to
replay the topic from the beginning. Defaults to `"latest"` if not
specified. On subsequent refreshes, Cube Store automatically resumes
from the last processed offset regardless of this setting.
* `unique_key_columns` — columns that uniquely identify a record, used
for deduplication (see [below](#unique-key-columns-and-deduplication)).
* `output_column_types` — declares the output column types for the Cube
Store table, required for Kafka streams mode (see
[below](#output-column-types)).
#### Primary key and ungrouped queries
For the streaming pre-aggregation to work in read-only mode, the
generated SQL must not contain a `GROUP BY` clause — Cube Store's stream
post-processing engine does not support aggregation.
Cube automatically omits the `GROUP BY` clause when the dimensions
included in the pre-aggregation contain a primary key. In that case, the
generated query becomes a simple `SELECT ... FROM ...` without grouping,
and measures are passed through as raw expressions rather than
aggregated. This is what makes the pre-aggregation eligible for the
read-only streaming path.
You must include **all** primary key columns of the cube in the
streaming pre-aggregation's `dimensions` list. If any primary key
dimension is missing, the query may not be recognized as ungrouped
and will fail to use the streaming path.
The `sql_table` or `sql` value should reference an existing ksqlDB stream
or table. Cube discovers its schema automatically. With Kafka streams
mode enabled, the streaming pre-aggregation reads the backing Kafka topic
directly — no objects are created in ksqlDB.
#### Topic name matching
In Kafka streams mode, Cube Store parses the `select_statement`
(generated from the cube's `sql` property) and matches the `FROM` table
name against the actual Kafka topic name. On managed platforms like
Confluent Cloud, the Kafka topic name often differs from the ksqlDB
stream or table name — for example, a ksqlDB stream called
`ORDER_EVENTS_STREAM` might be backed by a Kafka topic named
`pksqlc-abc123ORDER_EVENTS_STREAM`.
The cube's `sql` property must reference the **ksqlDB stream or table
name** (not the Kafka topic), because Cube uses ksqlDB `DESCRIBE` to
discover the schema and resolve the backing topic. However, Cube does
not currently rewrite the `FROM` clause in the generated
`select_statement` to use the resolved Kafka topic name. If the ksqlDB
object name differs from the Kafka topic name, Cube Store will fail
with:
> Topic table ORDER\_EVENTS\_STREAM is not found
**The ksqlDB stream (or table) name and the backing Kafka topic name
MUST be identical, including case.** Kafka streams mode will fail if
they differ in any way.
Take this into account when creating the stream — explicitly set the
topic name to match the stream name (and vice versa). For example:
```sql theme={"dark"}
CREATE STREAM ORDER_EVENTS_STREAM (...)
WITH (KAFKA_TOPIC='ORDER_EVENTS_STREAM', VALUE_FORMAT='JSON', ...);
```
The match is **case-sensitive**: `OrderEvents` and `ORDEREVENTS` are
treated as different names and will cause the build to fail with
`Topic table ... is not found`.
This is a known limitation of Kafka streams mode. It does not occur
when the ksqlDB object name and the Kafka topic name are the same,
which is the default behavior when ksqlDB creates a stream or table
with the default topic naming strategy.
### Unique key columns and deduplication
When `unique_key_columns` is set, Cube Store appends an internal
sequence column (`__seq`) to the table, populated from the Kafka
partition offset. The unique key columns together with `__seq` form the
sort key for all indexes on this table.
Entries in `unique_key_columns` are strings, not member references. Use
the dimension's own name (for example, `user_id`). To include the time
dimension as part of the unique key, use the form
`_` — for example, `created_at_second`
when `granularity` is `second`.
The naming convention for the time dimension inside `unique_key_columns`
is **not** the same as the column name used in
[`indexes`](/reference/data-modeling/pre-aggregations#indexes). Indexes
expect the fully-qualified alias from the generated `select_statement`
(`___`), while
`unique_key_columns` expects only `_`.
If the pre-aggregation defines `indexes`, every dimension referenced by
any index must also be listed in `unique_key_columns`. Cube Store uses
`unique_key_columns` (plus `__seq`) as the sort key for all indexes, so
an index column that is not part of the unique key cannot be sorted on
during compaction and the pre-aggregation build will fail.
Deduplication is not applied at ingestion time — all incoming records are
appended as they arrive. Instead, Cube Store deduplicates during
**reads** and **compaction**: rows are sorted by the unique key columns
and then by `__seq`, and only the **last row per unique key** (the one
with the highest sequence number) is kept. This means that if the same
key appears multiple times in the stream, the most recent version is
always the one returned by queries.
For Kafka messages, unique key column values can come from either the
message **payload** (the JSON value) or the message **key**. If a column
listed in `unique_key_columns` is missing from the payload, Cube Store
falls back to the Kafka message key: for a single unique key column, the
raw key value is used; for composite keys, the key is expected to be a
JSON object with matching field names.
### Output column types
In Kafka streams mode, Cube Store creates its internal pre-aggregation
table based on column type information. By default, column types are
inferred from the source ksqlDB stream using `DESCRIBE`. However, the
pre-aggregation's `select_statement` (generated from the rollup
definition) renames and transforms columns — for example, a source
column `CREATED_AT` becomes `order_events_stream__created_at_second` in
the output.
When this renaming happens, the raw source column types no longer match
the output column names, causing errors like:
> Key column `order_events_stream__id` not found among column definitions
To fix this, define `output_column_types` on the streaming
pre-aggregation. This tells Cube the exact output column types to use
for the Cube Store table, and separately passes the source schema so
Cube Store can deserialize the raw Kafka messages correctly.
```yaml title="YAML" theme={"dark"}
pre_aggregations:
- name: stream
type: rollup
read_only: true
measures:
- CUBE.count
- CUBE.total_amount
- CUBE.failed_count
dimensions:
- CUBE.order_id
- CUBE.user_id
- CUBE.status
unique_key_columns:
- order_id
time_dimension: CUBE.created_at
granularity: second
partition_granularity: day
# ksqlDB does not support NOW(); use CURRENT_TIMESTAMP with INTERVAL arithmetic instead.
build_range_start:
sql: "SELECT date_trunc('day', CURRENT_TIMESTAMP - INTERVAL '5 hour')"
build_range_end:
sql: "SELECT CURRENT_TIMESTAMP + INTERVAL '15 minute'"
refresh_key:
every: 1 minute
update_window: 1 hour
incremental: true
stream_offset: latest
output_column_types:
- member: CUBE.order_id
type: text
- member: CUBE.user_id
type: text
- member: CUBE.status
type: text
- member: CUBE.created_at.second
type: timestamp
- member: CUBE.count
type: int
- member: CUBE.total_amount
type: decimal
- member: CUBE.failed_count
type: int
```
```javascript title="JavaScript" theme={"dark"}
pre_aggregations: {
stream: {
type: `rollup`,
read_only: true,
measures: [CUBE.count, CUBE.total_amount, CUBE.failed_count],
dimensions: [CUBE.order_id, CUBE.user_id, CUBE.status],
unique_key_columns: [`order_id`],
time_dimension: CUBE.created_at,
granularity: `second`,
partition_granularity: `day`,
// ksqlDB does not support NOW(); use CURRENT_TIMESTAMP with INTERVAL arithmetic instead.
build_range_start: {
sql: `SELECT date_trunc('day', CURRENT_TIMESTAMP - INTERVAL '5 hour')`,
},
build_range_end: {
sql: `SELECT CURRENT_TIMESTAMP + INTERVAL '15 minute'`,
},
refresh_key: {
every: `1 minute`,
update_window: `1 hour`,
incremental: true,
},
stream_offset: `latest`,
output_column_types: [
{ member: CUBE.order_id, type: `text` },
{ member: CUBE.user_id, type: `text` },
{ member: CUBE.status, type: `text` },
{ member: CUBE.created_at.second, type: `timestamp` },
{ member: CUBE.count, type: `int` },
{ member: CUBE.total_amount, type: `decimal` },
{ member: CUBE.failed_count, type: `int` },
],
},
},
```
Each entry in `output_column_types` has two properties:
* `member` — a reference to a dimension or measure included in the
pre-aggregation. This must be a `CUBE` member reference (for example,
`CUBE.user_id`), not a string. For the time dimension, reference it
with the rollup granularity attached — for example,
`CUBE.created_at.second` when `granularity` is `second`.
* `type` — the Cube Store column type. Common values: `text`, `int`,
`bigint`, `decimal`, `float`, `boolean`, `timestamp`.
Include an entry for every dimension and measure in the pre-aggregation.
The time dimension must also be listed, with `type: timestamp` and the
granularity suffix on the member reference as shown above.
When `output_column_types` is defined, Cube uses the aliased column
names (matching the `select_statement`) for the Cube Store table
definition and passes the raw source schema separately via
`source_table`, so Cube Store knows how to deserialize incoming Kafka
messages. Without it, column names come from the raw ksqlDB `DESCRIBE`
output and will not match the aliased names in the `select_statement`
or `unique_key_columns`.
`output_column_types` is required for Kafka streams mode when the
pre-aggregation uses `unique_key_columns`. Without it, the unique key
column names will not match the table column definitions, causing the
pre-aggregation build to fail.
### Stream format
Cube Store expects Kafka messages to have a **JSON object** as their
value payload, with field names matching the column names defined in the
cube. For example, given the streaming cube above, each Kafka message
value should look like:
```json theme={"dark"}
{
"ORDER_ID": "ord_12345",
"USER_ID": "usr_789",
"STATUS": "completed",
"AMOUNT": 49.99,
"CREATED_AT": "2025-01-15T10:30:00.000"
}
```
Field names are case-sensitive and must match the column names used in
the `sql` property of each dimension and measure definition. Missing
fields default to `null`.
The message key is optional. When present and the value starts with `{`,
it is parsed as a JSON object and used as a fallback source for unique
key column values (see [above](#unique-key-columns-and-deduplication)).
#### Timestamp handling
For dimensions with `type: time`, Cube Store accepts timestamp values in
two formats:
* **String** — parsed using ISO 8601 / RFC 3339 formats. Supported
patterns include:
* `2025-01-15T10:30:00.000Z`
* `2025-01-15T10:30:00Z`
* `2025-01-15 10:30:00.000 UTC`
* `2025-01-15T10:30:00`
* `2025-01-15 10:30:00`
* `2025-01-15`
* **Number** — interpreted as **epoch milliseconds** (not seconds, not
microseconds). For example, `1736939400000` represents
`2025-01-15T10:30:00.000Z`.
If your Kafka topic produces timestamps as strings in a non-standard
format, you can use `PARSE_TIMESTAMP` in the cube's `sql` property to
convert them. In that case, define the source column as `type: string`
in a `source_table` and use the `select_statement` to transform it:
```javascript theme={"dark"}
sql: `SELECT PARSE_TIMESTAMP(TIMESTAMP_STR,
'yyyy-MM-dd''T''HH:mm:ss.SSSX', 'UTC') AS created_at,
ORDER_ID, USER_ID, STATUS, AMOUNT
FROM ORDER_EVENTS_STREAM`,
```
Time dimension truncation (controlled by the `granularity` property of
the pre-aggregation) is handled automatically. Cube generates the
appropriate `PARSE_TIMESTAMP(FORMAT_TIMESTAMP(CONVERT_TZ(...)))`
expression chain to truncate timestamps to the configured granularity
(e.g., `day`, `hour`, `minute`). Cube Store evaluates these expressions
natively on each micro-batch during ingestion. Standard SQL functions
like `date_trunc` are also available in the `select_statement`.
##### Converting `bigint` timestamps
When a source column stores a timestamp as a `bigint` (a common pattern
in Kafka topics — for example, a `created_at` field with values like
`1778800167128`), the value must be converted to a timestamp before it
can be used as a time dimension.
The conversion depends on the unit of the `bigint` value:
* **Milliseconds** (most common, e.g. `Date.now()` in JavaScript, Java's
`System.currentTimeMillis()`) — multiply by `1000` before casting,
because `CAST(... AS TIMESTAMP(6))` interprets numeric input as
**microseconds**. Without the multiplication, a millisecond value like
`1778800167128` is read as microseconds and produces a timestamp in
1970 instead of the intended date.
* **Microseconds** — cast directly to `TIMESTAMP(6)`.
* **Seconds** — multiply by `1000000` before casting.
Apply the conversion in the time dimension's `sql` property:
```javascript theme={"dark"}
dimensions: {
created_at: {
sql: `CAST(${CUBE}.created_at * 1000 AS TIMESTAMP(6))`,
type: `time`
}
}
```
The cast runs first; any granularity truncation generated by Cube
(`PARSE_TIMESTAMP(FORMAT_TIMESTAMP(CONVERT_TZ(...)))`) then operates on
the resulting `TIMESTAMP(6)` value.
When debugging a `bigint` timestamp issue, temporarily remove
`partition_granularity` from the pre-aggregation. This skips
partitioned builds and makes it easier to verify the conversion is
correct before re-enabling partitioning.
### Filtering on the stream
When the streaming cube defines a `sql` property with a `SELECT`
statement (rather than `sql_table`), Cube Store applies the projection
and any `WHERE` filters from that statement directly on each micro-batch
of incoming Kafka messages. This filtering happens inside Cube Store
using its query engine — it does not require ksqlDB to process the
filter. Only rows that pass the filter are ingested into the
pre-aggregation table.
This allows you to define a streaming cube that only ingests a subset of
the data from the underlying Kafka topic without creating any
server-side filter objects in ksqlDB.
#### Supported SQL syntax
The `SELECT` statement must follow a strict shape. Cube Store only
accepts plans that resolve to **Projection > Filter > TableScan** (where
the filter is optional). Any other query plan shape is rejected.
**Supported:**
* `SELECT` with column references (e.g., `SELECT col1, col2 FROM topic`)
* `SELECT *` wildcard
* Column aliases (`SELECT col1 AS my_alias`)
* `WHERE` clause with comparison operators (`=`, `!=`, `<`, `>`, `<=`,
`>=`)
* Boolean logic in `WHERE` (`AND`, `OR`, `NOT`)
* `IS NULL` and `IS NOT NULL`
* `IN` lists (`col IN (1, 2, 3)`)
* `BETWEEN` expressions
* `CASE ... WHEN ... THEN ... ELSE ... END` expressions
* `CAST(expr AS type)` type conversions
* `EXTRACT(field FROM expr)` for date/time parts
* `SUBSTRING(expr FROM start FOR length)`
* Scalar functions (e.g., `COALESCE`, `CONCAT`, arithmetic)
* `CONVERT_TZ` for timezone conversion (internally rewritten for
compatibility)
* `PARSE_TIMESTAMP` and `FORMAT_TIMESTAMP` for timestamp parsing and
formatting using ksql-style format strings (e.g.,
`yyyy-MM-dd'T'HH:mm:ss.SSS`)
* Nested expressions with parentheses
* `date_trunc` for timestamp truncation
**Not supported:**
* `JOIN` clauses — only a single `FROM` table is allowed
* Subqueries in `SELECT` or `WHERE`
* `GROUP BY`, `HAVING`, or aggregate functions (`SUM`, `COUNT`, `AVG`,
etc.)
* `ORDER BY` (rows are consumed in stream order)
* `LIMIT` and `OFFSET`
* `UNION`, `INTERSECT`, `EXCEPT`
* Window functions (`OVER`, `PARTITION BY`)
* Multiple `FROM` or multiple `WHERE` clauses
* Common Table Expressions (`WITH ... AS`)
All column expressions in the `SELECT` list that are not simple column
references must have explicit aliases. Unique key columns may reference
the source column through a scalar function (e.g.,
`CAST(id AS VARCHAR) AS id`), but not through arbitrary expressions.
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
[ref-streaming-pre-aggs]: /docs/pre-aggregations/using-pre-aggregations#streaming-pre-aggregations
[ref-decorated-env-vars]: /admin/connect-to-data/multiple-data-sources#under-the-hood
[ref-cube-data-source]: /reference/data-modeling/cube#data_source
[ref-lambda-pre-aggs]: /docs/pre-aggregations/lambda-pre-aggregations
# Materialize
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/materialize
The driver for Materialize is supported by its vendor. Please report any issues to their Slack.
The driver for Materialize is supported by its vendor. Please report any issues
to their [Slack][materialize-slack].
## Prerequisites
* The hostname for the [Materialize][materialize] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=materialize
CUBEJS_DB_HOST=my.materialize.host
CUBEJS_DB_PORT=6875
CUBEJS_DB_NAME=materialize
CUBEJS_DB_USER=materialize
CUBEJS_DB_PASS=materialize
CUBEJS_DB_MATERIALIZE_CLUSTER=quickstart
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ✅ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_DB_MATERIALIZE_CLUSTER`](/reference/configuration/environment-variables#cubejs_db_materialize_cluster) | The name of the Materialize cluster to connect to | A valid cluster name | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## SSL
To enable SSL-encrypted connections between Cube and Materialize, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
[materialize]: https://materialize.com/docs/
[materialize-slack]: https://materialize.com/s/chat
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
# MongoDB
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/mongodb
Query MongoDB from Cube through the BI Connector’s SQL interface, including community-driver and Atlas deprecation caveats.
[MongoDB](https://www.mongodb.com) is a popular document database. It can be
accessed using SQL via the [MongoDB Connector for BI][link-bi-connector], also
known as *BI Connector*.
The driver for MongoDB is community-supported and is not maintained by Cube or the database vendor.
BI Connector for MongoDB Atlas, cloud-based MongoDB service, is approaching
[end-of-life](https://www.mongodb.com/docs/atlas/bi-connection/#std-label-bi-connection).
It will be deprecated and no longer supported in June 2025.
## Prerequisites
To use Cube with MongoDB you need to install the [MongoDB Connector for
BI][mongobi-download]. [Learn more about setup for MongoDB
here][cube-blog-mongodb].
* [MongoDB Connector for BI][mongobi-download]
* The hostname for the [MongoDB][mongodb] database server
* The username/password for the [MongoDB][mongodb] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=mongobi
# The MongoBI connector host. If using on local machine, it should be either `localhost` or `127.0.0.1`:
CUBEJS_DB_HOST=my.mongobi.host
# The default port of the MongoBI connector service
CUBEJS_DB_PORT=3307
CUBEJS_DB_NAME=my_mongodb_database
CUBEJS_DB_USER=mongodb_server_user
CUBEJS_DB_PASS=mongodb_server_password
# MongoBI requires SSL connections, so set the following to `true`:
CUBEJS_DB_SSL=true
```
If you are connecting to a local MongoBI Connector, which is pointing to a local
MongoDB instance, If MongoBI Connector and MongoDB are both running locally,
then the above should work. To connect to a remote MongoDB instance, first
configure `mongosqld` appropriately. See [here for an example config
file][mongobi-with-remote-db].
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ✅ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
not be used in pre-aggregations when using MongoDB as a source database.
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ✅ | ✅ |
| Export Bucket | - | - |
By default, MongoDB uses [batching][self-preaggs-batching] to build
pre-aggregations.
### Batching
No extra configuration is required to configure batching for MongoDB.
### Export Bucket
MongoDB does not support export buckets.
## SSL
To enable SSL-encrypted connections between Cube and MongoDB, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
[mongodb]: https://www.mongodb.com/
[mongobi-with-remote-db]: https://docs.mongodb.com/bi-connector/current/reference/mongosqld/#example-configuration-file
[cube-blog-mongodb]: https://cube.dev/blog/building-mongodb-dashboard-using-nodejs
[mongobi-download]: https://www.mongodb.com/download-center/bi-connector
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-batching]: #batching
[link-bi-connector]: https://www.mongodb.com/docs/bi-connector/current/
# Microsoft Fabric
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/ms-fabric
Microsoft Fabric is an all-in-one analytics solution for enterprises. It offers a comprehensive suite of services, including a data warehouse.
[Microsoft Fabric][link-fabric] is an all-in-one analytics solution for
enterprises. It offers a comprehensive suite of services, including a
data warehouse.
Available on the [Enterprise plan](https://cube.dev/pricing).
[Contact us](https://cube.dev/contact) for details.
## Setup
When creating a new deployment in Cube Cloud, at the **Set up a database
connection** step, choose **Microsoft Fabric**.
Then, provide the **JDBC URL** connection string in the following
format. See below for tips on filling in ``, ``,
``, ``, and ``.
```text theme={"dark"}
jdbc:sqlserver://;serverName=.datawarehouse.pbidedicated.windows.net;database=;encrypt=true;Authentication=;UserName=;Password=
```
Optionally, fill in **Database** if you'd like to override the database
name from the JDBC URL.
### Server and database name
To obtain your data warehouse server name and database name, navigate to your
data warehouse in Microsoft Fabric and click on the cog icon to open
**Settings**:
On the **About** page, you can find the database name (``)
under **Name** and the server name (``) under
**SQL connection string**:
### Authentication
Microsoft Fabric supports two [authentication types][link-fabric-auth]:
* Use `ActiveDirectoryPassword` as `` to connect using a
Microsoft Entra principal name and password. Provide principal name as
`` and password as ``.
* Use `ActiveDirectoryServicePrincipal` as `` to connect
using the client ID and secret of a service principal identity. Provide
client ID as `` and secet as ``.
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
not be used in pre-aggregations when using Microsoft Fabric as a source
database.
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Simple | ✅ | ✅ |
| Batching | - | - |
| Export Bucket | - | - |
By default, Microsoft Fabric uses a [simple][self-preaggs-simple] strategy
to build pre-aggregations.
### Simple
No extra configuration is required to configure simple pre-aggregation builds
for Microsoft Fabric.
### Batching
Microsoft Fabric does not support batching.
### Export Bucket
Microsoft Fabric does not support export buckets.
[link-fabric]: https://www.microsoft.com/en-us/microsoft-fabric/
[link-fabric-auth]: https://learn.microsoft.com/en-us/sql/connect/jdbc/setting-the-connection-properties?view=sql-server-ver16#properties
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-simple]: #simple
# Microsoft SQL Server
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/ms-sql
Connect Cube to SQL Server with database host credentials, optional domain-scoped login, and SSL or pool tuning as needed.
## Prerequisites
* The hostname for the [Microsoft SQL Server][mssql] database server
* The username/password for the [Microsoft SQL Server][mssql] database server
* The name of the database to use within the [Microsoft SQL Server][mssql]
database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=mssql
CUBEJS_DB_HOST=my.mssql.host
CUBEJS_DB_NAME=my_mssql_database
CUBEJS_DB_USER=mssql_user
CUBEJS_DB_PASS=**********
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| --------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_DOMAIN`](/reference/configuration/environment-variables#cubejs_db_domain) | A domain name within the database to connect to | A valid domain name within a Microsoft SQL Server database | ❌ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_DB_MSSQL_USE_NAMED_TIMEZONES`](/reference/configuration/environment-variables#cubejs_db_mssql_use_named_timezones) | The flag to use time zone names (with DST-aware conversion) or numeric offsets for time zone conversion. Default is `false` | `true`, `false` | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
not be used in pre-aggregations when using Microsoft SQL Server as a source
database.
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Simple | ✅ | ✅ |
| Batching | - | - |
| Export Bucket | - | - |
By default, Microsoft SQL Server uses a [simple][self-preaggs-simple] strategy
to build pre-aggregations.
### Simple
No extra configuration is required to configure simple pre-aggregation builds
for Microsoft SQL Server.
### Batching
Microsoft SQL Server does not support batching.
### Export Bucket
Microsoft SQL Server does not support export buckets.
## SSL
To enable SSL-encrypted connections between Cube and Microsoft SQL Server, set
the [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how
to configure custom certificates, please check out [Enable SSL Connections to
the Database][ref-recipe-enable-ssl].
## Additional Configuration
### Windows Authentication
To connect to a Microsoft SQL Server database using Windows Authentication (also
sometimes known as `trustedConnection`), instantiate the driver with
`trustedConnection: true` in your `cube.js` configuration file:
```javascript theme={"dark"}
const MssqlDriver = require("@cubejs-backend/mssql-driver")
module.exports = {
driverFactory: ({ dataSource }) =>
new MssqlDriver({ database: dataSource, trustedConnection: true })
}
```
[mssql]: https://www.microsoft.com/en-gb/sql-server/sql-server-2019
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-simple]: #simple
# MySQL
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/mysql
Configure Cube’s MySQL client with TCP or Unix socket access, SSL, pooling limits, and optional timezone handling flags.
## Prerequisites
* The hostname for the [MySQL][mysql] database server
* The username/password for the [MySQL][mysql] database server
* The name of the database to use within the [MySQL][mysql] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=mysql
CUBEJS_DB_HOST=my.mysql.host
CUBEJS_DB_NAME=my_mysql_database
CUBEJS_DB_USER=mysql_user
CUBEJS_DB_PASS=**********
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| --------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_SOCKET_PATH`](/reference/configuration/environment-variables#cubejs_db_socket_path) | See [connecting via a Unix socket](#connecting-via-a-unix-socket) | A valid path to a socket | ❌ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_DB_MYSQL_USE_NAMED_TIMEZONES`](/reference/configuration/environment-variables#cubejs_db_mysql_use_named_timezones) | The flag to use time zone names or numeric offsets for time zone conversion. Default is `false` | `true`, `false` | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
not be used in pre-aggregations when using MySQL as a source database.
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ✅ | ✅ |
| Export Bucket | - | - |
By default, MySQL uses [batching][self-preaggs-batching] to build
pre-aggregations.
### Batching
No extra configuration is required to configure batching for MySQL.
### Export Bucket
MySQL does not support export buckets.
## SSL
To enable SSL-encrypted connections between Cube and MySQL, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
## Additional configuration
### Connecting via a Unix socket
To connect to a local MySQL database using a Unix socket, use the
[`CUBEJS_DB_SOCKET_PATH`](/reference/configuration/environment-variables#cubejs_db_socket_path) environment variable, for example:
```dotenv theme={"dark"}
CUBEJS_DB_SOCKET_PATH=/var/run/mysqld/mysqld.sock
```
When doing so, [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) will be ignored.
### Connecting to MySQL 8
MySQL 8 uses a new default authentication plugin (`caching_sha2_password`)
whereas previous version used `mysql_native_password`. Please
[see this answer on StackOverflow](https://stackoverflow.com/a/50377944)
for a workaround.
For additional details, check the [relevant issue](https://github.com/cube-js/cube/issues/3525) on GitHub.
[mysql]: https://www.mysql.com/
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-batching]: #batching
# Oracle
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/oracle
The driver for Oracle is community-supported and is not maintained by Cube or the database vendor.
The driver for Oracle is community-supported and is not maintained by Cube or the database vendor.
## Prerequisites
* The hostname for the [Oracle][oracle] database server
* The username/password for the [Oracle][oracle] database server
* The name of the database to use within the [Oracle][oracle] database server
The driver uses [`node-oracledb`][node-oracledb] in Thin mode, which supports Oracle Database
12.1 and later. Older servers (e.g. 11g) require [Thick mode][node-oracledb-thick], which the
driver does not currently enable — [track this issue](https://github.com/cube-js/cube/issues/10756).
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=oracle
CUBEJS_DB_HOST=my.oracle.host
CUBEJS_DB_NAME=my_oracle_database
CUBEJS_DB_USER=oracle_user
CUBEJS_DB_PASS=**********
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## SSL
To enable SSL-encrypted connections between Cube and Oracle, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
[oracle]: https://www.oracle.com/uk/index.html
[node-oracledb]: https://node-oracledb.readthedocs.io/en/latest/user_guide/introduction.html#architecture
[node-oracledb-thick]: https://node-oracledb.readthedocs.io/en/latest/user_guide/initialization.html#enabling-node-oracledb-thick-mode
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
# Apache Pinot
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/pinot
Broker-level connection details and required Pinot SQL features for wiring Cube to an Apache Pinot cluster via environment variables.
[Apache Pinot][link-pinot] is a real-time distributed OLAP datastore purpose-built
for low-latency, high-throughput analytics, and perfect for user-facing analytical
workloads. [StarTree][link-startree] is a fully-managed platform for Pinot.
## Prerequisites
* The hostname for the [Pinot][pinot] broker
* The port for the [Pinot][pinot] broker
Note that the following features should be enabled in your Pinot cluster:
* [Multi-stage query engine][link-pinot-msqe].
* [Advanced null value support][link-pinot-nvs].
## Setup
Currently, Apache Pinot is not shown in the list of available data sources in the UI.
However, you can still [configure](https://github.com/cube-js/cube/issues/8579#issuecomment-2453545048)
it using environment variables.
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=pinot
CUBEJS_DB_HOST=http[s]://pinot.broker.host
CUBEJS_DB_PORT=8099
CUBEJS_DB_USER=pinot_user
CUBEJS_DB_PASS=**********
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for your Pinot broker | A valid host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the broker | A valid username | ❌ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the broker | A valid password | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The database name for StarTree | A valid name | ❌ |
| [`CUBEJS_DB_PINOT_NULL_HANDLING`](/reference/configuration/environment-variables#cubejs_db_pinot_null_handling) | If `true`, enables null handling. Default is `false` | `true`, `false` | ❌ |
| [`CUBEJS_DB_PINOT_AUTH_TOKEN`](/reference/configuration/environment-variables#cubejs_db_pinot_auth_token) | The authentication token for StarTree | A valid token | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
be used in pre-aggregations when using Pinot as a source database. To learn more
about Pinot support for approximate aggregate functions, [click
here][pinot-docs-approx-agg-fns].
## Pre-aggregation build strategies
To learn more about pre-aggregation build strategies, \[head
here]\[ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Simple | ✅ | ✅ |
| Batching | - | - |
| Export bucket | - | - |
By default, Pinot uses a simple strategy to build pre-aggregations.
### Simple
No extra configuration is required to configure simple pre-aggregation builds
for Pinot.
### Batching
Pinot does not support batching.
### Export bucket
Pinot does not support export buckets.
## SSL
Cube does not require any additional configuration to enable SSL as Pinot connections are made over HTTPS.
[link-pinot]: https://pinot.apache.org/
[pinot]: https://docs.pinot.apache.org/
[link-pinot-msqe]: https://docs.pinot.apache.org/reference/multi-stage-engine
[link-pinot-nvs]: https://docs.pinot.apache.org/developers/advanced/null-value-support#advanced-null-handling-support
[pinot-docs-approx-agg-fns]: https://docs.pinot.apache.org/users/user-guide-query/query-syntax/how-to-handle-unique-counting
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[link-startree]: https://startree.ai
# Postgres
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/postgres
Supply PostgreSQL host, database name, credentials, and optional SSL or pool settings to connect Cube as a client.
## Prerequisites
* The hostname for the [Postgres][postgres] database server
* The username/password for the [Postgres][postgres] database server
* The name of the database to use within the [Postgres][postgres] database
server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=postgres
CUBEJS_DB_HOST=my.postgres.host
CUBEJS_DB_NAME=my_postgres_database
CUBEJS_DB_USER=postgres_user
CUBEJS_DB_PASS=**********
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
only be used in pre-aggregations when using the [Postgres HLL
extension][gh-postgres-hll] with Postgres as a source database.
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ✅ | ✅ |
| Export Bucket | - | - |
By default, Postgres uses [batching][self-preaggs-batching] to build
pre-aggregations.
### Batching
No extra configuration is required to configure batching for Postgres.
### Export Bucket
Postgres does not support export buckets.
## SSL
To enable SSL-encrypted connections between Cube and Postgres, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
## Additional Configuration
### AWS RDS
Use `CUBEJS_DB_SSL=true` to enable SSL if you have SSL enabled for your RDS
cluster. Download the new certificate [here][aws-rds-pem], and provide the
contents of the downloaded file to [`CUBEJS_DB_SSL_CA`](/reference/configuration/environment-variables#cubejs_db_ssl_ca). All other SSL-related
environment variables can be left unset. See [the SSL section][self-ssl] for
more details. More info on AWS RDS SSL can be found [here][aws-docs-rds-ssl].
### Google Cloud SQL
You can connect to an SSL-enabled MySQL database by setting [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) to
`true`. You may also need to set [`CUBEJS_DB_SSL_SERVERNAME`](/reference/configuration/environment-variables#cubejs_db_ssl_servername), depending on how
you are [connecting to Cloud SQL][gcp-docs-sql-connect].
### Heroku
Unless you're using a Private or Shield Heroku Postgres database, Heroku
Postgres does not currently support verifiable certificates. [Here is the
description of the issue from Heroku][heroku-postgres-issue].
As a workaround, you can set `rejectUnauthorized` option to `false` in the Cube
Postgres driver:
```javascript theme={"dark"}
const PostgresDriver = require("@cubejs-backend/postgres-driver")
module.exports = {
driverFactory: () =>
new PostgresDriver({
ssl: {
rejectUnauthorized: false
}
})
}
```
[aws-docs-rds-ssl]: https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/UsingWithRDS.SSL.html
[aws-rds-pem]: https://s3.amazonaws.com/rds-downloads/rds-ca-2019-root.pem
[gcp-docs-sql-connect]: https://cloud.google.com/sql/docs/postgres/connect-functions#connect
[gh-postgres-hll]: https://github.com/citusdata/postgresql-hll
[heroku-postgres-issue]: https://help.heroku.com/3DELT3RK/why-can-t-my-third-party-utility-connect-to-heroku-postgres-with-ssl
[postgres]: https://www.postgresql.org/
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-batching]: #batching
[self-ssl]: #ssl
# Presto
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/presto
Connect Cube to PrestoDB using hostname authentication together with catalog and schema selectors for your query engine layout.
## Prerequisites
* The hostname for the [Presto][presto] database server
* The username/password for the [Presto][presto] database server
* The name of the database to use within the [Presto][presto] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=prestodb
CUBEJS_DB_HOST=my.presto.host
CUBEJS_DB_USER=presto_user
CUBEJS_DB_PASS=**********
CUBEJS_DB_PRESTO_CATALOG=my_presto_catalog
CUBEJS_DB_SCHEMA=my_presto_schema
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_PRESTO_CATALOG`](/reference/configuration/environment-variables#cubejs_db_presto_catalog) | The catalog within Presto to connect to | A valid catalog name within a Presto database | ✅ |
| [`CUBEJS_DB_PRESTO_AUTH_TOKEN`](/reference/configuration/environment-variables#cubejs_db_presto_auth_token) | The authentication token to use when connecting to Presto/Trino. It will be sent in the `Authorization` header. | A valid authentication token | ❌ |
| [`CUBEJS_DB_SCHEMA`](/reference/configuration/environment-variables#cubejs_db_schema) | The schema within the database to connect to | A valid schema name within a Presto database | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_TYPE`](/reference/configuration/environment-variables#cubejs_db_export_bucket_type) | The export bucket type for pre-aggregations | `s3`, `gcs` | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET`](/reference/configuration/environment-variables#cubejs_db_export_bucket) | The export bucket to connect to | A valid bucket URL | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_AWS_KEY`](/reference/configuration/environment-variables#cubejs_db_export_bucket_aws_key) | The AWS Access Key ID to use for export bucket writes | A valid AWS Access Key ID | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET`](/reference/configuration/environment-variables#cubejs_db_export_bucket_aws_secret) | The AWS Secret Access Key to use for export bucket writes | A valid AWS Secret Access Key | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_AWS_REGION`](/reference/configuration/environment-variables#cubejs_db_export_bucket_aws_region) | The AWS region to use for export bucket writes | A valid AWS region | ❌ |
| [`CUBEJS_DB_EXPORT_GCS_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_export_gcs_credentials) | A Base64 encoded JSON key file for connecting to Google Cloud Storage | A valid Google Cloud JSON key file, encoded as a Base64 string | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
be used in pre-aggregations when using Presto as a source database. To learn
more about Presto support for approximate aggregate functions, [click
here][presto-docs-approx-agg-fns].
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Simple | ✅ | ✅ |
| Export Bucket | ❌ | ❌ |
By default, Presto uses a [simple][self-preaggs-simple] strategy to
build pre-aggregations.
### Simple
No extra configuration is required to configure simple pre-aggregation builds
for Presto.
### Export Bucket
Presto supports using both [AWS S3][aws-s3] and [Google Cloud Storage][google-cloud-storage] for export bucket functionality.
#### AWS S3
Ensure the AWS credentials are correctly configured in IAM to allow reads and
writes to the export bucket in S3.
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my.bucket.on.s3
CUBEJS_DB_EXPORT_BUCKET_AWS_KEY=
CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET=
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
#### Google Cloud Storage
When using an export bucket, remember to assign the **Storage Object Admin**
role to your Google Cloud credentials ([`CUBEJS_DB_EXPORT_GCS_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_export_gcs_credentials)).
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET=gs://presto-export-bucket
CUBEJS_DB_EXPORT_BUCKET_TYPE=gcs
CUBEJS_DB_EXPORT_GCS_CREDENTIALS=
```
## SSL
To enable SSL-encrypted connections between Cube and Presto, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
## Custom headers
The Presto driver supports forwarding custom HTTP headers on every request to
the Presto coordinator. This is useful for setting headers like
`X-Presto-Source`, `X-Presto-Client-Tags`, or any other custom header your
Presto coordinator expects.
Custom headers can't be configured via environment variables. Instead, use the
[`driver_factory`](/reference/configuration/config#driver_factory) configuration
option to pass a `headers` object to the driver:
```python title="Python" theme={"dark"}
from cube import config
@config('driver_factory')
def driver_factory(ctx: dict) -> dict:
return {
'type': 'prestodb',
'headers': {
'X-Presto-Source': 'cube',
'X-Presto-Client-Tags': 'user=alice@example.com'
}
}
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
driverFactory: ({ dataSource }) => ({
type: "prestodb",
headers: {
"X-Presto-Source": "cube",
"X-Presto-Client-Tags": "user=alice@example.com"
}
})
};
```
In multitenant deployments, you can use the [security
context](/embedding/authentication/security-context) to pass per-tenant headers,
for example to forward a user token from the API request down to Presto:
```python title="Python" theme={"dark"}
from cube import config
@config('driver_factory')
def driver_factory(ctx: dict) -> dict:
security_context = ctx['securityContext']
return {
'type': 'prestodb',
'headers': {
'X-Presto-Client-Tags': f"user={security_context['user_id']}",
'X-Custom-User-Token': security_context['token']
}
}
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
driverFactory: ({ securityContext }) => ({
type: "prestodb",
headers: {
"X-Presto-Client-Tags": `user=${securityContext.user_id}`,
"X-Custom-User-Token": securityContext.token
}
})
};
```
[aws-s3]: https://aws.amazon.com/s3/
[google-cloud-storage]: https://cloud.google.com/storage
[presto]: https://prestodb.io/
[presto-docs-approx-agg-fns]: https://prestodb.io/docs/current/functions/aggregate.html
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-simple]: #simple
# QuestDB
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/questdb
QuestDB is a high-performance time-series database which helps you overcome ingestion bottlenecks.
[QuestDB][questdb] is a high-performance [time-series database][time-series-database-glossary] which helps you overcome ingestion bottlenecks.
The driver for QuestDB is supported by its vendor. Please report any issues to
their [Slack][questdb-slack].
## Prerequisites
* The hostname for the QuestDB database server
* If QuestDB is not running, checkout the QuestDB [quick start][questdb-quick-start]
* [Docker][docker] (optional)
## Setup
### Docker
Create a dockerfile within the project directory:
```yaml title=docker-compose.yml theme={"dark"}
services:
cube:
environment:
- CUBEJS_DEV_MODE=true
image: "cubejs/cube:latest"
ports:
- "4000:4000"
volumes:
- ".:/cube/conf"
questdb:
container_name: questdb
hostname: questdb
image: "questdb/questdb:latest"
ports:
- "9000:9000"
- "8812:8812"
```
Within your project directory, create an `.env` file:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=questdb
CUBEJS_DB_HOST=my.questdb.host
CUBEJS_DB_PORT=8812
CUBEJS_DB_NAME=qdb
CUBEJS_DB_USER=admin
CUBEJS_DB_PASS=quest
```
Finally, bring it all up with Docker:
```bash title=shell theme={"dark"}
docker-compose up -d
```
Access Cube at [http://localhost:4000](http://localhost:4000) & QuestDB at [http://localhost:9000](http://localhost:9000)
Development mode is an authentication bypass. Cube is in development mode when
`CUBEJS_DEV_MODE=true` — which, under the `cubejs` CLI and the official Docker images,
also forces `NODE_ENV=development` and so switches off JWT verification on the
REST (JSON) and GraphQL APIs — and whenever `NODE_ENV` is not
`production`. Playground's endpoints are served with no authentication too: anyone who
can reach the instance is handed a ready-to-use API token, can mint others carrying any
security context, can read your data model files, and can overwrite your data model and
your `.env`. With `CUBEJS_DEV_MODE=true` and no
[`CUBEJS_SQL_PASSWORD`](/reference/configuration/environment-variables#cubejs_sql_password)
set, the SQL API accepts any credentials as well, allowing arbitrary SQL against
connected data sources. Use it only on a local development machine, never in
production. See
[`CUBEJS_DEV_MODE`](/reference/configuration/environment-variables#cubejs_dev_mode).
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ✅ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
[questdb]: https://questdb.io/
[questdb-slack]: https://slack.questdb.io/
[time-series-database-glossary]: https://questdb.io/glossary/time-series/database/
[questdb-quick-start]: https://questdb.io/docs/quick-start/
[docker]: https://docs.docker.com/get-docker/
# RisingWave
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/risingwave
Point Cube at RisingWave using the Postgres-compatible driver and your cluster host, database, and credential settings.
[RisingWave](https://risingwave.com) is a distributed streaming database that enables
processing and management of real-time data with a Postgres-style SQL syntax. It is
available as an [open-source edition](https://github.com/risingwavelabs/risingwave)
and a managed cloud service.
## Prerequisites
* [Credentials](https://docs.risingwave.com/docs/current/risingwave-docker-compose/#connect-to-risingwave) for the RisingWave cluster
## Setup
RisingWave provides a Postgres-compatible interface. You can connect Cube to RisingWave
as if it's a regular [Postgres][ref-postgres] data source.
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=postgres
CUBEJS_DB_HOST=risingwave_host
CUBEJS_DB_PORT=risingwave_port
CUBEJS_DB_NAME=risingwave_database
CUBEJS_DB_USER=risingwave_user
CUBEJS_DB_PASS=risingwave_password
```
## Environment variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables [SSL encryption](https://docs.risingwave.com/docs/current/secure-connections-with-ssl-tls/) for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-aggregation feature support
### `count_distinct_approx`
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
not be used in pre-aggregations when using RisingWave as a source database.
## Pre-aggregation build strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Simple | ✅ | ✅ |
| Batching | - | - |
| Export bucket | - | - |
By default, RisingWave uses a simple strategy to build pre-aggregations.
### Simple
No extra configuration is required to configure simple pre-aggregation builds
for RisingWave.
### Batching
RisingWave does not support batching.
### Export bucket
RisingWave does not support export buckets.
[ref-postgres]: /admin/connect-to-data/data-sources/postgres
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
# SingleStore
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/singlestore
SingleStore is a distributed SQL database that offers high-throughput transactions (inserts and upserts) and low-latency analytics.
[SingleStore][link-singlestore] is a distributed SQL database that offers
high-throughput transactions (inserts and upserts) and low-latency analytics.
Available on the [Enterprise plan](https://cube.dev/pricing).
[Contact us](https://cube.dev/contact) for details.
## Setup
When creating a new deployment in Cube Cloud, at the **Set up a database
connection** step, choose **SingleStore**.
Then, provide credetials: host name, database name, user name (`admin` by
default), and password.
### Host name
To obtain the host name, navigate to the necessary group and workspace
in the side bar, then click **Connect** and choose **CLI client**:
On the **Connect to Workspace** page, you can find the host name
(`svc--dml..svc.singlestore.com`):
### Database name
To obtain the database name, navigate to the necessary group and workspace
in the side bar. You'll see databases on the right:
### Password
You have set the password when creating the workspace. If you'd like to reset
the password, navigate to the **Access** tab within your workspace.
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
not be used in pre-aggregations when using SingleStore as a source
database.
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Simple | ✅ | ✅ |
| Batching | - | - |
| Export Bucket | - | - |
By default, SingleStore uses a [simple][self-preaggs-simple] strategy
to build pre-aggregations.
### Simple
No extra configuration is required to configure simple pre-aggregation builds
for SingleStore.
### Batching
SingleStore does not support batching.
### Export Bucket
SingleStore does not support export buckets.
[link-singlestore]: https://www.singlestore.com
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[self-preaggs-simple]: #simple
# Snowflake
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/snowflake
Snowflake is a popular cloud-based data platform.
[Snowflake][snowflake] is a popular cloud-based data platform.
## Prerequisites
In order to connect Cube to Snowflake, you need to grant certain permissions to the Snowflake role
used by Cube. **Read-only access is sufficient to query data.** Cube requires the role to have
`USAGE` on the warehouse, databases, and schemas, and `SELECT` on tables. An example configuration:
```sql theme={"dark"}
GRANT USAGE ON WAREHOUSE MY_WAREHOUSE TO ROLE XYZ;
GRANT USAGE ON DATABASE ABC TO ROLE XYZ;
GRANT USAGE ON ALL SCHEMAS IN DATABASE ABC TO ROLE XYZ;
GRANT USAGE ON FUTURE SCHEMAS IN DATABASE ABC TO ROLE XYZ;
GRANT SELECT ON ALL TABLES IN DATABASE ABC TO ROLE XYZ;
GRANT SELECT ON FUTURE TABLES IN DATABASE ABC TO ROLE XYZ;
```
Write access is only needed for [pre-aggregations](#pre-aggregation-build-strategies) built inside
Snowflake with the default [batching](#batching) strategy (`CREATE TABLE` on the pre-aggregations
schema). The [export bucket](#export-bucket) strategy writes to the bucket, not the source.
* [Account/Server URL][snowflake-docs-account-id] for Snowflake.
* User name and password or an RSA private key for the Snowflake account.
In Cube Cloud, you can authenticate with [OIDC workload
identity][ref-oidc-overview] instead — the same role grants apply to the
mapped service user.
* Optionally, the warehouse name, the user role, and the database name.
## Setup
If you're having Network error and Snowflake can't be reached please make sure you tried
[format 2 for an account id][snowflake-format-2].
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=snowflake
CUBEJS_DB_SNOWFLAKE_ACCOUNT=XXXXXXXXX.us-east-1
CUBEJS_DB_SNOWFLAKE_WAREHOUSE=MY_SNOWFLAKE_WAREHOUSE
CUBEJS_DB_NAME=my_snowflake_database
CUBEJS_DB_USER=snowflake_user
CUBEJS_DB_PASS=**********
```
### Cube Cloud
In some cases you'll need to allow connections from your Cube Cloud deployment
IP address to your database. You can copy the IP address from either the
Database Setup step in deployment creation, or from **Settings →
Configuration** in your deployment.
The following fields are required when creating a Snowflake connection:
#### OIDC workload identity
Instead of a password or key pair, Cube Cloud deployments can authenticate
to Snowflake with [OIDC workload identity][ref-oidc-overview]: a Snowflake
External OAuth integration trusts Cube's OIDC issuer, and the driver
presents a short-lived Cube-minted JWT — no long-lived secrets to
provision or rotate. Snowflake validates each connection's token against
your tenant's public JWKS, maps the token's `sub` claim to a Snowflake
user, and authorizes the session role through the token's `scp` claim.
Start with the [OIDC overview][ref-oidc-overview] for the concepts (issuer,
token configs, custom claims, target env var) and to enable OIDC for your
tenant, then:
In **Admin → OIDC**, click **Add Config** and create a config with:
| Field | Value |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------- |
| **Audience Type** | `Custom` |
| **Name** | `snowflake` |
| **Custom Audience** | `https://.snowflakecomputing.com` |
| **Subject Claim Format** | Keep the default, or pick a richer template — the rendered `sub` must match the Snowflake user's `LOGIN_NAME` below |
| **Custom Claims** | Claim `scp` with value `session:role-any` |
| **Target Env Var** | `CUBEJS_DB_SNOWFLAKE_OAUTH_TOKEN_PATH` |
The `scp` claim is the part that's easy to miss: Snowflake grants
session roles **exclusively** through it — a token without `scp`
authenticates but fails role authorization. `session:role-any` lets the
driver request any role granted to the mapped user; to pin the role
inside the token instead, use `session:role:` and
`EXTERNAL_OAUTH_ANY_ROLE_MODE = 'DISABLE'` below. The **Target Env Var**
is how the driver finds the token: Cube sets that env var to the token
file path in every execution context (deployed pods, dev mode, and test
connection each keep the file in a different place), so the path is
never written by hand.
In a Snowflake worksheet, as `ACCOUNTADMIN`:
```sql theme={"dark"}
CREATE SECURITY INTEGRATION CUBE_CLOUD_EXTERNAL_OAUTH
TYPE = EXTERNAL_OAUTH
ENABLED = TRUE
EXTERNAL_OAUTH_TYPE = CUSTOM
EXTERNAL_OAUTH_ISSUER = 'https://.cubecloud.dev'
EXTERNAL_OAUTH_JWS_KEYS_URL = 'https://.cubecloud.dev/.well-known/jwks.json'
EXTERNAL_OAUTH_AUDIENCE_LIST = ('https://.snowflakecomputing.com')
EXTERNAL_OAUTH_TOKEN_USER_MAPPING_CLAIM = 'sub'
EXTERNAL_OAUTH_SNOWFLAKE_USER_MAPPING_ATTRIBUTE = 'LOGIN_NAME'
EXTERNAL_OAUTH_ANY_ROLE_MODE = 'ENABLE';
```
The issuer and audience must match the token config exactly. Snowflake
fetches the JWKS from your tenant's public endpoint, so no key material
changes hands.
The integration maps the token's `sub` claim to a Snowflake user via
`LOGIN_NAME`. With the default subject claim format the rendered `sub`
is `cube:deployment:` — the deployment ID is the number
in your deployment's console URL. If you chose a different template,
use the token-config dialog's live preview to see the exact rendered
value.
```sql theme={"dark"}
CREATE ROLE CUBE_ROLE;
CREATE USER CUBE_SVC
TYPE = SERVICE
LOGIN_NAME = 'cube:deployment:'
DEFAULT_ROLE = CUBE_ROLE
DEFAULT_WAREHOUSE = ;
GRANT ROLE CUBE_ROLE TO USER CUBE_SVC;
GRANT USAGE ON WAREHOUSE TO ROLE CUBE_ROLE;
-- Grant the role read access to the data Cube serves, e.g.:
GRANT USAGE ON DATABASE TO ROLE CUBE_ROLE;
GRANT USAGE ON ALL SCHEMAS IN DATABASE TO ROLE CUBE_ROLE;
GRANT SELECT ON ALL TABLES IN DATABASE TO ROLE CUBE_ROLE;
```
`TYPE = SERVICE` blocks password logins for this user entirely — it
can only authenticate through the federation.
Set the `OAUTH` authenticator and omit `CUBEJS_DB_USER` /
`CUBEJS_DB_PASS`:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=snowflake
CUBEJS_DB_SNOWFLAKE_ACCOUNT=
CUBEJS_DB_SNOWFLAKE_WAREHOUSE=
CUBEJS_DB_NAME=
CUBEJS_DB_SNOWFLAKE_ROLE=CUBE_ROLE
CUBEJS_DB_SNOWFLAKE_AUTHENTICATOR=OAUTH
```
`CUBEJS_DB_SNOWFLAKE_OAUTH_TOKEN_PATH` is **not** set here — Cube
populates it automatically thanks to the **Target Env Var** from the
token config. The driver re-reads the file on every new connection, so
the broker's automatic refresh is picked up without restarts.
To verify, run any query against the Snowflake data source. On the
Snowflake side, a successful federation shows up in the login history with
`OAUTH_ACCESS_TOKEN` as the authentication factor:
```sql theme={"dark"}
SELECT event_timestamp, user_name, first_authentication_factor, is_success
FROM TABLE(SNOWFLAKE.INFORMATION_SCHEMA.LOGIN_HISTORY_BY_USER(USER_NAME => 'CUBE_SVC'))
ORDER BY event_timestamp DESC;
```
Common failure modes:
| Error | Cause |
| ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |
| `Invalid OAuth access token` | Issuer or audience mismatch between the token config and the security integration, or the JWKS URL is unreachable. |
| `The role … is not listed in the Access Token or was filtered` | The `scp` custom claim is missing from the token config, or the role isn't granted to the mapped user. |
| `User … not found` / mapping errors | The rendered `sub` doesn't match the Snowflake user's `LOGIN_NAME` — compare against the dialog's live preview. |
| `File … provided by CUBEJS_DB_SNOWFLAKE_OAUTH_TOKEN_PATH does not exist` | The **Target Env Var** isn't set on the token config (or a stale hand-written path is configured), or the deployment is still starting up. |
Cube Cloud also supports connecting to data sources within private VPCs
if [single-tenant infrastructure][ref-dedicated-infra] is used. Check out the
[VPC connectivity guide][ref-cloud-conf-vpc] for details.
### Query tagging
Set a [query tag][snowflake-docs-query-tag] on the connection so that queries
Cube sends to Snowflake are labeled in `QUERY_HISTORY`. There's no environment
variable for this; set `queryTag` from a custom [`driverFactory`][ref-driver-factory]
instead of returning a plain options object:
```javascript theme={"dark"}
const SnowflakeDriver = require('@cubejs-backend/snowflake-driver');
module.exports = {
driverFactory: () =>
new SnowflakeDriver({
queryTag: 'cube',
}),
};
```
The tag is set once per connection, not per query — all queries issued over
that connection carry the same tag.
[ref-dedicated-infra]: /admin/deployment/infrastructure#dedicated-infrastructure
[ref-cloud-conf-vpc]: /admin/deployment/dedicated
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| --------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------- | :-----------: |
| [`CUBEJS_DB_SNOWFLAKE_ACCOUNT`](/reference/configuration/environment-variables#cubejs_db_snowflake_account) | The Snowflake account identifier to use when connecting to the database | [A valid Snowflake account ID][snowflake-docs-account-id] | ✅ |
| [`CUBEJS_DB_SNOWFLAKE_REGION`](/reference/configuration/environment-variables#cubejs_db_snowflake_region) | The Snowflake region to use when connecting to the database | [A valid Snowflake region][snowflake-docs-regions] | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_WAREHOUSE`](/reference/configuration/environment-variables#cubejs_db_snowflake_warehouse) | The Snowflake warehouse to use when connecting to the database | [A valid Snowflake warehouse][snowflake-docs-warehouse] in the account | ✅ |
| [`CUBEJS_DB_SNOWFLAKE_ROLE`](/reference/configuration/environment-variables#cubejs_db_snowflake_role) | The Snowflake role to use when connecting to the database | [A valid Snowflake role][snowflake-docs-roles] in the account | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_CLIENT_SESSION_KEEP_ALIVE`](/reference/configuration/environment-variables#cubejs_db_snowflake_client_session_keep_alive) | If `true`, [keep the Snowflake connection alive indefinitely][snowflake-docs-connection-options] | `true`, `false` | ❌ |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅1 |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅1 |
| [`CUBEJS_DB_SNOWFLAKE_AUTHENTICATOR`](/reference/configuration/environment-variables#cubejs_db_snowflake_authenticator) | The type of authenticator to use with Snowflake. Use `SNOWFLAKE` with username/password, or `SNOWFLAKE_JWT` with key pairs. Defaults to `SNOWFLAKE` | `SNOWFLAKE`, `SNOWFLAKE_JWT`, `OAUTH` | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_PRIVATE_KEY`](/reference/configuration/environment-variables#cubejs_db_snowflake_private_key) | The content of the private RSA key | Content of the private RSA key (encrypted or not) | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_PRIVATE_KEY_PATH`](/reference/configuration/environment-variables#cubejs_db_snowflake_private_key_path) | The path to the private RSA key | A valid path to the private RSA key | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_PRIVATE_KEY_PASS`](/reference/configuration/environment-variables#cubejs_db_snowflake_private_key_pass) | The password for the private RSA key. Only required for encrypted keys | A valid password for the encrypted private RSA key | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_OAUTH_TOKEN`](/reference/configuration/environment-variables#cubejs_db_snowflake_oauth_token) | The OAuth token | A valid OAuth token (string) | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_OAUTH_TOKEN_PATH`](/reference/configuration/environment-variables#cubejs_db_snowflake_oauth_token_path) | The path to the valid oauth toket file | A valid path for the oauth token file | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_HOST`](/reference/configuration/environment-variables#cubejs_db_snowflake_host) | Host address to which the driver should connect | A valid hostname | ❌ |
| [`CUBEJS_DB_SNOWFLAKE_QUOTED_IDENTIFIERS_IGNORE_CASE`](/reference/configuration/environment-variables#cubejs_db_snowflake_quoted_identifiers_ignore_case) | Whether or not [quoted identifiers should be case insensitive][link-snowflake-quoted-identifiers]. Default is `false` | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `20` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
1 Required when using password-based authentication. Not required with `SNOWFLAKE_JWT` (key pair) or with `OAUTH` — including [OIDC workload identity](#oidc-workload-identity) in Cube Cloud, where the driver reads a Cube-minted token from [`CUBEJS_DB_SNOWFLAKE_OAUTH_TOKEN_PATH`](/reference/configuration/environment-variables#cubejs_db_snowflake_oauth_token_path).
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
be used in pre-aggregations when using Snowflake as a source database. To learn
more about Snowflake's support for approximate aggregate functions, [click
here][snowflake-docs-approx-agg-fns].
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Batching | ❌ | ✅ |
| Export Bucket | ❌ | ❌ |
By default, Snowflake uses [batching][self-preaggs-batching] to build
pre-aggregations.
### Batching
No extra configuration is required to configure batching for Snowflake.
### Export Bucket
Snowflake supports using both [AWS S3][aws-s3] and [Google Cloud
Storage][google-cloud-storage] for export bucket functionality.
The export bucket strategy unloads with `COPY INTO` directly from the query — no
temporary tables. The role only needs the read-only grants above plus `USAGE` on
the storage integration.
#### AWS S3
Ensure proper IAM privileges are configured for S3 bucket reads and writes, using either
storage integration or user credentials for Snowflake and either IAM roles/IRSA or user
credentials for Cube Store, with mixed configurations supported.
Using IAM user credentials for both:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my.bucket.on.s3
CUBEJS_DB_EXPORT_BUCKET_AWS_KEY=
CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET=
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
Using a [Storage Integration][snowflake-docs-aws-integration] to write to export buckets and
user credentials to read from Cube Store:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my.bucket.on.s3
CUBEJS_DB_EXPORT_INTEGRATION=aws_int
CUBEJS_DB_EXPORT_BUCKET_AWS_KEY=
CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET=
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
Using a Storage Integration to write to export bucket and IAM role/IRSA to read from Cube Store:\*\*
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my.bucket.on.s3
CUBEJS_DB_EXPORT_INTEGRATION=aws_int
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
#### Google Cloud Storage
When using an export bucket, remember to assign the **Storage Object Admin**
role to your Google Cloud credentials ([`CUBEJS_DB_EXPORT_GCS_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_export_gcs_credentials)).
Before configuring Cube, an [integration must be created and configured in
Snowflake][snowflake-docs-gcs-integration]. Take note of the integration name
(`gcs_int` from the example link) as you'll need it to configure Cube.
Once the Snowflake integration is set up, configure Cube using the following:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET=snowflake-export-bucket
CUBEJS_DB_EXPORT_BUCKET_TYPE=gcs
CUBEJS_DB_EXPORT_GCS_CREDENTIALS=
CUBEJS_DB_EXPORT_INTEGRATION=gcs_int
```
#### Azure
To use Azure Blob Storage as an export bucket, follow [the guide on
using a Snowflake storage integration (Option 1)][snowflake-docs-azure].
Take note of the integration name (`azure_int` from the example link)
as you'll need it to configure Cube.
[Retrieve the storage account access key][azure-bs-docs-get-key] from your Azure
account.
Once the Snowflake integration is set up, configure Cube using the following:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=azure
CUBEJS_DB_EXPORT_BUCKET=wasbs://my-container@my-storage-account.blob.core.windows.net
CUBEJS_DB_EXPORT_BUCKET_AZURE_KEY=
CUBEJS_DB_EXPORT_INTEGRATION=azure_int
```
## SSL
Cube does not require any additional configuration to enable SSL as Snowflake
connections are made over HTTPS.
[aws-s3]: https://aws.amazon.com/s3/
[google-cloud-storage]: https://cloud.google.com/storage
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
[ref-oidc-overview]: /admin/deployment/oidc
[ref-driver-factory]: /reference/configuration/config#driver_factory
[self-preaggs-batching]: #batching
[snowflake]: https://www.snowflake.com/
[snowflake-docs-account-id]: https://docs.snowflake.com/en/user-guide/admin-account-identifier.html
[snowflake-docs-connection-options]: https://docs.snowflake.com/en/developer-guide/node-js/nodejs-driver-options#additional-connection-options
[snowflake-docs-query-tag]: https://docs.snowflake.com/en/developer-guide/node-js/nodejs-driver-options#additional-connection-options
[snowflake-docs-aws-integration]: https://docs.snowflake.com/en/user-guide/data-load-s3-config-storage-integration
[snowflake-docs-gcs-integration]: https://docs.snowflake.com/en/user-guide/data-load-gcs-config.html
[snowflake-docs-regions]: https://docs.snowflake.com/en/user-guide/intro-regions.html
[snowflake-docs-roles]: https://docs.snowflake.com/en/user-guide/security-access-control-overview.html#roles
[snowflake-docs-approx-agg-fns]: https://docs.snowflake.com/en/sql-reference/functions/approx_count_distinct.html
[snowflake-docs-warehouse]: https://docs.snowflake.com/en/user-guide/warehouses.html
[snowflake-format-2]: https://docs.snowflake.com/en/user-guide/admin-account-identifier#format-2-account-locator-in-a-region
[snowflake-docs-azure]: https://docs.snowflake.com/en/user-guide/data-load-azure-config#option-1-configuring-a-snowflake-storage-integration
[azure-bs-docs-get-key]: https://docs.microsoft.com/en-us/azure/storage/common/storage-account-keys-manage?toc=%2Fazure%2Fstorage%2Fblobs%2Ftoc.json&tabs=azure-portal#view-account-access-keys
[link-snowflake-quoted-identifiers]: https://docs.snowflake.com/en/sql-reference/identifiers-syntax#double-quoted-identifiers
# SQLite
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/sqlite
The driver for SQLite is community-supported and is not maintained by Cube or the database vendor.
The driver for SQLite is community-supported and is not maintained by Cube or the database vendor.
## Prerequisites
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=sqlite
CUBEJS_DB_NAME=my_sqlite_database
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | --------------------- | :------: |
| [`CUBEJS_DB_NAME`](/reference/configuration/environment-variables#cubejs_db_name) | The name of the database to connect to | A valid database name | ✅ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## SSL
SQLite does not support SSL connections.
# Trino
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/trino
Map Cube to Trino with host credentials plus catalog, schema, SSL, and optional bearer-token variables on the shared Presto driver.
## Prerequisites
* The hostname for the [Trino][trino] database server
* The username/password for the [Trino][trino] database server
* The name of the database to use within the [Trino][trino] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=trino
CUBEJS_DB_HOST=my.trino.host
CUBEJS_DB_USER=trino_user
CUBEJS_DB_PASS=**********
CUBEJS_DB_PRESTO_CATALOG=my_trino_catalog
CUBEJS_DB_SCHEMA=my_trino_schema
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_PRESTO_CATALOG`](/reference/configuration/environment-variables#cubejs_db_presto_catalog) | The catalog within Presto to connect to | A valid catalog name within a Presto database | ✅ |
| [`CUBEJS_DB_PRESTO_AUTH_TOKEN`](/reference/configuration/environment-variables#cubejs_db_presto_auth_token) | The authentication token to use when connecting to Presto/Trino. It will be sent in the `Authorization` header. | A valid authentication token | ❌ |
| [`CUBEJS_DB_SCHEMA`](/reference/configuration/environment-variables#cubejs_db_schema) | The schema within the database to connect to | A valid schema name within a Presto database | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_TYPE`](/reference/configuration/environment-variables#cubejs_db_export_bucket_type) | The export bucket type for pre-aggregations | `s3`, `gcs` | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET`](/reference/configuration/environment-variables#cubejs_db_export_bucket) | The export bucket to connect to | A valid bucket URL | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_AWS_KEY`](/reference/configuration/environment-variables#cubejs_db_export_bucket_aws_key) | The AWS Access Key ID to use for export bucket writes | A valid AWS Access Key ID | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET`](/reference/configuration/environment-variables#cubejs_db_export_bucket_aws_secret) | The AWS Secret Access Key to use for export bucket writes | A valid AWS Secret Access Key | ❌ |
| [`CUBEJS_DB_EXPORT_BUCKET_AWS_REGION`](/reference/configuration/environment-variables#cubejs_db_export_bucket_aws_region) | The AWS region to use for export bucket writes | A valid AWS region | ❌ |
| [`CUBEJS_DB_EXPORT_GCS_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_export_gcs_credentials) | A Base64 encoded JSON key file for connecting to Google Cloud Storage | A valid Google Cloud JSON key file, encoded as a Base64 string | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## Pre-Aggregation Feature Support
### count\_distinct\_approx
Measures of type
[`count_distinct_approx`][ref-schema-ref-types-formats-countdistinctapprox] can
be used in pre-aggregations when using Trino as a source database. To learn more
about Trino support for approximate aggregate functions, [click
here][trino-docs-approx-agg-fns].
## Pre-Aggregation Build Strategies
To learn more about pre-aggregation build strategies, [head
here][ref-caching-using-preaggs-build-strats].
| Feature | Works with read-only mode? | Is default? |
| ------------- | :------------------------: | :---------: |
| Simple | ✅ | ✅ |
| Export Bucket | ❌ | ❌ |
By default, Trino uses a [simple][self-preaggs-simple] strategy to
build pre-aggregations.
### Simple
No extra configuration is required to configure simple pre-aggregation builds
for Trino.
### Export Bucket
Trino supports using both [AWS S3][aws-s3] and [Google Cloud Storage][google-cloud-storage] for export bucket functionality.
#### AWS S3
Ensure the AWS credentials are correctly configured in IAM to allow reads and
writes to the export bucket in S3.
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my.bucket.on.s3
CUBEJS_DB_EXPORT_BUCKET_AWS_KEY=
CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET=
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=
```
#### Google Cloud Storage
When using an export bucket, remember to assign the **Storage Object Admin**
role to your Google Cloud credentials ([`CUBEJS_DB_EXPORT_GCS_CREDENTIALS`](/reference/configuration/environment-variables#cubejs_db_export_gcs_credentials)).
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET=gs://trino-export-bucket
CUBEJS_DB_EXPORT_BUCKET_TYPE=gcs
CUBEJS_DB_EXPORT_GCS_CREDENTIALS=
```
## SSL
To enable SSL-encrypted connections between Cube and Trino, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
## Custom headers
The Trino driver supports forwarding custom HTTP headers (e.g., `X-Trino-Source`,
`X-Trino-Routing-Group`, `X-Trino-Client-Tags`, or any other custom header) on
every request to the Trino coordinator. See the [Trino client
protocol][trino-docs-client-protocol] for the list of headers accepted by Trino.
Custom headers can't be configured via environment variables. Instead, use the
[`driver_factory`](/reference/configuration/config#driver_factory) configuration
option to pass a `headers` object to the driver:
```python title="Python" theme={"dark"}
from cube import config
@config('driver_factory')
def driver_factory(ctx: dict) -> dict:
return {
'type': 'trino',
'headers': {
'X-Trino-Source': 'cube',
'X-Trino-Routing-Group': 'etl',
'X-Trino-Client-Tags': 'user=alice@example.com'
}
}
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
driverFactory: ({ dataSource }) => ({
type: "trino",
headers: {
"X-Trino-Source": "cube",
"X-Trino-Routing-Group": "etl",
"X-Trino-Client-Tags": "user=alice@example.com"
}
})
};
```
In multitenant deployments, you can use the [security
context](/embedding/authentication/security-context) to pass per-tenant headers, for example
to forward a user token from the API request down to Trino:
```python title="Python" theme={"dark"}
from cube import config
@config('driver_factory')
def driver_factory(ctx: dict) -> dict:
security_context = ctx['securityContext']
return {
'type': 'trino',
'headers': {
'X-Trino-Client-Tags': f"user={security_context['user_id']}",
'X-Custom-User-Token': security_context['token']
}
}
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
driverFactory: ({ securityContext }) => ({
type: "trino",
headers: {
"X-Trino-Client-Tags": `user=${securityContext.user_id}`,
"X-Custom-User-Token": securityContext.token
}
})
};
```
[trino-docs-client-protocol]: https://trino.io/docs/current/develop/client-protocol.html
[aws-s3]: https://aws.amazon.com/s3/
[google-cloud-storage]: https://cloud.google.com/storage
[ref-caching-using-preaggs-build-strats]: /docs/pre-aggregations/using-pre-aggregations#pre-aggregation-build-strategies
[self-preaggs-simple]: #simple
[trino]: https://trino.io/
[trino-docs-approx-agg-fns]: https://trino.io/docs/current/functions/aggregate.html#approximate-aggregate-functions
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
[ref-schema-ref-types-formats-countdistinctapprox]: /reference/data-modeling/measures#type
# Vertica
Source: https://docs.cube.dev/admin/connect-to-data/data-sources/vertica
OpenText Analytics Database (Vertica) is a columnar database designed for big data analytics.
[OpenText Analytics Database][opentext-adb] (Vertica) is a columnar database
designed for big data analytics.
## Prerequisites
* The hostname for the [Vertica][vertica] database server
* The username/password for the [Vertica][vertica] database server
* The name of the database to use within the [Vertica][vertica] database server
## Setup
### Manual
Add the following to a `.env` file in your Cube project:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=vertica
CUBEJS_DB_HOST=my.vertica.host
CUBEJS_DB_USER=vertica_user
CUBEJS_DB_PASS=**********
CUBEJS_DB_SCHEMA=my_vertica_schema
```
## Environment Variables
| Environment Variable | Description | Possible Values | Required |
| ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | -------------------------------------------- | :------: |
| [`CUBEJS_DB_HOST`](/reference/configuration/environment-variables#cubejs_db_host) | The host URL for a database | A valid database host URL | ✅ |
| [`CUBEJS_DB_PORT`](/reference/configuration/environment-variables#cubejs_db_port) | The port for the database connection | A valid port number | ❌ |
| [`CUBEJS_DB_USER`](/reference/configuration/environment-variables#cubejs_db_user) | The username used to connect to the database | A valid database username | ✅ |
| [`CUBEJS_DB_PASS`](/reference/configuration/environment-variables#cubejs_db_pass) | The password used to connect to the database | A valid database password | ✅ |
| [`CUBEJS_DB_SCHEMA`](/reference/configuration/environment-variables#cubejs_db_schema) | The schema within the database to connect to | A valid schema name within a Presto database | ✅ |
| [`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) | If `true`, enables SSL encryption for database connections from Cube | `true`, `false` | ❌ |
| [`CUBEJS_DB_MAX_POOL`](/reference/configuration/environment-variables#cubejs_db_max_pool) | The maximum number of concurrent database connections to pool. Default is `8` | A valid number | ❌ |
| [`CUBEJS_CONCURRENCY`](/reference/configuration/environment-variables#cubejs_concurrency) | The number of [concurrent queries][ref-data-source-concurrency] to the data source | A valid number | ❌ |
[ref-data-source-concurrency]: /admin/connect-to-data/concurrency#data-source-concurrency
## SSL
To enable SSL-encrypted connections between Cube and Verica, set the
[`CUBEJS_DB_SSL`](/reference/configuration/environment-variables#cubejs_db_ssl) environment variable to `true`. For more information on how to
configure custom certificates, please check out [Enable SSL Connections to the
Database][ref-recipe-enable-ssl].
[opentext-adb]: https://www.opentext.com/products/analytics-database?o=vert
[vertica]: https://www.vertica.com/documentation/vertica/all/
[ref-recipe-enable-ssl]: /recipes/configuration/using-ssl-connections-to-data-source
# Connect to data
Source: https://docs.cube.dev/admin/connect-to-data/index
Connect your data warehouse or database to Cube and configure connection settings.
Connect your data warehouse, database, or query engine to Cube. Once connected,
you can configure multiple data sources, tune concurrency, and set up
multitenancy.
## Getting started
Choose and configure a data warehouse, query engine, or other data source.
Connect to multiple databases so different cubes reference different sources.
Optimize query queue settings and control load on your data sources.
## Recipes
Enable TLS to upstream databases with custom CA bundles and client certificates.
Reduce warehouse spend through pre-aggregation strategy and workload-aware settings.
Route each tenant to its own database while reusing a single data model.
# Connecting to multiple data sources
Source: https://docs.cube.dev/admin/connect-to-data/multiple-data-sources
Manage one or more data source connections per Cube Cloud deployment from a single place.
A Cube Cloud deployment always has a single **default** data source plus,
optionally, one or more **named** data sources. [Cubes](/reference/data-modeling/cube)
reference whichever source they need via the
[`data_source`](/reference/data-modeling/cube#data_source) property.
Manage data sources from your deployment's **Settings → Data Sources** page.
## Adding a data source
Click **+ Add data source**, pick the driver, give the source a short
name (see [Source name rules](#source-name-rules)), and fill in the
connection form — the same form the deployment wizard uses.
Cube Cloud tests the connection before saving; on success it persists the
env vars and restarts the development environment, on failure it surfaces
the driver's error and saves nothing.
## Editing and deleting
Use the **Edit** action on a row to update credentials, or the **Delete**
action to remove a named source. The **default** source can be edited but
not deleted — to remove it, clear the bare `CUBEJS_DB_*` variables from
[**Settings → Environment Variables**][ref-config-ref-env] by hand.
## Source name rules
* Letters, digits, and underscores; must start with a letter.
* `default` is reserved; use the **Use as default** button on the naming
step instead.
* Flat names (`analytics`, `warehouse`) are safer than names with
underscores, which can collide with the env-var parser when they end in
`_db`, `_aws`, or `_jdbc`.
## Using a data source from a cube
Reference a source from the
[`data_source`](/reference/data-modeling/cube#data_source) property using
the name you gave it when you created it:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders_from_other_data_source
# ...
data_source: analytics
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders_from_other_data_source`, {
// ...
data_source: `analytics`
})
```
Cubes that don't set `data_source` use the default source — there's no
need to write `data_source: default` explicitly, though you can if you
prefer to be explicit.
A single query can also span sources: see [querying across data
sources](/recipes/data-modeling/cross-data-source-queries) for appending rows
from cubes in different databases into one result set.
For multitenancy scenarios where the data source is selected dynamically
per request, see [`driver_factory`](/reference/configuration/config#driver_factory)
and the [multitenancy guide][ref-config-multitenancy].
## Workbooks
Sources show up immediately in
[Workbooks](/docs/explore-analyze/workbooks) — open a
[Source SQL tab](/docs/explore-analyze/workbooks/source-sql-tabs) and they
appear in the data sources sidebar without needing a cube to reference
them first.
## Under the hood
The UI is a friendlier surface over the deployment's environment
variables — inspect or hand-edit them from
[**Settings → Environment Variables**][ref-config-ref-env] at any time.
The default source uses bare `CUBEJS_DB_*` variables; each named source
uses a `CUBEJS_DS__*` prefix. `CUBEJS_DATASOURCES` is auto-maintained
whenever at least one named source exists, with `default` first and named
entries lowercased:
```dotenv theme={"dark"}
CUBEJS_DATASOURCES=default,analytics
CUBEJS_DB_TYPE=postgres
CUBEJS_DB_HOST=localhost
CUBEJS_DS_ANALYTICS_DB_TYPE=postgres
CUBEJS_DS_ANALYTICS_DB_HOST=remotehost
```
For the full list of variables that support decoration, see the
[environment variables reference][ref-config-ref-env].
Each source can also build and store its pre-aggregations on a dedicated connection
using the `CUBEJS_DS__PRE_AGGREGATIONS_DB_*` variables. See [pre-aggregation data
source][ref-preagg-data-source] for details.
[ref-config-ref-env]: /reference/configuration/environment-variables
[ref-preagg-data-source]: /docs/pre-aggregations/refreshing-pre-aggregations#pre-aggregation-data-source
[ref-config-multitenancy]: /embedding/multitenancy#multitenancy-multitenancy-vs-multiple-data-sources
# Set up per-user OAuth
Source: https://docs.cube.dev/admin/connect-to-data/oauth-authentication
Configure Cube to authenticate each user with their own OAuth token, falling back to a service account for liveness checks.
This feature is in beta. Reach out to your account manager to have it
enabled for your Cube Cloud deployment.
## Use case
You want each user's queries to run under their own database identity
using OAuth tokens managed by Cube Cloud. When a user's token is
unavailable or expired, Cube falls back to a service account so that
connectivity checks and background operations still work.
This pattern applies to any data source that supports OAuth, including
[Databricks][ref-databricks-jdbc] and [Snowflake][ref-snowflake]. The
examples below use Databricks; switch the `userCredentials` key and
driver options for any other OAuth-capable data source.
Because every user connects with different credentials, you also need
per-user query orchestrator state. Without this, one user's cached
connection could leak to another.
## Prerequisites
* A [Cube Cloud][ref-cube-cloud] deployment connected to an
OAuth-capable data source
* OAuth configured in your data source so that Cube Cloud can
obtain per-user tokens (via the **User Credentials** feature)
* A service account credential (token or password) stored as an
environment variable for fallback connectivity
The service account credential is used only as a fallback for Cube's
internal liveness checks and background operations. Grant it the minimum
permissions necessary — ideally read-only access to the required schemas —
to limit exposure if the credential is compromised.
## Set up the OAuth app
Before configuring Cube to use per-user OAuth, register your data
source as an OAuth app in Cube Cloud:
In Cube Cloud, go to **Admin → Integrations → OAuth apps** and click
**Add**.
Provide the OAuth app metadata from your data source: **Name**,
**Auth URL**, **Token URL**, **Client ID**, **Client Secret**, and any
required **Scopes**. Copy the **Redirect URI** shown in this form and
register it with your data source's OAuth provider, then click
**Create**.
Open the sidebar and go to **Connected apps**. Find your OAuth app
and click **Authorize** to generate an access token.
You'll need to repeat this step whenever the token expires.
## Configuration
The configuration uses two options from the
[configuration file reference][ref-config]:
* [`driver_factory`][ref-driver-factory] — dynamically selects the
authentication credential per request
* [`context_to_orchestrator_id`][ref-context-to-orchestrator-id] — gives
each user their own query orchestrator instance (database connections,
execution queues, pre-aggregation table caches)
### Environment variables
Set the environment variables for your data source. The examples below
show Databricks and Snowflake; adapt them to your specific setup.
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=databricks-jdbc
CUBEJS_DB_DATABRICKS_URL=jdbc:databricks://dbc-XXXXXXX-XXXX.cloud.databricks.com:443/default;transportMode=http;ssl=1;httpPath=sql/protocolv1/o/XXXXX/XXXXX;AuthMech=3;UID=token
CUBEJS_DB_DATABRICKS_TOKEN=dapi_service_account_token
CUBEJS_DB_DATABRICKS_ACCEPT_POLICY=true
# Optional: specify a catalog
CUBEJS_DB_DATABRICKS_CATALOG=my_catalog
```
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=snowflake
CUBEJS_DB_SNOWFLAKE_ACCOUNT=XXXXXXXXX.us-east-1
CUBEJS_DB_SNOWFLAKE_WAREHOUSE=MY_SNOWFLAKE_WAREHOUSE
CUBEJS_DB_NAME=my_snowflake_database
CUBEJS_DB_USER=service_account_user
CUBEJS_DB_PASS=service_account_password
CUBEJS_DB_SNOWFLAKE_ROLE=MY_ROLE
```
### Configuration file
The examples below use Databricks. To target a different data source,
swap `userCredentials.databricks` for the matching key (for example,
`userCredentials.snowflake`) and update the `driver_factory` return
value with the correct `type` and driver-specific options. See the
[data sources reference][ref-data-sources] for available drivers.
```python cube.py theme={"dark"}
from cube import config
import os
@config("driver_factory")
def driver_factory(ctx: dict) -> dict:
# Extract the Cube Cloud security context, which contains
# per-user OAuth credentials when available.
# For other data sources, swap "databricks" for "snowflake", etc.
databricks_creds = (
ctx
.get("securityContext", {})
.get("cubeCloud", {})
.get("userCredentials", {})
.get("databricks", {})
)
# Only use the OAuth token when the credential status is "active".
# An expired or revoked token falls back to the service account.
oauth_token = (
databricks_creds.get("accessToken")
if databricks_creds.get("status") == "active"
else None
)
return {
"type": "databricks-jdbc",
"url": os.environ["CUBEJS_DB_DATABRICKS_URL"],
# Prefer the user's OAuth token; fall back to the service account token
"token": oauth_token or os.environ["CUBEJS_DB_DATABRICKS_TOKEN"],
"acceptPolicy": True,
"catalog": os.environ.get("CUBEJS_DB_DATABRICKS_CATALOG"),
}
@config("context_to_orchestrator_id")
def context_to_orchestrator_id(ctx: dict) -> str:
# Give each user a separate orchestrator instance (DB connections,
# execution queues, pre-aggregation caches)
username = (
ctx
.get("securityContext", {})
.get("cubeCloud", {})
.get("username", "default")
)
return f"CUBE_APP_{username}"
```
```javascript cube.js theme={"dark"}
module.exports = {
driverFactory: ({ securityContext }) => {
// Extract the Cube Cloud security context, which contains
// per-user OAuth credentials when available.
// For other data sources, swap `databricks` for `snowflake`, etc.
const databricksCreds =
securityContext?.cubeCloud?.userCredentials?.databricks ?? {};
// Only use the OAuth token when the credential status is "active".
// An expired or revoked token falls back to the service account.
const oauthToken =
databricksCreds.status === "active"
? databricksCreds.accessToken
: null;
return {
type: "databricks-jdbc",
url: process.env.CUBEJS_DB_DATABRICKS_URL,
// Prefer the user's OAuth token; fall back to the service account token
token: oauthToken || process.env.CUBEJS_DB_DATABRICKS_TOKEN,
acceptPolicy: true,
catalog: process.env.CUBEJS_DB_DATABRICKS_CATALOG,
};
},
// Give each user a separate orchestrator instance (DB connections,
// execution queues, pre-aggregation caches)
contextToOrchestratorId: ({ securityContext }) => {
const username = securityContext?.cubeCloud?.username ?? "default";
return `CUBE_APP_${username}`;
},
};
```
## How it works
1. **User makes a request** — Cube Cloud attaches the user's OAuth
credentials to `securityContext.cubeCloud.userCredentials.`
(for example, `.databricks` or `.snowflake`).
2. **`driver_factory` resolves the credential** — If the credential status
is `active`, the user's OAuth token is used. Otherwise, Cube falls back
to the service account credential stored in environment variables.
3. **Per-user orchestrator** —
[`context_to_orchestrator_id`][ref-context-to-orchestrator-id] returns
a unique key per username, so each user gets their own database
connection pool, execution queues, and pre-aggregation table cache.
Without this, Cube would share a single cached connection across all
users, causing one user's credentials to be reused for another user's
queries.
[ref-config]: /reference/configuration/config
[ref-driver-factory]: /reference/configuration/config#driver_factory
[ref-context-to-orchestrator-id]: /reference/configuration/config#context_to_orchestrator_id
[ref-databricks-jdbc]: /admin/connect-to-data/data-sources/databricks-jdbc
[ref-snowflake]: /admin/connect-to-data/data-sources/snowflake
[ref-data-sources]: /admin/connect-to-data/data-sources
[ref-cube-cloud]: /docs/introduction
# App theme
Source: https://docs.cube.dev/admin/customization/app-theme
Brand Cube with your own colors, logo, and fonts — applied to embedded surfaces and, optionally, across the entire Cube Cloud console.
The **app theme** controls the brand appearance of Cube — the accent color, background and text colors, logo, and fonts. You pick a handful of colors per color scheme and Cube generates a complete, accessible light and dark palette from them.
By default the app theme brands [embedded surfaces](/embedding/iframe/creator-mode) (Creator Mode, published dashboards, and embedded chat) so they match your product. You can also turn it on for the entire Cube Cloud console so everyone in your account sees your branding.
## Requirements
Configuring the app theme requires the **Admin** [role][ref-roles]. The theme is account-wide: there is a single app theme per account, and changes apply to every deployment in it.
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
The app theme is distinct from **[dashboard themes](/admin/customization/dashboard-themes)**, which style individual dashboards (backgrounds, widget cards, borders). The app theme sets your account's overall brand; dashboard themes are a narrower override.
## What you can customize
The app theme is configured for two color schemes — **Light** and **Dark** — plus shared **Typography**.
| Section | What it controls |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| **Accent color** | Your primary brand color. Cube derives the full palette (buttons, links, active states, highlights) from its hue and saturation. |
| **Background color** | The main surface and page background. |
| **Foreground color** | The primary text color. |
| **Contrast** | A 0–100 control (45 is neutral) for how strongly the generated palette separates surfaces and text. Higher values increase contrast. |
| **Logo** | The mark shown in the navigation header. |
| **Typography** | The web fonts used for body and heading text (shared across both schemes). |
You only choose the accent, background, and foreground colors — Cube generates every other token (borders, muted text, success/danger states, hover and active variants) for both light and dark mode, keeping contrast accessible automatically.
## Configure the theme
Go to **Admin → Customization → App Theme**. The page has a **Light theme** and a **Dark theme** editor on the left and a live preview on the right.
For each scheme, pick the **Accent**, **Background**, and **Foreground** colors and adjust the **Contrast** slider. The preview updates as you edit.
Paste an HTTPS URL to your logo image (see [Logo](#logo) below).
Add one or more web fonts and choose which to use for body and heading text (see [Fonts](#fonts) below).
Use the **Light / Auto / Dark** toggle above the preview to check both schemes. **Auto** follows your current OS/app appearance.
Click **Save**. To start over, use **Restore defaults** (and **Revert** to undo a restore before saving).
### Logo
The logo is referenced by URL and must be served over **HTTPS**. Use a small, square mark (around **24×24**) in **PNG**, **SVG**, or **JPG** format — it renders in the navigation header. If you set a logo only for the light scheme, it is reused in dark mode unless you set a dark-mode logo as well.
### Fonts
Fonts are shared across the light and dark schemes. You register one or more fonts, then choose which font and weight to use for body and heading text.
You can add a font in two ways:
* **A direct font file** — an HTTPS URL ending in a font file: `woff2`, `woff`, `ttf` (TrueType), or `otf` (OpenType).
* **A CSS stylesheet** — an HTTPS URL to a CSS file containing `@font-face` declarations (for example, a Google Fonts stylesheet link). A single stylesheet can register multiple font families at once.
In the **Typography** section, paste the font or stylesheet URL and add it. Cube validates the URL and reads the font's metadata; a stylesheet may register several families in one step.
Choose a font from your registry for **Body** and for **Heading** text, and set the **font weight** for each.
Font URLs must be HTTPS and publicly reachable. Cube fetches the font to validate it; private, non-HTTPS, or oversized files are rejected.
## Where the theme is applied
### Embedded surfaces
When configured, the app theme automatically brands your [embedded](/embedding/iframe/customization) experiences — the [Creator Mode](/embedding/iframe/creator-mode) app, published dashboards, and embedded chat — including overlays such as dialogs and menus. No extra configuration is required beyond saving the theme.
### Across Cube Cloud
Turn on **Apply across Cube Cloud** (the switch at the top of the App Theme page) to apply the same theme to the **entire authenticated Cube Cloud console** — the sidebar, pages, and overlays — for every user in the account.
* It takes effect immediately for all users on their next load.
* It respects each user's light/dark preference, switching schemes accordingly.
* Unauthenticated pages (such as sign-in) keep Cube's default appearance.
Turn the switch off to return the console to Cube's default look while keeping your embedded surfaces branded.
# Chart palettes
Source: https://docs.cube.dev/admin/customization/chart-palettes
Create custom color palettes so charts across your account stay consistent and on-brand.
Chart palettes control the colors used in chart visualizations across [workbooks][ref-workbooks] and [dashboards][ref-dashboards]. Cube ships with a set of built-in palettes, and you can also define your own custom palettes — for example, to match your company brand — and reuse them on any chart in the account.
Custom palettes are managed centrally and are available across every deployment in the account.
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-dashboards]: /docs/explore-analyze/dashboards
## Requirements
Custom chart palettes are managed under the **Manage chart palettes** permission. The built-in **Admin** [role][ref-roles] has this permission by default; for other users, grant the **Manage chart palettes** global permission via a [custom role][ref-custom-roles].
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
## Browse palettes
Go to **Admin → Customization → Chart Palettes** to see every palette available in your account. The list combines:
* **Built-in palettes** shipped with Cube — these can be used in charts but cannot be edited or deleted.
* **Custom palettes** created in your account — these can be edited, renamed, and deleted.
Each row shows the palette name and a preview of its colors.
## Create a custom palette
1. Go to **Admin → Customization → Chart Palettes**.
2. Click **Add palette**.
3. Give the palette a **Name** and provide its colors as a comma-separated list of hex codes (for example, `#1F77B4, #FF7F0E, #2CA02C`).
4. Click **Save**.
The new palette appears in the list immediately and becomes selectable in the chart styling controls.
## Edit a custom palette
1. Go to **Admin → Customization → Chart Palettes**.
2. Click the row action on the palette you want to change and select **Edit**.
3. Update the name or colors and click **Save**.
Existing charts that use the palette pick up the new colors automatically.
## Delete a custom palette
1. Go to **Admin → Customization → Chart Palettes**.
2. Click the row action on the palette you want to remove and select **Delete**.
Charts that were using the deleted palette fall back to the built-in **Default** palette.
## Apply a palette to a chart
Palettes are applied per chart from the chart's styling controls:
1. Open a workbook and select a chart.
2. Open the **Style** panel for the chart.
3. Use the **Palette** dropdown to pick a built-in or custom palette.
The dropdown shows palettes that match the chart's color encoding:
* **Discrete** palettes appear when the color encoding is categorical (for example, color by `status` or `category`).
* **Continuous** palettes appear when the color encoding is quantitative or temporal (for example, color by `total_sale_price` or `created_at`).
If no palette is explicitly selected, charts use the built-in **Default** palette.
# Dashboard themes
Source: https://docs.cube.dev/admin/customization/dashboard-themes
Style dashboards with reusable themes that control backgrounds, widget cards, borders, titles, and text.
Dashboard themes control the look and feel of [dashboards][ref-dashboards] — including the dashboard background, widget cards, borders, titles, and text. Cube ships with a built-in theme, and you can also create your own custom themes and reuse them across dashboards in the account.
Themes are designed and saved from the dashboard builder, then managed centrally in the admin area.
[ref-dashboards]: /docs/explore-analyze/dashboards
## Requirements
Custom dashboard themes are managed under the **Manage dashboard themes** permission. The built-in **Admin** [role][ref-roles] has this permission by default; for other users, grant the **Manage dashboard themes** global permission via a [custom role][ref-custom-roles].
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
## What you can theme
A theme defines styles for the following parts of a dashboard:
| Section | Properties |
| ------------- | ----------------------------------------------------- |
| **Dashboard** | Background color, padding |
| **Widgets** | Card background color, padding |
| **Borders** | Color, width, style, corner radius |
| **Titles** | Color, font size, font weight, font family |
| **Text** | Color, secondary color, font family, code font family |
The built-in **Cube** theme provides Cube's default look and is always available. It can be selected on dashboards but cannot be edited, renamed, or deleted.
## Browse themes
Go to **Admin → Customization → Dashboard Themes** to see every theme available in your account. The list combines the built-in **Cube** theme and any custom themes saved from dashboards.
For each custom theme you can:
* **Rename** the theme.
* **Delete** the theme.
Themes are created from the dashboard builder rather than from this page — see [Create a custom theme](#create-a-custom-theme) below.
## Create a custom theme
Custom themes are saved from the dashboard builder's **Styling** tab.
Open a dashboard in a workbook and switch to the dashboard builder.
Click the dashboard options button and switch to the **Styling** tab.
Adjust the dashboard background, widget card styles, borders, titles, and text until the dashboard looks the way you want.
Click **Save as new**, give the theme a name, and confirm.
The new theme appears in the **Admin → Customization → Dashboard Themes** list and becomes selectable in the **Styling** tab on every other dashboard in the account.
## Apply a theme to a dashboard
1. Open the dashboard in the dashboard builder.
2. Open dashboard options and switch to the **Styling** tab.
3. Pick a theme from the theme picker.
You can also tweak any styling properties on top of the selected theme — those changes are stored as overrides on the dashboard and do not modify the underlying theme.
## Update a saved theme
After tweaking styling on a dashboard that uses a custom theme, you can push those changes back into the theme so every other dashboard using it picks them up.
1. From the dashboard's **Styling** tab, make your changes.
2. Click **Save** to write the current effective styles back into the theme.
Saving a theme updates **every dashboard** that uses it. Already-published dashboard snapshots keep their previous styling until they are re-published.
## Delete a custom theme
1. Go to **Admin → Customization → Dashboard Themes**.
2. Click the row action on the theme you want to remove and select **Delete**.
Dashboards that were using the deleted theme revert to the built-in **Cube** theme.
## Theme scope
* **Themes** are defined once and shared across all deployments in the account.
* **Theme assignment** is per dashboard — different dashboards in the same workbook can use different themes.
* New dashboards use the built-in **Cube** theme until you assign a custom theme.
# Continuous deployment
Source: https://docs.cube.dev/admin/deployment/continuous-deployment
This guide covers features and tools you can use to deploy your Cube project to Cube Cloud.
Each Cube deployment is configured to use one of the following deploy modes:
* **Deploy with Git** — Cube Cloud builds and deploys your project from a Git
repository. Use this mode when you want commits to the production branch
(or merges performed from the Cube Cloud UI) to automatically trigger a
production build.
* **Deploy with CLI** — you push your project to Cube Cloud from your local
machine or CI/CD pipeline by running the `cubejs-cli deploy` command. The
production branch is **never** built automatically; it only changes when
somebody explicitly runs `cubejs-cli deploy`.
You can change the deploy mode at any time on the **Build & Deploy** tab of
the **Settings** screen.
Independently of the deploy mode, you can [connect a GitHub
repository](#deploy-with-github) to the deployment. Non-production branches
pushed to GitHub are then auto-synced into Cube Cloud and reflected in their
[staging environments][ref-environments-staging]. The production branch
follows the deploy mode: it auto-deploys in Git mode and stays manual in CLI
mode.
## Deploy with Git
Continuous deployment works by connecting a Git repository to a Cube Cloud
deployment and keeping the two in sync.
First, go to the **Build & Deploy** tab on the **Settings** screen and
select **Deploy with Git** under **Deploy with**. From the same screen you
can pick the production branch, connect a Git repository, and copy the
commands that set up Cube Cloud as a Git remote:
Click **Generate Git credentials** to obtain Git credentials. The instructions
to set up Cube Cloud as a Git remote are also available on the same screen:
```bash theme={"dark"}
git config credential.helper store
git remote add cubecloud
git push cubecloud master
```
## Deploy with GitHub
You can connect a GitHub repository to your deployment by clicking the
**Connect to GitHub** button on the **Build & Deploy** tab of the
**Settings** screen and selecting your repository. This works with both the
**Deploy with Git** and **Deploy with CLI** deploy modes.
If your organization uses SAML SSO for GitHub authentication, make sure to start an active SAML
session prior to connecting to your GitHub account from Cube.
What happens after connecting depends on the deploy mode:
* In **Deploy with Git** mode, Cube Cloud will automatically deploy from the
specified production branch (`master` by default) on every push, and also
sync non-production branches into their staging environments.
* In **Deploy with CLI** mode, only non-production branches are auto-synced
into their staging environments. See [Connecting GitHub on a CLI
deployment](#connecting-github-on-a-cli-deployment) for details.
## Deploy with CLI
When a deployment is configured to deploy with CLI, production builds are
**only** triggered when you (or your CI/CD pipeline) push the project to
Cube Cloud by running:
```bash theme={"dark"}
npx cubejs-cli deploy --token
```
You can find your Deploy Token, along with copy-paste instructions and an
example CI/CD pipeline config (such as a GitHub Actions workflow), on the
**Build & Deploy** tab of the **Settings** screen. The same screen lets you
regenerate the token if it needs to be rotated.
Cube Cloud still maintains an internal Git repository for the deployment so
that features such as [development mode][ref-dev-mode] and branching can work.
This internal repository is overwritten on every successful CLI deploy — only
the contents of your most recent `cubejs-cli deploy` invocation are used to
build and run production.
### Connecting GitHub on a CLI deployment
You can still [connect a GitHub repository](#deploy-with-github) to a CLI
deployment. When you do:
* Pushes to **non-production branches** (typically feature branches) are
picked up by the GitHub webhook and automatically synced into Cube Cloud,
so the corresponding [staging environments][ref-environments-staging]
reflect the latest commits without any manual step.
* Pushes to the **production branch** are intentionally **not**
auto-deployed in CLI mode. The production branch only changes when
somebody runs `cubejs-cli deploy` against the deployment.
This gives you GitHub-driven previews for feature work while keeping the
production deploy gated on an explicit CLI invocation (for example, from your
CI/CD pipeline after a release is approved).
### Development mode with CLI deployments
[Development mode][ref-dev-mode] is fully available on CLI deployments:
developers can enter dev mode, switch branches, save changes, commit them,
and even merge branches into the production branch from the Cube Cloud UI.
However, in CLI mode, **none of these actions trigger a production build or
redeploy**. Specifically:
* Saving or committing changes in dev mode only updates the per-developer
[development environment][ref-environments-dev] and the corresponding
branch in the internal repository. Production is not affected.
* Merging a branch into the production branch from the Cube Cloud UI updates
the production branch in the internal repository but **does not** rebuild
or redeploy the [production environment][ref-environments-prod].
* Any changes made through the UI are **overwritten on the next
`cubejs-cli deploy`**, which uploads the contents of your local project to
Cube Cloud and triggers a production build.
In other words, a CLI-mode deployment will not change production just because
somebody pressed **Enter Development Mode** and merged a work-in-progress
branch — production only changes when somebody (or some pipeline) explicitly
runs `cubejs-cli deploy` against the deployment.
If you want changes made in the Cube Cloud UI to flow to production
automatically, switch the deployment to [Deploy with Git](#deploy-with-git).
[ref-dev-mode]: /docs/data-modeling/dev-mode
[ref-environments-dev]: /admin/deployment/environments#development-environments
[ref-environments-prod]: /admin/deployment/environments#production-environment
[ref-environments-staging]: /admin/deployment/environments#staging-environments
# Custom domains
Source: https://docs.cube.dev/admin/deployment/custom-domains
Maps your own DNS hostname to Cube Cloud deployment APIs using the custom domain flow and required CNAME records.
By default, Cube Cloud deployments and their API endpoints use auto-generated
anonymized domain names, e.g., `emerald-llama.gcp-us-central1.cubecloudapp.dev`.
You can also assign a custom domain to your deployment.
Available on [Premium and above plans](https://cube.dev/pricing).
To set up a custom domain, go to your deployment's settings page. Under
the **Custom Domain** section, type in your custom domain and
click **Add**:
After doing this, copy the provided domain and add the following `CNAME` records
to your DNS provider:
| Name | Value |
| ------------- | ------------------------------------------- |
| `YOUR_DOMAIN` | The copied value from Cube Cloud's settings |
Using the example subdomain from the screenshot, a `CNAME` record for `acme.dev`
would look like:
| Name | Value |
| ---------- | ----------------------------------------------------- |
| `insights` | `colossal-crowville.gcp-us-central1.cubecloudapp.dev` |
DNS changes can sometimes take up to 15 minutes to propagate, please wait at
least 15 minutes and/or try using another DNS provider to verify the `CNAME`
record correctly before raising a new support ticket.
# Bring Your Own Cloud on AWS
Source: https://docs.cube.dev/admin/deployment/dedicated/aws/byoc
Prerequisites, IAM access, and provisioning steps for deploying BYOC inside your own AWS account, managed by the Cube Operator.
With Bring Your Own Cloud (BYOC) on AWS, all the components interacting with private data are deployed on
the customer infrastructure on AWS and managed by the Cube Control Plane via the Cube Operator.
This document provides step-by-step instructions for deploying Cube BYOC on AWS.
Available on the [Enterprise plan](https://cube.dev/pricing).
[Contact us](https://cube.dev/contact) for details. For private API access from your
applications and BI tools, see [Private API Connectivity][private-api-connectivity].
[private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
## Prerequisites
The bulk of provisioning work will be done remotely by Cube automation.
However, to get started, you'll need to provide Cube with the necessary access
along with some additional information that includes:
* **AWS Account ID:** The AWS account ID of the target deployment account
[the AWS Console][aws-console].
* **AWS Region:** [The AWS region][aws-docs-regions] where the BYOC resources
should be deployed.
In addition to that, you'll need to make sure you have sufficient access to create
the `CubeCloudBYOC` IAM role that would allow Cube to:
* Create and manage a VPC
* Create one or more EKS clusters
* Create and maintain necessary IAM roles and policies
* Configure VPC networking
* Run ec2 instances
* Manage ec2 autoscaling
* Manage S3 buckets
* Manage CloudWatch Logs
* Create and manage RDS PostgreSQL instances
**Restrictive SCPs may block RDS provisioning.** Amazon RDS needs a service-linked
role named `AWSServiceRoleForRDS` in the account. The policy below grants
`iam:CreateServiceLinkedRole` for `rds.amazonaws.com`, so Cube creates it during
provisioning — but if your organization's SCPs deny IAM writes, the grant is
present in the policy without being effective in the account, and provisioning
fails with a message that names a parameter rather than a permission:
```
InvalidParameterValue: Unable to create the resource.
Verify that you have permission to create service linked role.
```
To rule this out, run the following once in the target account before granting
Cube access. `InvalidInput: ... has been taken in this account` means the role
already exists and no action is needed.
```bash theme={"dark"}
aws iam create-service-linked-role --aws-service-name rds.amazonaws.com
```
If the command itself is denied by an SCP, the same denial will block Cube during
provisioning — ask your AWS organization administrator to permit
`iam:CreateServiceLinkedRole` for `rds.amazonaws.com` in this account.
## Provisioning access
### Create a CubeCloudBYOC policy
Navigate to **IAM->Policies** and create a new policy called `CubeCloudBYOC`
with the following JSON content. Please substitute `AWS_ACCOUNT_ID` with your
actual account ID.
```JSON theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"autoscaling:DescribeAutoScalingGroups",
"autoscaling:DescribeAutoScalingInstances",
"autoscaling:DescribeLaunchConfigurations",
"autoscaling:DescribeTags",
"ec2:DescribeAddresses",
"ec2:DescribeAddressesAttribute",
"ec2:DescribeAvailabilityZones",
"ec2:DescribeInternetGateways",
"ec2:DescribeLaunchTemplateVersions",
"ec2:DescribeLaunchTemplates",
"ec2:DescribeNatGateways",
"ec2:DescribeNetworkInterfaces",
"ec2:DescribePrefixLists",
"ec2:DescribeRegions",
"ec2:DescribeRouteTables",
"ec2:DescribeSecurityGroupRules",
"ec2:DescribeSecurityGroups",
"ec2:DescribeSubnets",
"ec2:DescribeVpcAttribute",
"ec2:DescribeVpcClassicLink",
"ec2:DescribeVpcClassicLinkDnsSupport",
"ec2:DescribeVpcEndpointServiceConfigurations",
"ec2:DescribeVpcEndpoints",
"ec2:DescribeVpcPeeringConnections",
"ec2:DescribeVpcs",
"ec2:DescribeVolumes",
"ec2:RunInstances",
"eks:DescribeCluster",
"eks:DescribeNodegroup",
"eks:ListClusters",
"iam:GetRole",
"sts:DecodeAuthorizationMessage",
"rds:DescribeDBInstances",
"rds:DescribeDBSubnetGroups",
"rds:DescribeDBParameterGroups",
"rds:ListTagsForResource"
],
"Resource": "*"
},
{
"Effect": "Allow",
"Action": ["s3:*"],
"Resource": ["arn:aws:s3:::cube-store-*"]
},
{
"Effect": "Allow",
"Action": [
"ec2:AllocateAddress",
"ec2:CreateInternetGateway",
"ec2:CreateLaunchTemplate",
"ec2:CreateNatGateway",
"ec2:CreateRoute",
"ec2:CreateRouteTable",
"ec2:CreateSecurityGroup",
"ec2:CreateSubnet",
"ec2:CreateTags",
"ec2:CreateVpc",
"ec2:CreateVpcEndpoint",
"ec2:CreateVpcEndpointServiceConfiguration",
"ec2:CreateVpcPeeringConnection",
"ec2:CreateVolume",
"eks:CreateCluster",
"eks:CreateNodegroup",
"iam:CreateOpenIDConnectProvider",
"iam:PassRole",
"iam:TagOpenIDConnectProvider",
"logs:CreateLogDelivery",
"kms:TagResource",
"kms:CreateKey",
"rds:CreateDBInstance",
"rds:CreateDBSubnetGroup",
"rds:AddTagsToResource"
],
"Resource": "*",
"Condition": {
"StringEquals": {
"aws:RequestTag/Created-By": "CubeCloud"
}
}
},
{
"Effect": "Allow",
"Action": [
"iam:AddRoleToInstanceProfile",
"iam:AttachRolePolicy",
"iam:CreateInstanceProfile",
"iam:CreateOpenIDConnectProvider",
"iam:CreatePolicy",
"iam:CreatePolicyVersion",
"iam:CreateRole",
"iam:CreateServiceLinkedRole",
"iam:DeleteInstanceProfile",
"iam:DeleteOpenIDConnectProvider",
"iam:DeletePolicy",
"iam:DeletePolicyVersion",
"iam:DeleteRole",
"iam:DeleteRolePolicy",
"iam:DeleteServiceLinkedRole",
"iam:DetachRolePolicy",
"iam:GetInstanceProfile",
"iam:GetOpenIDConnectProvider",
"iam:GetPolicy",
"iam:GetPolicyVersion",
"iam:GetRole",
"iam:GetRolePolicy",
"iam:ListAttachedRolePolicies",
"iam:ListInstanceProfilesForRole",
"iam:ListOpenIDConnectProviderTags",
"iam:ListPolicyVersions",
"iam:ListRolePolicies",
"iam:PassRole",
"iam:PutRolePolicy",
"iam:RemoveRoleFromInstanceProfile",
"iam:TagInstanceProfile",
"iam:TagOpenIDConnectProvider",
"iam:TagPolicy",
"iam:TagRole",
"iam:UpdateOpenIDConnectProviderThumbprint"
],
"Resource": [
"arn:aws:iam::{AWS_ACCOUNT_ID}:instance-profile/CubeCloud*",
"arn:aws:iam::{AWS_ACCOUNT_ID}:instance-profile/cubeapp-*",
"arn:aws:iam::{AWS_ACCOUNT_ID}:instance-profile/cube-store-*",
"arn:aws:iam::{AWS_ACCOUNT_ID}:oidc-provider/oidc.eks.*",
"arn:aws:iam::{AWS_ACCOUNT_ID}:policy/CubeCloud*",
"arn:aws:iam::{AWS_ACCOUNT_ID}:policy/cubeapp-*",
"arn:aws:iam::{AWS_ACCOUNT_ID}:role/CubeCloud*",
"arn:aws:iam::{AWS_ACCOUNT_ID}:role/cubeapp-*",
"arn:aws:iam::{AWS_ACCOUNT_ID}:role/cube-store-*"
],
"Condition": {
"StringEquals": {
"iam:ResourceTag/Created-By": "CubeCloud"
}
}
},
{
"Effect": "Allow",
"Action": "iam:CreateServiceLinkedRole",
"Resource": "*",
"Condition": {
"StringEquals": {
"iam:AWSServiceName": [
"eks.amazonaws.com",
"eks-nodegroup.amazonaws.com",
"eks-fargate.amazonaws.com",
"rds.amazonaws.com"
]
}
}
},
{
"Effect": "Allow",
"Action": [
"autoscaling:SetDesiredCapacity",
"autoscaling:TerminateInstanceInAutoScalingGroup"
],
"Resource": ["*"],
"Condition": {
"StringEquals": {
"autoscaling:ResourceTag/Created-By": "CubeCloud"
}
}
},
{
"Effect": "Allow",
"Action": ["eks:*", "kms:*"],
"Resource": ["*"],
"Condition": {
"StringEquals": {
"aws:ResourceTag/Created-By": "CubeCloud"
}
}
},
{
"Effect": "Allow",
"Action": ["ec2:*"],
"Resource": ["*"],
"Condition": {
"StringEquals": {
"ec2:ResourceTag/Created-By": "CubeCloud"
}
}
},
{
"Effect": "Allow",
"Action": [
"rds:ModifyDBInstance",
"rds:DeleteDBInstance",
"rds:DeleteDBSubnetGroup",
"rds:RebootDBInstance",
"rds:StartDBInstance",
"rds:StopDBInstance",
"rds:AddTagsToResource",
"rds:RemoveTagsFromResource"
],
"Resource": ["*"],
"Condition": {
"StringEquals": {
"aws:ResourceTag/Created-By": "CubeCloud"
}
}
}
]
}
```
### Creating a role
Navigate to **IAM->Roles** and create a new Role called `CubeCloudBYOC`. Select
**AWS Account** as the Trusted entity. Type and enter
`arn:aws:iam::307491255751:root`, which is the Cube BYOC provisioner
account. On the **Add permissions** page, find and select the `CubeCloudBYOC`
policy you created earlier. On the final **Review and create** page, edit the
**Trust Policy** to make it look like this.
```JSON theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"AWS": "arn:aws:iam::307491255751:root"
},
"Action": "sts:AssumeRole",
"Condition": {
"StringEquals": {
"sts:ExternalId": "cube-cloud-byoc"
}
}
}
]
}
```
Make sure to include `"sts:ExternalId": "cube-cloud-byoc"` in the Condition section.
## Deployment
The actual deployment will be done by Cube automation. All that's left to
do is notify your Cube contact point that access has been granted, and pass
along your Region/AWS Account ID information.
[aws-console]: https://console.aws.amazon.com/
[aws-docs-regions]: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-regions-availability-zones.html#concepts-available-regions
# Dedicated Infrastructure on AWS
Source: https://docs.cube.dev/admin/deployment/dedicated/aws/index
Connect Cube's Dedicated Infrastructure on AWS to your VPCs and corporate networks, or deploy the entire data plane inside your own AWS account via BYOC.
On AWS, Cube offers single-tenant Dedicated Infrastructure operated by Cube,
and Bring Your Own Cloud (BYOC) operated inside your own AWS account. Both
options support private network connectivity so that data source traffic — and,
optionally, Cube API traffic — never traverses the public internet.
## Backend connectivity (Cube → your network)
Use these options to give Cube private access to your data sources, auth
providers, BI APIs targeted by Semantic Layer Sync, and anything else Cube
needs to query. See [Backend and frontend connectivity][backend-frontend] for
the full picture.
* [**AWS PrivateLink**][aws-private-link] — let Cube reach data sources (your
own services, Snowflake, Databricks, etc.) over PrivateLink endpoint services
you expose.
* [**VPC Peering**][aws-vpc-peering] — establish a VPC peering connection
between the Cube VPC and your own VPC for access to internal data sources.
## Frontend connectivity (your clients → Cube)
Use this option to expose Cube's APIs to your applications, browsers, BI
tools, embedded analytics clients, and Semantic Layer Sync-generated configs
over a private network. Public endpoints can be disabled entirely on request.
* [**Private API Connectivity**][aws-private-api-connectivity] — expose Cube's
HTTP and SQL APIs over AWS PrivateLink.
## Bring Your Own Cloud
If you'd like the entire Cube data plane to live inside your own AWS account,
see [Bring Your Own Cloud on AWS][aws-byoc].
[aws-private-link]: /admin/deployment/dedicated/aws/private-link
[aws-vpc-peering]: /admin/deployment/dedicated/aws/vpc-peering
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[aws-byoc]: /admin/deployment/dedicated/aws/byoc
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Private API Connectivity on AWS
Source: https://docs.cube.dev/admin/deployment/dedicated/aws/private-api-connectivity
Expose Cube's HTTP and SQL APIs to your AWS network and internal BI tools over AWS PrivateLink so traffic from your applications, browsers, and BI clients never traverses the public internet.
This page covers **frontend connectivity** — exposing Cube's HTTP and SQL APIs
to your applications, browsers, BI tools, embedded analytics clients, and
Semantic Layer Sync-generated configs over a private network. For **backend
connectivity** (letting Cube reach into your network to query data sources,
auth providers, BI APIs targeted by SLS, and other upstream services), see
[AWS PrivateLink][aws-private-link] or [VPC Peering][aws-vpc-peering].
With Dedicated Infrastructure and Bring Your Own Cloud on AWS, Cube supports
establishing **AWS PrivateLink** connections from your AWS accounts to the Cube
API endpoints. This lets your applications, internal BI tools, and end-user
browsers reach the Cube HTTP, SQL, and AI APIs entirely over private AWS
networking — never touching the public internet. When private connectivity is
in place, the public API endpoints can be disabled completely on request.
Available on the [Enterprise plan](https://cube.dev/pricing) with Dedicated
Infrastructure or BYOC on AWS. [Contact us](https://cube.dev/contact) to enable
private API connectivity for your tenant.
## Architecture
Cube runs as two cooperating planes:
* The **control plane** powers the Cube UI and the product surface area
(deployment management, schema editor, dashboards, Semantic Layer Sync,
etc.) and is always served from Cube's public domain at
`https://.cubecloud.dev`. The control plane itself does not need
private connectivity — it is a SaaS UI like any other.
* The **data plane** runs your Cube deployments and serves all Cube HTTP and
SQL **data API** traffic — REST/GraphQL queries from your applications,
live data calls issued by the Cube UI while rendering charts, SQL
connections from BI tools, and AI traffic to the [Chat API][chat-api] and
other AI endpoints (including data-plane-hosted AI Engineer agents and
external agentic clients [calling the Chat API as a tool][agent-to-agent]).
The data plane is the part that talks to your databases and returns query
results, and the part this document teaches you how to expose privately.
[AWS PrivateLink][aws-privatelink] lets a service running in one VPC be
consumed from another VPC over the AWS internal network, without VPC peering
or routing through the internet. The service owner publishes a
[VPC Endpoint Service][aws-endpoint-service] in front of an internal load
balancer; the consumer creates a corresponding
[interface VPC Endpoint][aws-interface-endpoint] in their VPC, which appears
as a set of private ENIs that route traffic to the service over the AWS
backbone.
Cube exposes two VPC Endpoint Services per [Cube Region][cube-region] — one
fronting the HTTP (REST/GraphQL) data API and one fronting the SQL data API.
Each endpoint service sits in front of an internal
[Network Load Balancer][aws-nlb] (NLB) in the Cube VPC.
On the HTTP side, the NLB forwards traffic to Cube's ingress controller, which
terminates TLS using a certificate that only covers
`*..cubecloudapp.dev`. Because that certificate is bound to the
Cube-managed hostname, you have two options for presenting HTTPS to your own
clients (covered in detail [below](#tls-options-for-https)):
* **Re-terminate TLS on your side** with your own certificate for the hostname
your clients will dial (`http.cube.internal`, `cube.example.com`, …),
forwarding upstream to the HTTP VPC endpoint.
* **Reuse the Cube hostname privately** by setting up an internal DNS override
for `.cubecloudapp.dev` that points at the VPC endpoint. Cube's
certificate is then valid for all clients without any additional cert work
on your side.
On the SQL side, TLS is terminated inside the SQL API service behind the NLB,
and SQL clients connect using `sslmode=require` (or equivalent) directly to
the VPC endpoint hostname.
On your side, you create a **VPC Endpoint** (interface endpoint) in your VPC
that connects to each Cube endpoint service, and bind a DNS name to it that
resolves to the endpoint's private IPs from inside your VPN-routable network.
Your applications and BI tools then connect to that private hostname exactly
as they would to a public Cube endpoint, except the traffic flows through
PrivateLink instead of the internet.
## Cube Region
The endpoint services described here are scoped to a **Cube Region** — the
unit of infrastructure that hosts one or more of your deployments. Each region
has a stable identifier of the form `--`
(e.g. `aws-us-east-1-t-12345-prod` for a single-tenant Dedicated region, or
`aws-us-east-1-t-12345-byoc` for a BYOC region). The hostname you override on
your side — either privately remapping `.cubecloudapp.dev` or
fronting it with your own domain — applies to **every** deployment in that
region. See [Cube Regions][cube-region] for the full reference, including
how to find the exact region identifier for your tenant.
## How queries are routed
A single private endpoint per region serves all deployments inside that
region; there is no per-deployment subdomain or per-deployment endpoint to
provision separately.
On Cube's shared, public-facing infrastructure, traffic is routed to a
deployment by **subdomain** — `..cubecloudapp.dev`
maps to a specific deployment. When private connectivity is enabled, Cube
switches to **path-based routing** so that a single private hostname can serve
every deployment in the region.
For the HTTP API, the prefix has the shape:
```
https:///deployment//cubejs-api/v1/...
```
where `` is the leftmost label of the deployment's Cube-issued
hostname (e.g. `thirsty-raccoon` from
`thirsty-raccoon.aws-us-east-1-t-12345-prod.cubecloudapp.dev`). The same prefix
applies to all HTTP endpoints — `/cubejs-api/v1/load`, `/livez`, the GraphQL
endpoint, and so on:
```bash theme={"dark"}
curl https://http.cube.internal/deployment/thirsty-raccoon/cubejs-api/v1/load \
-H "Authorization: $CUBE_JWT" \
-d '{"query":{"measures":["orders.count"]}}'
```
The SQL API uses a TCP protocol that does not carry an HTTP path, so SQL
connections are routed at the protocol layer (via the database name and
credentials in the connection string) rather than by URL prefix. SQL clients
connect to the same private hostname on port 5432 and identify the deployment
through the SQL connection parameters.
## TLS options for HTTPS
Cube's ingress controller can terminate TLS only for the Cube-issued domain
`*..cubecloudapp.dev`. Pick the option that fits your DNS and
certificate posture:
### Option A — Reuse the Cube hostname via private DNS
Create a private DNS zone for `.cubecloudapp.dev` —
typically a [Route 53 private hosted zone][aws-route53-phz] associated with
your VPC, or an equivalent override in your corporate resolver — and point the
following records at the VPC endpoints:
| Record name | Type | Target |
| -------------------------------------- | ---- | ------------------------------ |
| `.cubecloudapp.dev` | A | Alias to the HTTP VPC endpoint |
| `*..cubecloudapp.dev` | A | Alias to the HTTP VPC endpoint |
| `sql..cubecloudapp.dev` | A | Alias to the SQL VPC endpoint |
| `*.sql..cubecloudapp.dev` | A | Alias to the SQL VPC endpoint |
Clients then dial:
```
https://.cubecloudapp.dev/deployment//cubejs-api/v1/...
```
TLS is terminated by Cube's ingress controller using Cube's own certificate
for `*.cubecloudapp.dev` — you don't need to manage a certificate yourself.
This is the simplest setup if you control DNS resolution on the networks your
clients live on. Configure the same hostname in Cube's admin interface so
that the UI and Semantic Layer Sync generate links and configs against it.
### Option B — Front the endpoint with your own domain and certificate
If you'd rather present clients a hostname inside your own domain (e.g.
`cube.example.com`), stand up a customer-side proxy — typically an internal
NLB with a TLS listener bound to a certificate you own, or any reverse proxy
such as nginx, HAProxy, Envoy, or an ALB — that terminates TLS with your
certificate and **re-encrypts** traffic to the Cube HTTP VPC endpoint. Because
Cube's ingress only has a cert for `*.cubecloudapp.dev`, the proxy must do
its own TLS termination; it cannot pass-through your custom-domain TLS to
Cube.
The rest of this section walks each step with both **AWS Console** and
**AWS CLI** instructions. The CLI commands use AWS CLI v2 — run them with
credentials for the account that holds the VPC endpoint, and set your region
with `--region` or the `AWS_REGION` environment variable. If you front the
endpoint with a different reverse proxy (nginx, Envoy, HAProxy, an ALB) instead
of an NLB, the same two principles apply: terminate your own certificate at the
proxy, and re-encrypt to the endpoint over TLS (HTTPS) on port 443 rather than
forwarding plaintext.
Create a target group for the HTTP VPC endpoint with these settings:
* **Target type**: **IP addresses**. A target group cannot reference a VPC
endpoint service name directly, so you register the endpoint's private
IPs instead — one ENI per Availability Zone, so register the address from
every AZ. These IPs are stable for the lifetime of the endpoint.
* **Protocol / port**: **TLS / 443**. This must be TLS, not TCP — see the
callout below.
* **VPC**: the same VPC that holds your HTTP VPC endpoint.
* **Health check**: **TCP** on port 443 (a TLS health check works too).
**Console** — under **EC2 → Target groups → Create target group**, choose
the settings above, then register the endpoint ENI IPs on the **Register
targets** screen. Find the ENI IPs on the endpoint's **VPC → Endpoints →
(your endpoint) → Subnets** view, or via the CLI below.
**AWS CLI**:
```bash theme={"dark"}
# Look up the endpoint's ENI private IPs (one per AZ)
ENIS=$(aws ec2 describe-vpc-endpoints --vpc-endpoint-ids \
--query 'VpcEndpoints[0].NetworkInterfaceIds' --output text)
aws ec2 describe-network-interfaces --network-interface-ids $ENIS \
--query 'NetworkInterfaces[].PrivateIpAddress' --output text
# → one private IP per Availability Zone, e.g. 10.0.1.15 10.0.2.15
TG_ARN=$(aws elbv2 create-target-group \
--name cube-http-tg \
--protocol TLS --port 443 \
--vpc-id \
--target-type ip \
--health-check-protocol TCP --health-check-port 443 \
--query 'TargetGroups[0].TargetGroupArn' --output text)
aws elbv2 register-targets --target-group-arn "$TG_ARN" \
--targets Id=10.0.1.15,Port=443 Id=10.0.2.15,Port=443
```
Create an internal NLB with these settings:
* **Scheme**: **internal**.
* **VPC / subnets**: the same VPC, with subnets in the same Availability
Zones as the endpoint's ENIs (see [Availability Zone
alignment](#availability-zone-alignment)). Enable **cross-zone load
balancing** if you want an NLB node in one AZ to reach endpoint ENIs in
another.
* **Listener**: protocol **TLS**, port **443**, forwarding to the target
group from the previous step.
* **Default certificate**: your own [ACM certificate][aws-acm] covering the
hostname your clients will dial (e.g. `cube.example.com`).
* **Security policy**: a current TLS 1.2+ policy such as
`ELBSecurityPolicy-TLS13-1-2-2021-06`.
Also allow inbound **TCP/443** from the NLB's subnets on the HTTP VPC
endpoint's security group so the re-encrypted traffic can reach the
endpoint ENIs.
**Console** — under **EC2 → Load balancers → Create load balancer →
Network Load Balancer**, apply the settings above and add the **TLS**
listener; then add the endpoint security-group rule under **VPC → Security
groups**.
**AWS CLI**:
```bash theme={"dark"}
NLB_ARN=$(aws elbv2 create-load-balancer \
--name cube-http-nlb \
--type network --scheme internal \
--subnets \
--query 'LoadBalancers[0].LoadBalancerArn' --output text)
aws elbv2 wait load-balancer-available --load-balancer-arns "$NLB_ARN"
# Optional: let an NLB node reach endpoint ENIs in another AZ
aws elbv2 modify-load-balancer-attributes --load-balancer-arn "$NLB_ARN" \
--attributes Key=load_balancing.cross_zone.enabled,Value=true
# TLS listener on 443 with your ACM certificate, forwarding to the TG
aws elbv2 create-listener --load-balancer-arn "$NLB_ARN" \
--protocol TLS --port 443 \
--certificates CertificateArn= \
--ssl-policy ELBSecurityPolicy-TLS13-1-2-2021-06 \
--default-actions Type=forward,TargetGroupArn="$TG_ARN"
# Allow re-encrypted traffic from the NLB subnets into the endpoint
aws ec2 authorize-security-group-ingress --group-id \
--protocol tcp --port 443 --cidr
```
The NLB receives an internal DNS name of the form
`-.elb..amazonaws.com`. Point your chosen hostname
(e.g. `cube.example.com`) at it with an `ALIAS`/`CNAME` record in a
[Route 53 private hosted zone][aws-route53-phz] or your corporate resolver,
resolvable from every network that reaches Cube (corporate VPN included).
**Console** — copy the NLB's DNS name from its detail page under **EC2 →
Load balancers**, then create the `ALIAS`/`CNAME` record under **Route 53 →
Hosted zones** (or in your corporate DNS).
**AWS CLI** (Route 53 private hosted zone):
```bash theme={"dark"}
NLB_DNS=$(aws elbv2 describe-load-balancers --load-balancer-arns "$NLB_ARN" \
--query 'LoadBalancers[0].DNSName' --output text)
aws route53 change-resource-record-sets --hosted-zone-id \
--change-batch "{
\"Changes\": [{
\"Action\": \"UPSERT\",
\"ResourceRecordSet\": {
\"Name\": \"cube.example.com\",
\"Type\": \"CNAME\",
\"TTL\": 60,
\"ResourceRecords\": [{ \"Value\": \"$NLB_DNS\" }]
}
}]
}"
```
Then set the same hostname as the domain override in Cube's admin panel so
generated UI links and SLS configs use it — see
[Configuring the private domain](#configuring-the-private-domain).
The target group's protocol must be **TLS**, not **TCP**. Cube's HTTP endpoint
expects a TLS handshake on port 443, so the NLB has to **re-encrypt** on the
upstream leg: with a [TLS target group][aws-nlb-tls] it terminates your
listener's TLS and opens a fresh TLS connection to the endpoint, whereas a TCP
target group would forward the already-decrypted bytes to a port that expects
TLS — so requests never complete. AWS does **not** validate the backend
certificate on a TLS target group, so Cube's `*.cubecloudapp.dev` certificate
is accepted even though it does not match your custom hostname; it is used
purely to encrypt the upstream hop, keeping traffic encrypted end-to-end across
PrivateLink.
## Reaching Cube from every client
The same private hostname is used by three distinct classes of client, and
each imposes a requirement on how the name resolves:
1. **Your application servers** inside the same VPC as the VPC endpoint
resolve the private hostname via the VPC's DNS resolver and connect over
PrivateLink directly.
2. **End-user browsers loading the Cube UI.** When a developer opens
`https://.cubecloud.dev` and renders a dashboard, the page issues
live queries against the Cube data API from the user's browser. Those
calls originate on the user's laptop, not from a Cube backend, so the URL
embedded in the UI must be a hostname the user's machine can resolve and
reach. For this to use PrivateLink, the hostname has to resolve to the
VPC endpoint over your corporate VPN's DNS and route over the VPN to your
VPC. A purely internal name that only resolves inside the VPC will fail
when the user is on the corporate network but not in the VPC; a public
name that resolves outside the VPN will bypass PrivateLink entirely.
3. **Internal BI tools configured by Semantic Layer Sync (SLS).** If you
publish Cube datasets to BI tools via SLS, Cube generates connection
configs that embed the Cube API hostname. Those configs are used by BI
desktop clients, BI gateways, or other clients running inside your
network — so SLS must be configured with a hostname those clients can
resolve and reach over your VPN.
In short, the **private hostname must be visible and routable from every
place the Cube API needs to be reached** — your VPCs, your corporate
VPN-connected laptops, and any BI gateway hosts. Cube cannot infer your
internal DNS; you tell us the hostname you want to use, and you point it at
the VPC endpoint on your side.
If a user opens the Cube UI and sees a **"Network Error"** banner while
loading charts, dashboards, or other data-driven views, it means the
browser cannot reach the data plane over your private network. See
[Troubleshooting "Network Error" in the Cube UI](#troubleshooting-network-error-in-the-cube-ui)
below for steps to isolate the failure.
## Chrome local network access permission
Chromium-based browsers block requests from a page served over the public
internet to a host that resolves to a **private** address, unless the site has
the **Local network access** permission. The Cube UI is served from the public
`https://.cubecloud.dev` origin while your private API hostname
resolves to an RFC 1918 address over the VPN — so every browser that loads the
Cube UI needs this permission before charts, the Semantic Model IDE, the
playground, and any other data-driven view will render.
This is a browser policy that Cube cannot grant on your behalf, and it applies
to both TLS options above: the browser decides from the **resolved IP address**,
not the hostname, so reusing `.cubecloudapp.dev` does not avoid it.
Non-browser clients — `curl`, `psql`, your application servers, BI gateways — are
unaffected, which is why `curl` from a laptop can succeed while the Cube UI on
the same laptop shows **"Network Error"**.
Chrome enforces this as [Local Network Access][chrome-lna] from Chrome 142
onwards; earlier versions enforced its predecessor,
[Private Network Access][chrome-pna]. A blocked call appears in DevTools as a
failed request with `ERR_BLOCKED_BY_PRIVATE_NETWORK_ACCESS_CHECKS`, or as a
CORS-style console error naming a private or local target address space.
**Individual users** can accept Chrome's *"wants to look for and connect to any
device on your local network"* prompt, or set **Local network access** to
**Allow** under the padlock icon → **Site settings** on the Cube UI origin.
**Managed fleets** should pre-grant the permission with a Chrome Enterprise
policy, since the prompt is easy to dismiss and reappears per profile. Set the
policy value to your Cube UI origin (e.g. `https://.cubecloud.dev`):
| Chrome version | Policy |
| -------------- | ------------------------------------------------------------------- |
| 142 and later | [`LocalNetworkAccessAllowedForUrls`][chrome-policy-lna] |
| Before 142 | [`InsecurePrivateNetworkRequestsAllowedForUrls`][chrome-policy-pna] |
If some users still fail after the policy is rolled out, have them check
`chrome://policy` and confirm the policy is listed with the expected value. A
missing entry means it has not synced to that device or profile — most often
because the user is signed in to Chrome with a personal, unmanaged Google
profile, which enterprise policies do not reach. That is the usual explanation
for the same page working for one colleague and failing for another on the same
VPN.
## Configuring the private domain
Configure the private hostname for each Cube Region from the Cube admin panel,
under **Regions → \ → Path-based routing → Domain override**.
The domain override tells Cube **where the data plane lives from the
caller's point of view**. It does *not* re-route traffic on Cube's side and
it does *not* open a new tunnel — it is purely a string Cube embeds when
generating outbound references to the data API. Specifically, Cube uses the
override value when generating:
* Links inside the Cube UI (used by **browsers** when rendering dashboards
and the schema editor — every Cube API call the UI issues goes to this
hostname directly from the user's machine).
* Connection configs generated by Semantic Layer Sync (used by BI
desktop clients and gateways inside your network).
* API connection details surfaced to your applications on the
deployment's API page.
Because every one of those callers — the browser, the BI client, your
application — connects to the hostname **directly**, the override only
works if that hostname is resolvable and reachable from each of those
networks. Setting the override does not give Cube's control plane a new
way to talk to your data plane; it tells *your* clients to stop using the
public Cube domain and start using the private one *you* have published.
Use the hostname your clients will actually dial — the custom-domain front-end
in Option B (e.g. `http.cube.internal`), or the Cube-issued
`.cubecloudapp.dev` if you went with Option A. Leaving the field
empty falls back to the default public hostname, and clients will not use
PrivateLink even if the endpoint service is connected.
HTTP and SQL APIs are typically published under two separate hostnames
(`http.cube.internal` and `sql.cube.internal` in the diagram above), since
the HTTP endpoint sits behind your TLS-terminating proxy on port 443 while
the SQL endpoint is accessed directly on port 5432.
## Provisioning checklist
1. **Request PrivateLink for your tenant.** Contact Cube and provide your
tenant name and the AWS account ID(s) that will create the VPC endpoints.
2. **Receive endpoint service names.** Cube provisions one HTTP and one SQL
VPC Endpoint Service per region and shares the service names
(`com.amazonaws.vpce..vpce-svc-…`) along with the Cube Region
identifier.
3. **Create VPC endpoints in your account.** In the AWS Console under
**VPC → Endpoints**, create an [interface endpoint][aws-interface-endpoint]
for each service. Place the endpoint in subnets that are reachable from
both your application workloads and your corporate VPN, and attach a
security group that allows inbound traffic on 443 (HTTP) or 5432 (SQL)
from those clients.
4. **Accept the connection.** Cube accepts the endpoint connection request
on the provider side once visibility is confirmed.
5. **Bind the private hostname.** Pick a TLS option above and create the
corresponding DNS records — either a [Route 53 private hosted
zone][aws-route53-phz] for `.cubecloudapp.dev` (Option A) or
a single record for your custom hostname pointing at your proxy
(Option B). Make sure the records resolve from every network that needs
to reach Cube, corporate VPN included.
6. **Set the domain in Cube.** In the admin panel under
**Regions → \ → Path-based routing → Domain override**,
enter the chosen hostname so that Cube generates UI links and SLS configs
using that name.
7. **(Optional) Disable public endpoints.** Once private connectivity is
verified end-to-end, ask Cube support to disable the public HTTP and SQL
endpoints for your deployment.
## Availability Zone alignment
Cube's dedicated infrastructure in each AWS region is deployed across a
**limited, fixed set of Availability Zones**.
[AWS PrivateLink requires][aws-privatelink-az] the consumer's VPC endpoint
to have a subnet in **at least one of the same AZs** where the provider's
endpoint service is exposed. If none of the subnets in your VPC live in a
Cube-supported AZ for that region, the VPC endpoint will fail to attach.
The remedy is to create an additional subnet in one of the Cube-supported
AZs inside your VPC, attach the VPC endpoint to that subnet, and route
traffic from your other AZs to it. AWS-internal AZ IDs (e.g. `use1-az2`)
differ per-account, so coordinate with the Cube team to confirm the
correct AZ IDs for your region before adding the subnet.
## Verifying connectivity
From a host inside the consumer VPC or attached to the corporate VPN:
```bash theme={"dark"}
# DNS resolves the HTTP and SQL names to their VPC endpoints' private IPs
dig +short http.cube.internal
dig +short sql.cube.internal
# HTTPS reachable over PrivateLink, routed to a specific deployment by path prefix
curl -v https://http.cube.internal/deployment//livez
# SQL API reachable
psql "host=sql.cube.internal port=5432 user=… dbname=… sslmode=require" -c "select 1;"
```
If DNS resolves but connections hang, check the VPC endpoint state, security
group rules on the endpoint ENIs, and that your VPN's route tables include
the VPC's CIDR.
## Troubleshooting "Network Error" in the Cube UI
When the domain override is configured, the Cube UI stops calling the
public Cube data API and starts calling the private hostname you supplied
**directly from the user's browser**. If that hostname is not reachable
from the machine the UI is loaded on, the UI surfaces a red **"Network
Error"** banner on charts, the schema editor, the playground, and any
other view that issues data queries — even though direct API calls (e.g.
`curl` from inside the VPC, or your application servers) keep working.
The fix is usually somewhere along the network path between the browser
and the VPC endpoint. Use the browser's developer tools to pinpoint which
hop is failing:
Reload the page that shows the **"Network Error"** banner with the
Network tab open. Filter for `Fetch/XHR` requests and look for the
failing call — it will be a request to the private hostname you
configured as the domain override (for example
`https://http.cube.internal/deployment//cubejs-api/v1/load`),
not to `*.cubecloud.dev`.
If the failing request still points at the **public** `*.cubecloud.dev`
hostname, the domain override has not been picked up yet — re-check
the value in **Regions → \ → Path-based routing → Domain
override** and reload the UI with cache disabled.
Right-click the failing request and copy its full URL, then open it
in a new browser tab. This strips the Cube UI out of the picture so
you can see what the browser itself sees when talking to the private
hostname.
Common outcomes:
* **The page does not load at all / "This site can't be reached" /
`ERR_NAME_NOT_RESOLVED`** — DNS for the private hostname is not
reachable from the user's machine. Confirm the user is on the
corporate VPN, that the VPN pushes the private DNS zone (or that
the corporate resolver answers for the override hostname), and
that the record points at the VPC endpoint's private IPs.
* **`ERR_CONNECTION_TIMED_OUT` / `ERR_CONNECTION_REFUSED`** — DNS
resolves but TCP to the VPC endpoint is blocked. Check that the
VPN route table covers the VPC CIDR and that the security group on
the VPC endpoint ENIs allows 443 from the VPN's client CIDR.
* **`NET::ERR_CERT_COMMON_NAME_INVALID` /
`NET::ERR_CERT_AUTHORITY_INVALID` / browser interstitial about an
untrusted certificate** — TLS termination is misconfigured. With
Option A (reuse the Cube hostname), make sure the private DNS
record really is for `.cubecloudapp.dev` so Cube's
`*.cubecloudapp.dev` certificate matches. With Option B (custom
domain), make sure your proxy is presenting a certificate whose
CN/SAN matches the override hostname and is signed by a CA the
user's machine trusts.
* **`ERR_BLOCKED_BY_PRIVATE_NETWORK_ACCESS_CHECKS`, or the URL
loads fine in its own tab but still fails from the Cube UI** —
the browser is blocking the call because the private hostname
resolves to a private address and the site lacks Chrome's
**Local network access** permission. See
[Chrome local network access permission](#chrome-local-network-access-permission).
* **HTTP 4xx/5xx from Cube** — the network path is fine; the error
is on the API itself. Inspect the response body in the new tab to
see Cube's error message and proceed as you would for any other
API error.
From a host inside the consumer VPC, run the same request with
`curl -v` (see [Verifying connectivity](#verifying-connectivity)
above). If it succeeds from the VPC but fails from the user's
laptop, the data plane is healthy and the problem is strictly in the
VPN → VPC endpoint path for that user.
If every request from every machine fails, re-open
**Regions → \ → Path-based routing → Domain override**
and confirm:
* The hostname has no scheme prefix (no `https://`) and no path.
* The hostname matches the DNS record you actually published.
* For Option B, the hostname matches the certificate on your TLS
proxy.
* Clearing the field falls back to the public hostname — useful as
a quick smoke test to confirm the UI itself is healthy and the
problem is the private path.
[cube-region]: /admin/deployment/infrastructure#understanding-cube-cloud-region
[aws-private-link]: /admin/deployment/dedicated/aws/private-link
[aws-vpc-peering]: /admin/deployment/dedicated/aws/vpc-peering
[chat-api]: /reference/embed-apis/chat-api
[agent-to-agent]: /recipes/ai/agent-to-agent
[aws-privatelink]: https://docs.aws.amazon.com/vpc/latest/privatelink/what-is-privatelink.html
[aws-endpoint-service]: https://docs.aws.amazon.com/vpc/latest/privatelink/configure-endpoint-service.html
[aws-interface-endpoint]: https://docs.aws.amazon.com/vpc/latest/privatelink/create-interface-endpoint.html
[aws-nlb]: https://docs.aws.amazon.com/elasticloadbalancing/latest/network/introduction.html
[aws-nlb-tls]: https://docs.aws.amazon.com/elasticloadbalancing/latest/network/create-tls-listener.html
[aws-acm]: https://docs.aws.amazon.com/acm/latest/userguide/acm-overview.html
[aws-privatelink-az]: https://docs.aws.amazon.com/vpc/latest/privatelink/configure-endpoint-service.html#endpoint-service-availability-zones
[aws-route53-phz]: https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/hosted-zones-private.html
[chrome-lna]: https://developer.chrome.com/blog/local-network-access
[chrome-pna]: https://developer.chrome.com/blog/private-network-access-update
[chrome-policy-lna]: https://chromeenterprise.google/policies/local-network-access-allowed-for-urls/
[chrome-policy-pna]: https://chromeenterprise.google/policies/insecure-private-network-requests-allowed-for-urls/
# Setting up AWS PrivateLink
Source: https://docs.cube.dev/admin/deployment/dedicated/aws/private-link
How to expose an AWS endpoint service and coordinate PrivateLink so Cube's Dedicated Infrastructure reaches your VPC privately.
This page covers **backend connectivity** — Cube reaching into your network to
query data sources, auth providers, BI APIs targeted by Semantic Layer Sync,
and other upstream services. See
[Backend and frontend connectivity][backend-frontend] for the full picture.
For **frontend connectivity** (exposing Cube's APIs to your applications,
browsers, BI tools, and embedded analytics clients), see
[Private API Connectivity on AWS][aws-private-api-connectivity].
[AWS PrivateLink][aws-docs-private-link] provides private connectivity between
virtual private clouds (VPCs), supported services and resources, and your
on-premises networks, without exposing your traffic to the public internet.
To set up a PrivateLink connection between Cube's Dedicated Infrastructure
and your own VPC, you'll need to prepare an Endpoint Service, share service
details with the Cube team, and accept the incoming connection request.
**Dedicated Infrastructure vs. Bring Your Own Cloud.** The flow described on
this page — sharing service details with the Cube team and letting Cube create
the VPC endpoint and DNS overrides — applies to
[Dedicated Infrastructure][cube-region] operated by Cube.
In a [Bring Your Own Cloud (BYOC)][aws-byoc] deployment, the Cube VPC lives in
**your own AWS account**, so you own the networking. The IAM role granted to
the Cube Operator intentionally does not include `route53:*` permissions,
which means Cube cannot create the VPC interface endpoint or the Route 53
private hosted zone needed for the DNS override on your behalf.
For BYOC, set up PrivateLink yourself inside the Cube VPC:
1. Create the VPC interface endpoint in the Cube VPC against the provider's
Endpoint Service Name.
2. Create a Route 53 private hosted zone for the TLS hostname and associate
it with the Cube VPC, with an `A` ALIAS record pointing at the interface
endpoint.
3. Confirm the security groups on both ends allow the required ports.
If you'd prefer Cube to do this for you in BYOC, you can grant the Cube
Operator role `route53:*` (and the matching `ec2:*VpcEndpoint*` permissions)
on the BYOC role — but most customers keep this networking in their own
hands.
## Preparing the Endpoint Service
There are two common scenarios for preparing the Endpoint Service:
* Connecting to a service in your AWS infrastructure
* Connecting to a service provided by a third party such as Snowflake,
Databricks, Altinity Cloud, etc.
In the case of your own infrastructure, please follow the
[official AWS documentation][aws-docs-endpoint-service] to configure the
Endpoint Service pointing at your data source.
If your data source is hosted in a third-party infrastructure, please follow
the vendor's documentation for creating and managing an Endpoint Service.
## Allowing the Cube principal
Cube needs to be added to the list of principals allowed to discover your
Endpoint Service. To do so, please go to **AWS Console** → **VPC** →
**Endpoint Services** → **Your service** → **Allow principals** and add
`arn:aws:iam::331376342520:root` to the list.
`331376342520` is the AWS account ID of Cube's PrivateLink consumer account.
Adding its root principal authorizes Cube to discover your endpoint service
and create a private endpoint against it; nothing else in Cube's AWS estate
gains access to your network.
## Gathering required information
To request establishing a PrivateLink connection, please share the following
information with the Cube team:
* **Service Name** (such as `com.amazonaws.vpce.us-west-2.vpce-svc-abcde`)
* **Reference Name** for the record (such as "Snowflake-prod" or
"clickhouse-dev")
* **Ports**: a list of ports that will be accessed through this connection
* **DNS Name(s)**: see [DNS and TLS](#dns-and-tls) below
* **Cube Region:** PrivateLink requires Cube to be hosted on
[Dedicated Infrastructure][cube-region]. Specify which Cube Region should
host your Dedicated Infrastructure.
## DNS and TLS
How your data source is addressed inside Cube depends on whether it speaks
TLS:
* **If the service uses TLS** (HTTPS, JDBC `sslmode=require`, etc.), share
the **DNS name(s)** the certificate is issued for — typically the same
hostname your in-network clients already use to reach it. Cube creates
internal DNS overrides inside the Dedicated Infrastructure so that the same
hostname resolves to the PrivateLink endpoint. Keeping the original
hostname is what preserves TLS validity: the certificate's CN/SAN keeps
matching what Cube dials.
* **If the service does not use TLS** and you don't supply a DNS name, the
Cube team will share back an internal endpoint hostname (e.g. an
AWS-assigned interface-endpoint DNS name) that you can configure as the
upstream when you wire the connection into Cube.
## Accepting the connection
The Cube team will notify you once the connection request is sent. You can
accept it by going to **AWS Console** → **VPC** → **Endpoint Services** →
**Your Service** → **Endpoint Connections** and clicking
**Accept Connection Request**.
## Using the connection
Once the connection is established, you can access your data source by
addressing it via the DNS name(s) you supplied (TLS case) or the internal
endpoint hostname returned to you by the Cube team (non-TLS case).
## Supported Regions
AWS PrivateLink is available in all AWS commercial regions where Dedicated
Infrastructure can be provisioned. AWS China (`cn-north-1`, `cn-northwest-1`)
and AWS GovCloud (`us-gov-east-1`, `us-gov-west-1`) are not supported.
[aws-docs-private-link]: https://aws.amazon.com/privatelink/
[aws-docs-endpoint-service]: https://docs.aws.amazon.com/vpc/latest/privatelink/configure-endpoint-service.html
[cube-region]: /admin/deployment/infrastructure#understanding-cube-cloud-region
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[aws-byoc]: /admin/deployment/dedicated/aws/byoc
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Setting up VPC Peering on AWS
Source: https://docs.cube.dev/admin/deployment/dedicated/aws/vpc-peering
End-to-end checklist for VPC peering Cube's Dedicated Infrastructure with your AWS VPC for private data access.
This page covers **backend connectivity** — Cube reaching into your network to
query data sources, auth providers, BI APIs targeted by Semantic Layer Sync,
and other upstream services. See
[Backend and frontend connectivity][backend-frontend] for the full picture.
For **frontend connectivity** (exposing Cube's APIs to your applications,
browsers, BI tools, and embedded analytics clients), see
[Private API Connectivity on AWS][aws-private-api-connectivity].
To set up AWS VPC Peering between Cube's Dedicated Infrastructure and your
VPC, you collect the information below and hand it over to your Cube
representative. Next, you accept a VPC peering request initiated by Cube, then
configure security groups and route tables so that Cube can reach your data
source.
## Information required by Cube
To allow Cube to peer with a [VPC on AWS][aws-docs-vpc], please share the
following with the Cube team:
* **AWS Account ID:** The AWS account ID of the VPC owner. This can be found
in the top-right corner of [the AWS Console][aws-console].
* **AWS Region:** [The AWS region][aws-docs-regions] that the VPC resides in.
* **AWS VPC ID:** The ID of the VPC that Cube will connect to, for example
`vpc-0099aazz`.
* **AWS VPC CIDR:** The [CIDR block][wiki-cidr-block] of the VPC that Cube
will connect to, for example `10.0.0.0/16`.
* **Cube Region:** VPC Peering requires Cube to be hosted on
[Dedicated Infrastructure][cube-region]. Specify which Cube Region should
host your Dedicated Infrastructure.
## Setup
### Accepting the peering request
After receiving the information above, Cube will send a
[VPC peering request][aws-docs-vpc-peering] that must be accepted. This can
be done either through the [AWS Web Console][aws-console] or through an
infrastructure-as-code tool.
To [accept the VPC peering request][aws-docs-vpc-peering-accept] through the
AWS Web Console, follow the instructions below:
1. Open the [Amazon VPC console](https://console.aws.amazon.com/vpc/).
Ensure you have the necessary permissions to accept a VPC peering
request. If you are unsure, please contact your AWS administrator.
2. Use the Region selector to choose the Region of the accepter VPC.
3. In the navigation pane, choose **Peering connections**.
4. Select the pending VPC peering connection (the status should be
`pending-acceptance`), then choose **Actions**, followed by
**Accept request**.
Ensure the peering request is from Cube by checking that the **AWS
account ID**, **region**, and **VPC IDs** match those provided by your
CSM.
5. When prompted for confirmation, choose **Accept request**.
6. Choose **Modify my route tables now** to add a route to the VPC route
table so that you can send and receive traffic across the peering
connection.
For more information about peering connection lifecycle statuses, check out
the [VPC peering connection lifecycle on AWS][aws-docs-vpc-peering-lifecycle].
### Updating security groups
The initial VPC setup will not allow traffic from Cube; this is because
[the security group][aws-docs-vpc-security-group] for the database will need
to allow access from the Cube VPC CIDR block.
This can be achieved by adding a new security group rule:
| Protocol | Port Range | Source/Destination |
| -------- | ---------- | ------------------------------------------- |
| TCP | 3306 | The Cube VPC CIDR block for the AWS region. |
The Cube VPC CIDR block is shared with you by the Cube team alongside the
peering request, and is also visible in the AWS Console on the **Peering
connections** → **\** → **Details** page as the
**Requester VPC CIDR**.
### Updating route tables
The final step is to update route tables in your VPC to allow traffic from
Cube to reach your database. The Cube VPC CIDR block must be added to the
route tables of all subnets that connect to the database. To do this, follow
the instructions on [the AWS documentation][aws-docs-vpc-peering-routing].
## Troubleshooting
Database connection issues with misconfigured VPCs often manifest as
connection timeouts. If you are experiencing connection issues, please check
the following:
* Verify that
[all security groups allow traffic](#updating-security-groups) from the
Cube VPC CIDR block.
* Verify that
[a route exists to the Cube VPC CIDR block](#updating-route-tables) from
the subnets that connect to the database.
## Supported Regions
VPC Peering is available in all AWS commercial regions where Dedicated
Infrastructure can be provisioned. AWS China (`cn-north-1`, `cn-northwest-1`)
and AWS GovCloud (`us-gov-east-1`, `us-gov-west-1`) are not supported.
[aws-console]: https://console.aws.amazon.com/
[aws-docs-regions]: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-regions-availability-zones.html#concepts-available-regions
[aws-docs-vpc]: https://docs.aws.amazon.com/vpc/latest/userguide/what-is-amazon-vpc.html
[aws-docs-vpc-peering]: https://docs.aws.amazon.com/vpc/latest/peering/what-is-vpc-peering.html
[aws-docs-vpc-peering-accept]: https://docs.aws.amazon.com/vpc/latest/peering/create-vpc-peering-connection.html#different-account-different-region
[aws-docs-vpc-peering-lifecycle]: https://docs.aws.amazon.com/vpc/latest/peering/vpc-peering-basics.html#vpc-peering-lifecycle
[aws-docs-vpc-peering-routing]: https://docs.aws.amazon.com/vpc/latest/peering/vpc-peering-routing.html
[aws-docs-vpc-security-group]: https://docs.aws.amazon.com/vpc/latest/userguide/security-group-rules.html
[wiki-cidr-block]: https://en.wikipedia.org/wiki/Classless_Inter-Domain_Routing#CIDR_blocks
[cube-region]: /admin/deployment/infrastructure#understanding-cube-cloud-region
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Bring Your Own Cloud on Azure
Source: https://docs.cube.dev/admin/deployment/dedicated/azure/byoc
Azure BYOC architecture, provisioner app access, and AKS-based onboarding for deploying Cube inside your own Azure subscription.
With Bring Your Own Cloud (BYOC) on Azure, all the components interacting with private data are deployed on
the customer infrastructure on Azure and managed by the Cube Control Plane via the Cube Operator.
This document provides step-by-step instructions for deploying Cube BYOC on Azure.
## Overall Design
Cube will gain access to your Azure account via the Cube Provisioner Enterprise App.
It will leverage a dedicated subscription where it will create a new Resource
Group and bootstrap all the necessary infrastructure. At the center of the BYOC
infrastructure are two AKS clusters that provide compute resources for Cube
Store and all Cube deployments you configure in the Cube UI. These AKS
clusters will have a Cube Operator installed in them that is connected to
the Cube Control Plane. The Cube Operator receives instructions from the
Control Plane and dynamically creates or destroys all the necessary
Kubernetes resources required to support your Cube deployments.
## Prerequisites
The bulk of provisioning work will be done remotely by Cube automation.
However, to get started, you'll need to provide Cube with the necessary access
along with some additional information that includes:
* **Azure Tenant ID** - the Entra ID of your Azure account
* **Azure Subscription ID** - The target subscription where Cube will be granted admin permissions to provision the BYOC infrastructure
* **Region** - The target Azure region where Cube BYOC will be installed
## Provisioning access
### Add Cube tenant to your organization
First you should add the Cube tenant to your organization. To do this,
open the [Azure Portal][azure-console] and go to **Azure Active
Directory** → **External Identities** → **Cross-tenant
access settings** → **Organizational Settings**
→ **Add Organization**.
For Tenant ID, enter `197e5263-87f4-4ce1-96c4-351b0c0c714a`.
Make sure that **B2B Collaboration** → **Inbound Access**
→ **Applications** is set to **Allows access**.
### Register Cube service principal at your organization
To register the Cube service principal for your organization, follow these
steps:
1. Log in with an account that has permissions to register Enterprise
applications.
2. Open a browser tab and go to the following URL, replacing `` with
your tenant ID:
`https://login.microsoftonline.com//oauth2/authorize?client_id=0c5d0d4b-6cee-402e-9a08-e5b79f199481&response_type=code&redirect_uri=https%3A%2F%2Fwww.microsoft.com%2F`
3. The Cube service principal has specific credentials. Check that the
following details match exactly what you see on the dialog box that pops up:
* Client ID: `d1c59948-4d4a-43dc-8d04-c0df8795ae19`
* Name: `cube-cloud-byoc-provisioner`
Once you have confirmed that all the information is correct,
select **Consent on behalf of your organization** and
click **Accept**.
### Grant admin permissions on your BYOC Azure Subscription to the cube-cloud-byoc-provisioner
On the [Azure Portal][azure-console], go to **Subscriptions**
→ *Your BYOC Subscription* → **IAM**→ **Role Assignment**
and assing `Contributor` and `Role Based Access Control Administrator` to the `cube-cloud-byoc-provisioner`
Service Principal.
## Deployment
The actual deployment will be done by Cube automation. All that's left to
do is notify your Cube contact point that access has been granted, and pass
along your Azure Tenant/Subscription/Region information.
[azure-console]: https://portal.azure.com
# Dedicated Infrastructure on Azure
Source: https://docs.cube.dev/admin/deployment/dedicated/azure/index
Connect Cube's Dedicated Infrastructure on Azure to your VNets and corporate networks, or deploy the entire data plane inside your own Azure subscription via BYOC.
On Azure, Cube offers single-tenant Dedicated Infrastructure operated by Cube,
and Bring Your Own Cloud (BYOC) operated inside your own Azure subscription.
Both options support private network connectivity to your data sources.
## Backend connectivity (Cube → your network)
Use these options to give Cube private access to your data sources, auth
providers, BI APIs targeted by Semantic Layer Sync, and anything else Cube
needs to query. See [Backend and frontend connectivity][backend-frontend] for
the full picture.
* [**Azure Private Link**][azure-private-link] — connect to data sources
exposed through Azure Private Link Services without routing traffic over
the public internet.
* [**VNet Peering**][azure-vnet-peering] — establish a VNet peering connection
between the Cube VNet and your own VNet.
## Frontend connectivity (your clients → Cube)
Expose Cube's APIs to your applications, browsers, BI tools, embedded
analytics clients, and Semantic Layer Sync-generated configs over a private
network. The pattern mirrors the AWS implementation documented in
[Private API Connectivity on AWS][aws-private-api-connectivity];
[contact us](https://cube.dev/contact) to enable the equivalent on Azure for
your tenant.
## Bring Your Own Cloud
If you'd like the entire Cube data plane to live inside your own Azure
subscription, see [Bring Your Own Cloud on Azure][azure-byoc].
[azure-private-link]: /admin/deployment/dedicated/azure/private-link
[azure-vnet-peering]: /admin/deployment/dedicated/azure/vpc-peering
[azure-byoc]: /admin/deployment/dedicated/azure/byoc
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Setting up Azure Private Link
Source: https://docs.cube.dev/admin/deployment/dedicated/azure/private-link
How to publish an Azure Private Link Service and coordinate the connection so Cube's Dedicated Infrastructure reaches your VNet privately.
This page covers **backend connectivity** — Cube reaching into your network to
query data sources, auth providers, BI APIs targeted by Semantic Layer Sync,
and other upstream services. See
[Backend and frontend connectivity][backend-frontend] for the full picture.
For **frontend connectivity** (exposing Cube's APIs to your applications,
browsers, BI tools, and embedded analytics clients), see
[Private API Connectivity on AWS][aws-private-api-connectivity]; the
equivalent pattern is available on Azure on request.
[Azure Private Link][azure-docs-private-link] enables you to access Azure
PaaS services and Azure-hosted customer-owned/partner services over a private
endpoint in your virtual network. To set up a Private Link connection between
Cube's Dedicated Infrastructure and your own VNet, you'll need to prepare a
Private Link Service, share service details with the Cube team, and approve
the incoming connection request.
**Dedicated Infrastructure vs. Bring Your Own Cloud.** The flow described on
this page — sharing service details with the Cube team and letting Cube create
the private endpoint and DNS overrides — applies to
[Dedicated Infrastructure][cube-region] operated by Cube.
In a [Bring Your Own Cloud (BYOC)][azure-byoc] deployment, the Cube VNet lives
in **your own Azure subscription**, so you own the networking. The role granted
to the Cube Operator does not include permissions to manage Private DNS Zones,
which means Cube cannot create the private endpoint or the DNS override on
your behalf. In BYOC, create the private endpoint in the Cube VNet against the
provider's Private Link Service Resource ID yourself, then create a Private DNS
Zone for the TLS hostname and link it to the Cube VNet.
## Preparing the Private Link Service
There are two common scenarios for preparing the Private Link Service:
* Connecting to a service in your Azure infrastructure
* Connecting to a service provided by a third party such as Snowflake,
Databricks, Confluent Cloud, etc.
In the case of your own infrastructure, please follow the
[official Azure documentation][azure-docs-private-link-service] to configure
the Private Link Service behind a standard Azure Load Balancer.
If your data source is hosted in a third-party infrastructure, please follow
the vendor's documentation for creating and managing a Private Link Service.
## Configuring service visibility
Azure Private Link Service enables you to control the visibility of your
private endpoint. You'll need to configure access permissions to allow Cube
to connect to your service.
To allow Cube access, please go to **Azure Portal** → **Private Link
Services** → **Your service** → **Manage visibility** and add the following
subscription ID to the allowed list: `cd69336e-c628-4a88-a56e-86900a0df732`.
This is the Azure subscription ID of Cube's Private Link consumer
subscription. Adding it authorizes Cube to discover your Private Link
Service and create a private endpoint against it; nothing else in Cube's
Azure estate gains access to your network.
Alternatively, you can configure auto-approval for faster connection
establishment by adding the same subscription ID to the auto-approval list
under **Manage auto-approval**.
## Gathering required information
To request establishing a Private Link connection, please share the following
information with the Cube team:
* **Private Link Service Resource ID** (such as
`/subscriptions/abc123/resourceGroups/myResourceGroup/providers/Microsoft.Network/privateLinkServices/myservice`)
* **Reference Name** for the record (such as "Snowflake-prod" or
"databricks-dev")
* **Ports**: a list of ports that will be accessed through this connection
* **DNS Name(s)**: see [DNS and TLS](#dns-and-tls) below
* **Cube Region:** Private Link requires Cube to be hosted on
[Dedicated Infrastructure][cube-region]. Specify which Cube Region should
host your Dedicated Infrastructure.
## DNS and TLS
How your data source is addressed inside Cube depends on whether it speaks
TLS:
* **If the service uses TLS** (HTTPS, JDBC `Encrypt=true`, etc.), share the
**DNS name(s)** the certificate is issued for — typically the same
hostname your in-network clients already use to reach it. Cube creates
internal DNS overrides inside the Dedicated Infrastructure so that the
same hostname resolves to the Private Endpoint. Keeping the original
hostname is what preserves TLS validity: the certificate's CN/SAN keeps
matching what Cube dials.
* **If the service does not use TLS** and you don't supply a DNS name, the
Cube team will share back an internal endpoint hostname (e.g. an
Azure-assigned private-endpoint DNS name) that you can configure as the
upstream when you wire the connection into Cube.
## Approving the connection
The connection approval process depends on your visibility configuration:
### Manual approval
If you haven't configured auto-approval, the Cube team will notify you once
the Private Endpoint connection request is sent. You can approve it by:
1. Going to **Azure Portal** → **Private Link Center** → **Private Link
Services** → **Your Service** → **Private endpoint connections**.
2. Finding the pending connection from Cube.
3. Clicking **Approve** and optionally providing an approval message.
Alternatively, you can approve the connection from the resource itself if it
supports Private Link natively (e.g., Storage Accounts, SQL Databases).
### Auto-approval
If you've added Cube's subscription ID to the auto-approval list, the
connection will be automatically approved upon creation and no manual action
is required.
## Using the connection
Once the connection is established, you can access your data source by
addressing it via the DNS name(s) you supplied (TLS case) or the internal
endpoint hostname returned to you by the Cube team (non-TLS case).
## Supported Regions
Azure Private Link is available in all Azure commercial regions where
Dedicated Infrastructure can be provisioned. Azure operated by 21Vianet
(China) and Azure Government regions are not supported.
[azure-docs-private-link]: https://docs.microsoft.com/azure/private-link/
[azure-docs-private-link-service]: https://docs.microsoft.com/azure/private-link/create-private-link-service-portal
[cube-region]: /admin/deployment/infrastructure#understanding-cube-cloud-region
[azure-byoc]: /admin/deployment/dedicated/azure/byoc
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Setting up VNet Peering on Azure
Source: https://docs.cube.dev/admin/deployment/dedicated/azure/vpc-peering
End-to-end checklist for VNet peering Cube's Dedicated Infrastructure with your Azure VNet for private data access.
This page covers **backend connectivity** — Cube reaching into your network to
query data sources, auth providers, BI APIs targeted by Semantic Layer Sync,
and other upstream services. See
[Backend and frontend connectivity][backend-frontend] for the full picture.
For **frontend connectivity** (exposing Cube's APIs to your applications,
browsers, BI tools, and embedded analytics clients), see
[Private API Connectivity on AWS][aws-private-api-connectivity]; the
equivalent pattern is available on Azure on request.
For cross-tenant peering in Azure, you assign the peering role to the service
principal of the peering party. Using the steps outlined below, you would
register the Cube tenant in your organization, grant peering access to the
Cube service principal, and hand over the information Cube needs to initiate
the peering.
## Granting peering access to Cube
### Add the Cube tenant to your organization
First, add the Cube tenant to your organization. Open the
[Azure Portal][azure-console] and go to **Azure Active Directory** →
**External Identities** → **Cross-tenant access settings** →
**Organizational Settings** → **Add Organization**.
For Tenant ID, enter `197e5263-87f4-4ce1-96c4-351b0c0c714a`.
Make sure that **B2B Collaboration** → **Inbound Access** →
**Applications** is set to **Allows access**.
### Register the Cube service principal at your organization
To register the Cube service principal for your organization, follow these
steps:
1. Log in with an account that has permissions to register Enterprise
applications.
2. Open a browser tab and go to the following URL, replacing ``
with your tenant ID:
`https://login.microsoftonline.com//oauth2/authorize?client_id=7f3afcf3-e061-4e1b-8261-f396646d7fc7&response_type=code&redirect_uri=https%3A%2F%2Fwww.microsoft.com%2F`
3. The Cube service principal has specific credentials. Check that the
following details match exactly what you see on the dialog box that
pops up:
* Client ID: `7f3afcf3-e061-4e1b-8261-f396646d7fc7`
* Name: `cube-dedicated-infra-peering-sp`
Once you have confirmed that all the information is correct,
select **Consent on behalf of your organization** and
click **Accept**.
### Grant peering permissions on your virtual network
As the peering role you can use the built-in `Network Contributor` role or
create a custom role (e.g. `cube-peering-role`) with the following
permissions:
* `Microsoft.Network/virtualNetworks/virtualNetworkPeerings/write`
* `Microsoft.Network/virtualNetworks/peer/action`
* `Microsoft.ClassicNetwork/virtualNetworks/peer/action`
* `Microsoft.Network/virtualNetworks/virtualNetworkPeerings/read`
* `Microsoft.Network/virtualNetworks/virtualNetworkPeerings/delete`
On the [Azure Portal][azure-console], go to **Virtual networks** →
*Virtual Network Name* → **Access Control (IAM)** → **Add** →
**Add role assignment** and fill in the following details:
* Role: `Network Contributor` or `cube-peering-role`
* Members: `cube-dedicated-infra-peering-sp`
## Information required by Cube
When reaching out to Cube support, please provide the following information:
* **Virtual Network ID:** Find this at **Virtual Networks** →
*Virtual Network Name* → **Overview** → **JSON view** →
**Resource ID** on the [Azure Portal][azure-console].
* **Virtual Network Address Spaces:** Find this at **Virtual Networks**
→ *Virtual Network Name* → **Overview** → **JSON view** →
**properties** → **addressSpace** on the [Azure Portal][azure-console].
* **Tenant ID:** Find this in **Azure Active Directory** →
**Properties** → **Tenant ID** section of the
[Azure Portal][azure-console].
* **Cube Region:** VNet Peering requires Cube to be hosted on
[Dedicated Infrastructure][cube-region]. Specify which Cube Region should
host your Dedicated Infrastructure.
## Firewall and routing
Once the peering is established, allow traffic from Cube's VNet CIDR block to
reach your data source:
1. **Network Security Groups (NSGs)** attached to the data-source subnet (or
the resource itself) must include an inbound rule that permits TCP traffic
from Cube's VNet CIDR on the database port. For example, for PostgreSQL:
| Priority | Source | Source Port | Destination | Service / Port | Action |
| -------- | ----------------------------- | ----------- | ---------------- | -------------- | ------ |
| 1000 | Cube VNet CIDR (e.g. 10.x/16) | `*` | `VirtualNetwork` | TCP / 5432 | Allow |
Cube's VNet CIDR is shared with you alongside the peering request and is
also visible in the Azure Portal on the **Virtual networks** →
**\** → **Peerings** → **\** → **Address
space** field.
2. **Azure Firewall / third-party NVAs**: if traffic between your subnets
transits a firewall, add a rule permitting TCP from the Cube VNet CIDR to
the data source's IP and port.
3. **User-defined routes (UDRs)**: confirm that the route tables on your
subnets do not blackhole Cube's CIDR via `0.0.0.0/0` next-hop appliances.
Ensure traffic destined for Cube's VNet CIDR is routed to the **Virtual
network peering** next-hop.
4. **Data source firewall**: if the resource has its own firewall (e.g. an
Azure SQL Server firewall or a PaaS-level allow-list), add Cube's VNet
CIDR there as well.
## Supported Regions
VNet Peering is available in all Azure commercial regions where Dedicated
Infrastructure can be provisioned. Azure operated by 21Vianet (China) and
Azure Government regions are not supported.
[azure-console]: https://portal.azure.com
[cube-region]: /admin/deployment/infrastructure#understanding-cube-cloud-region
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Bring Your Own Cloud on GCP
Source: https://docs.cube.dev/admin/deployment/dedicated/gcp/byoc
Project setup, permissions, and provisioning flow for deploying Cube BYOC inside a dedicated GCP project.
With Bring Your Own Cloud (BYOC) on Google Cloud Platform (GCP), all the components interacting with private data are deployed on
the customer infrastructure on GCP and managed by the Cube Control Plane via the Cube Operator.
This document provides step-by-step instructions for deploying Cube BYOC on GCP.
## Prerequisites
The bulk of provisioning work will be done remotely by Cube automation.
However, to get started, you'll need:
### Required Information
* **GCP Project ID:** A dedicated GCP project ID that will exclusively host Cube-managed infrastructure.
This should be a new, isolated project created specifically for Cube BYOC.
* **GCP Region:** [The GCP region][gcp-docs-regions] where the BYOC resources
should be deployed.
### Required Permissions
You'll need to have the following permissions in your GCP organization/folder to complete the setup:
* **Project Creator** (`roles/resourcemanager.projectCreator`) - To create a new dedicated project
* **Project IAM Admin** (`roles/resourcemanager.projectIamAdmin`) - To grant permissions in the project
* **Billing Account User** (`roles/billing.user`) - To link billing to the new project
If you don't have these permissions, contact your GCP organization administrator.
## Provisioning access
### Step 1: Create a dedicated GCP project
We strongly recommend creating a dedicated GCP project that will exclusively host
Cube-managed infrastructure. This project isolation approach simplifies permission
management and provides clear resource boundaries.
1. Navigate to the [GCP Console][gcp-console]
2. Click **Create Project**
3. Enter a project name (e.g., "cube-cloud-byoc")
4. Note the **Project ID** (not the project name) - you'll need this for subsequent steps
5. Select your billing account
6. Click **Create**
Make sure billing is enabled for the project. You can verify this by navigating to
**Billing** in the GCP Console and confirming the project is linked to an active billing account.
### Step 2: Enable required APIs
Before granting permissions, enable the necessary GCP APIs in your dedicated project.
This ensures that subsequent API calls will work correctly.
**Required APIs:**
* **Compute Engine API** (`compute.googleapis.com`) - For VPC networks and compute resources
* **Kubernetes Engine API** (`container.googleapis.com`) - For GKE clusters
* **Cloud Storage API** (`storage.googleapis.com`) - For Cube Store buckets
* **IAM API** (`iam.googleapis.com`) - For service account management
* **Cloud Resource Manager API** (`cloudresourcemanager.googleapis.com`) - For project IAM operations
* **Service Networking API** (`servicenetworking.googleapis.com`) - For private service connectivity
**Note:** DNS and Artifact Registry APIs are not required in your project. Cube manages DNS in its own project,
and container images are pulled from Cube's Artifact Registry using Cube-provided credentials.
You can enable these APIs through the [API Library][gcp-api-library] in the GCP Console,
or use the `gcloud` command:
```bash theme={"dark"}
# Set your project ID
export PROJECT_ID="your-cube-byoc-project-id"
# Enable all required APIs
gcloud services enable \
compute.googleapis.com \
container.googleapis.com \
storage.googleapis.com \
iam.googleapis.com \
cloudresourcemanager.googleapis.com \
servicenetworking.googleapis.com \
--project=$PROJECT_ID
```
### Step 3: Grant IAM permissions
In order to manage resources in the Cube-dedicated GCP project, the Cube service principal
needs to be granted administrative permissions to a set of services.
Navigate to **IAM & Admin > IAM** in your dedicated project and add the following IAM
binding for the Cube service account:
**Principal:** `cube-cloud-byoc-installer@cube-cloud-byoc.iam.gserviceaccount.com`
**Roles:**
* **Compute Admin** (`roles/compute.admin`) - Allows creation and management of VPC networks, subnets, routers, NAT gateways, firewall rules, IP addresses, and Private Service Connect endpoints
* **Kubernetes Engine Admin** (`roles/container.admin`) - Allows creation and management of GKE clusters and node pools
* **Storage Admin** (`roles/storage.admin`) - Allows creation and management of Cloud Storage buckets for Cube Store
* **Service Account Admin** (`roles/iam.serviceAccountAdmin`) - Allows creation and management of service accounts for cluster nodes and workload identity
* **Service Account Key Admin** (`roles/iam.serviceAccountKeyAdmin`) - Allows creation and management of service account keys for Cube Store authentication
* **Project IAM Admin** (`roles/resourcemanager.projectIamAdmin`) - Allows granting IAM permissions to created resources (e.g., bucket access for service accounts)
You can grant these permissions through the Google Cloud Console UI or using the
`gcloud` command-line tool:
```bash theme={"dark"}
# Set your project ID (replace with your actual project ID)
export PROJECT_ID="your-cube-byoc-project-id"
# Set the Cube service account (use this exact value)
export CUBE_SA="cube-cloud-byoc-installer@cube-cloud-byoc.iam.gserviceaccount.com"
# Grant all required roles
gcloud projects add-iam-policy-binding $PROJECT_ID \
--member="serviceAccount:$CUBE_SA" \
--role="roles/compute.admin"
gcloud projects add-iam-policy-binding $PROJECT_ID \
--member="serviceAccount:$CUBE_SA" \
--role="roles/container.admin"
gcloud projects add-iam-policy-binding $PROJECT_ID \
--member="serviceAccount:$CUBE_SA" \
--role="roles/storage.admin"
gcloud projects add-iam-policy-binding $PROJECT_ID \
--member="serviceAccount:$CUBE_SA" \
--role="roles/iam.serviceAccountAdmin"
gcloud projects add-iam-policy-binding $PROJECT_ID \
--member="serviceAccount:$CUBE_SA" \
--role="roles/iam.serviceAccountKeyAdmin"
gcloud projects add-iam-policy-binding $PROJECT_ID \
--member="serviceAccount:$CUBE_SA" \
--role="roles/resourcemanager.projectIamAdmin"
```
### Step 4: Grant Service Account User permissions
Additionally, the Cube service account needs permission to use the default Compute Engine service account for GKE node pools.
Make sure you have the `PROJECT_ID` and `CUBE_SA` environment variables set from Step 3 before running these commands.
Run the following command to grant the necessary permissions:
```bash theme={"dark"}
# Get the project number
export PROJECT_NUMBER=$(gcloud projects describe $PROJECT_ID --format='value(projectNumber)')
# Grant the Cube service account permission to use the default compute service account
gcloud iam service-accounts add-iam-policy-binding \
${PROJECT_NUMBER}-compute@developer.gserviceaccount.com \
--member="serviceAccount:$CUBE_SA" \
--role="roles/iam.serviceAccountUser" \
--project=$PROJECT_ID
```
This allows the Cube service account to create GKE clusters that use the project's default compute service account for worker nodes.
### Step 5: Verify setup
Before notifying Cube, verify that all permissions and APIs are correctly configured:
```bash theme={"dark"}
# Verify APIs are enabled
gcloud services list --enabled --project=$PROJECT_ID | grep -E '(compute|container|storage|iam|cloudresourcemanager|servicenetworking)'
# Verify IAM bindings for the Cube service account
gcloud projects get-iam-policy $PROJECT_ID \
--flatten="bindings[].members" \
--format="table(bindings.role)" \
--filter="bindings.members:serviceAccount:cube-cloud-byoc-installer@cube-cloud-byoc.iam.gserviceaccount.com"
# Verify Service Account User permission
gcloud iam service-accounts get-iam-policy \
${PROJECT_NUMBER}-compute@developer.gserviceaccount.com \
--project=$PROJECT_ID
```
If all commands return the expected results, you're ready to proceed with deployment.
## Deployment
The actual deployment will be done by Cube automation. All that's left to
do is notify your Cube contact point that access has been granted, and pass
along your GCP Project ID and Region information.
After deployment, Cube will manage the following resources in your dedicated project:
* A VPC network with subnets, Cloud Router, and Cloud NAT for outbound connectivity
* A GKE cluster with node pools for running Cube applications
* Cloud Storage buckets for Cube Store data
* Service accounts and IAM bindings for secure resource access
* Firewall rules and network policies for security
[gcp-console]: https://console.cloud.google.com/
[gcp-docs-regions]: https://cloud.google.com/compute/docs/regions-zones
[gcp-api-library]: https://console.cloud.google.com/apis/library
# Dedicated Infrastructure on GCP
Source: https://docs.cube.dev/admin/deployment/dedicated/gcp/index
Connect Cube's Dedicated Infrastructure on GCP to your VPC networks via Private Service Connect or VPC Peering, or deploy the entire data plane inside your own GCP project via BYOC.
On GCP, Cube offers single-tenant Dedicated Infrastructure operated by Cube,
and Bring Your Own Cloud (BYOC) operated inside your own GCP project. Both
options support private network connectivity to your data sources.
## Backend connectivity (Cube → your network)
Use these options to give Cube private access to your data sources, auth
providers, BI APIs targeted by Semantic Layer Sync, and anything else Cube
needs to query. See [Backend and frontend connectivity][backend-frontend] for
the full picture.
* [**Private Service Connect**][gcp-private-service-connect] — connect to
data sources exposed through GCP Service Attachments without routing traffic
over the public internet.
* [**VPC Peering**][gcp-vpc-peering] — establish a VPC peering connection
between the Cube VPC and your own VPC.
## Frontend connectivity (your clients → Cube)
Expose Cube's APIs to your applications, browsers, BI tools, embedded
analytics clients, and Semantic Layer Sync-generated configs over a private
network. The pattern mirrors the AWS implementation documented in
[Private API Connectivity on AWS][aws-private-api-connectivity];
[contact us](https://cube.dev/contact) to enable the equivalent on GCP for
your tenant.
## Bring Your Own Cloud
If you'd like the entire Cube data plane to live inside your own GCP project,
see [Bring Your Own Cloud on GCP][gcp-byoc].
[gcp-private-service-connect]: /admin/deployment/dedicated/gcp/private-service-connect
[gcp-vpc-peering]: /admin/deployment/dedicated/gcp/vpc-peering
[gcp-byoc]: /admin/deployment/dedicated/gcp/byoc
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Setting up Google Private Service Connect
Source: https://docs.cube.dev/admin/deployment/dedicated/gcp/private-service-connect
How to publish a Service Attachment and coordinate Private Service Connect so Cube's Dedicated Infrastructure reaches your VPC privately.
This page covers **backend connectivity** — Cube reaching into your network to
query data sources, auth providers, BI APIs targeted by Semantic Layer Sync,
and other upstream services. See
[Backend and frontend connectivity][backend-frontend] for the full picture.
For **frontend connectivity** (exposing Cube's APIs to your applications,
browsers, BI tools, and embedded analytics clients), see
[Private API Connectivity on AWS][aws-private-api-connectivity]; the
equivalent pattern is available on GCP on request.
[Private Service Connect][gcp-docs-psc] (PSC) provides private connectivity
between VPC networks in different projects or organizations, without VPC
peering or exposing your traffic to the public internet. To set up a PSC
connection between Cube's Dedicated Infrastructure and your own VPC, you'll
need to publish a Service Attachment, share its details with the Cube team,
and approve the incoming connection request.
**Dedicated Infrastructure vs. Bring Your Own Cloud.** The flow described on
this page — sharing service details with the Cube team and letting Cube create
the PSC endpoint and DNS overrides — applies to
[Dedicated Infrastructure][cube-region] operated by Cube.
In a [Bring Your Own Cloud (BYOC)][gcp-byoc] deployment, the Cube VPC lives in
**your own GCP project**, so you own the networking. The service account
granted to the Cube Operator does not include Cloud DNS admin permissions,
which means Cube cannot create the PSC endpoint or the private DNS zone needed
for the hostname override on your behalf. In BYOC, create the PSC endpoint in
the Cube VPC against the provider's Service Attachment yourself, then create a
Cloud DNS private zone for the TLS hostname and attach it to the Cube VPC.
## Preparing the Service Attachment
There are two common scenarios for preparing the Service Attachment:
* Connecting to a service in your GCP infrastructure
* Connecting to a service provided by a third party such as Snowflake,
Databricks, Confluent Cloud, etc.
In the case of your own infrastructure, follow the
[official GCP documentation][gcp-docs-publish-service] to publish a Service
Attachment that points at an
[internal passthrough or proxy Network Load Balancer][gcp-docs-internal-lb]
in front of your data source.
If your data source is hosted in a third-party infrastructure, follow the
vendor's documentation for creating and managing a Service Attachment.
## Allowing the Cube consumer project
PSC service attachments can restrict which consumer projects are allowed to
create a PSC endpoint against them. Cube's PSC consumer project is
`cube-cloud-dedicated`.
In the GCP Console, go to **Network services → Private Service Connect →
Published services → \** and add `cube-cloud-dedicated` to
**Accepted projects**. For faster connection establishment, you can also
add the same project to the **auto-accept** list so the connection is
approved automatically when Cube initiates it.
`cube-cloud-dedicated` is the GCP project Cube uses to host Dedicated
Infrastructure PSC endpoints. Adding it to your accepted-projects list
authorizes Cube to create a private endpoint against your Service
Attachment; nothing else in Cube's GCP estate gains access to your network.
## Gathering required information
To request establishing a PSC connection, please share the following
information with the Cube team:
* **Service Attachment URI** (such as
`projects//regions//serviceAttachments/`)
* **Reference Name** for the record (such as "Snowflake-prod" or
"clickhouse-dev")
* **Ports**: a list of ports that will be accessed through this connection
* **DNS Name(s)**: see [DNS and TLS](#dns-and-tls) below
* **Cube Region:** PSC requires Cube to be hosted on
[Dedicated Infrastructure][cube-region]. Specify which Cube Region should
host your Dedicated Infrastructure.
## DNS and TLS
How your data source is addressed inside Cube depends on whether it speaks
TLS:
* **If the service uses TLS** (HTTPS, JDBC `sslmode=require`, etc.), share
the **DNS name(s)** the certificate is issued for — typically the same
hostname your in-network clients already use to reach it. Cube creates
internal DNS overrides inside the Dedicated Infrastructure so that the
same hostname resolves to the PSC endpoint. Keeping the original hostname
is what preserves TLS validity: the certificate's CN/SAN keeps matching
what Cube dials.
* **If the service does not use TLS** and you don't supply a DNS name, the
Cube team will share back an internal endpoint hostname that you can
configure as the upstream when you wire the connection into Cube.
## Accepting the connection
The approval flow depends on how your Service Attachment is configured:
* **Manual acceptance.** Cube will notify you once the connection request
has been sent. Approve it in the GCP Console under **Network services →
Private Service Connect → Published services → \ →
Connected endpoints**, then select the pending connection and click
**Accept**.
* **Auto-accept.** If you added `cube-cloud-dedicated` to the auto-accept
list, the connection is approved automatically upon creation and no
further action is required.
## Using the connection
Once the connection is established, you can access your data source by
addressing it via the DNS name(s) you supplied (TLS case) or the internal
endpoint hostname returned to you by the Cube team (non-TLS case).
## Supported Regions
Private Service Connect is available in all GCP commercial regions where
Dedicated Infrastructure can be provisioned. GCP regions in mainland China
(serviced by partner providers) are not supported.
[gcp-docs-psc]: https://cloud.google.com/vpc/docs/private-service-connect
[gcp-docs-publish-service]: https://cloud.google.com/vpc/docs/configure-private-service-connect-producer
[gcp-docs-internal-lb]: https://cloud.google.com/load-balancing/docs/internal
[cube-region]: /admin/deployment/infrastructure#understanding-cube-cloud-region
[gcp-byoc]: /admin/deployment/dedicated/gcp/byoc
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Setting up VPC Peering on GCP
Source: https://docs.cube.dev/admin/deployment/dedicated/gcp/vpc-peering
End-to-end checklist for VPC peering Cube's Dedicated Infrastructure with your GCP VPC network for private data access.
This page covers **backend connectivity** — Cube reaching into your network to
query data sources, auth providers, BI APIs targeted by Semantic Layer Sync,
and other upstream services. See
[Backend and frontend connectivity][backend-frontend] for the full picture.
For **frontend connectivity** (exposing Cube's APIs to your applications,
browsers, BI tools, and embedded analytics clients), see
[Private API Connectivity on AWS][aws-private-api-connectivity]; the
equivalent pattern is available on GCP on request.
VPC Peering requires Cube to be hosted on
[Dedicated Infrastructure][cube-region]. Let the Cube team know which Cube
Region should host your Dedicated Infrastructure.
Cube will provision the Dedicated VPC and provide the following information
you can use to create the peering request:
* **GCP Project ID:** `cube-cloud-dedicated` (the project Cube uses to host
Dedicated VPCs).
* **VPC Network Name:** shared with you by the Cube team once the Dedicated
VPC is provisioned.
## Setup
### Creating the peering connection
After receiving the information above, create a
[VPC peering request][gcp-docs-vpc-peering], either through the
[GCP Web Console][gcp-console] or an infrastructure-as-code tool. To send a
VPC peering request through the Google Cloud Console, follow
[the instructions here][gcp-docs-create-vpc-peering], with the following
amendments:
* In Step 6, use the project ID `cube-cloud-dedicated` and the network name
provided by Cube.
* In Step 7, ensure **Import custom routes** and **Export custom routes** are
selected so that the necessary routes are created.
### Firewall and routing
Once the peering is established, configure your VPC firewall rules to allow
inbound TCP traffic from Cube's VPC CIDR block to your data source on the
database port. Cube's VPC CIDR is shared with you alongside the peering
request and is also visible in the GCP Console on the **VPC network** →
**\** → **VPC network peering** → **\** page as
the **Peer VPC network** subnet ranges.
If your data source is in a different project or subnet that transits a
firewall or Cloud NAT, add a matching allow rule for Cube's CIDR there as
well.
## Cloud SQL
Google Cloud SQL databases
[can only be peered to a VPC within the same GCP project][gcp-docs-vpc-peering-restrictions].
If you need Cube to reach a Cloud SQL instance, prefer
[Private Service Connect][gcp-private-service-connect] (Cloud SQL supports
PSC natively), or alternatively provision a small VM in your GCP project
running the [Cloud SQL Auth Proxy][gcp-cloudsql-auth-proxy].
## Supported Regions
VPC Peering is available in all GCP commercial regions where Dedicated
Infrastructure can be provisioned. GCP regions in mainland China (serviced
by partner providers) are not supported.
[gcp-cloudsql-auth-proxy]: https://cloud.google.com/sql/docs/mysql/connect-admin-proxy
[gcp-console]: https://console.cloud.google.com/
[gcp-docs-create-vpc-peering]: https://cloud.google.com/vpc/docs/using-vpc-peering#creating_a_peering_configuration
[gcp-docs-vpc-peering]: https://cloud.google.com/vpc/docs/vpc-peering
[gcp-docs-vpc-peering-restrictions]: https://cloud.google.com/vpc/docs/vpc-peering#restrictions
[cube-region]: /admin/deployment/infrastructure#understanding-cube-cloud-region
[gcp-private-service-connect]: /admin/deployment/dedicated/gcp/private-service-connect
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[backend-frontend]: /admin/deployment/dedicated#backend-and-frontend-connectivity
# Dedicated Infrastructure
Source: https://docs.cube.dev/admin/deployment/dedicated/index
Run Cube on dedicated single-tenant infrastructure (managed by Cube) or on your own AWS, Azure, or GCP account (BYOC), with private connectivity to your data sources and APIs.
Cube offers two flavors of single-tenant deployment: **Dedicated Infrastructure**
managed by Cube in our cloud accounts, and **Bring Your Own Cloud (BYOC)** managed
by Cube inside your own cloud account. Both options give you isolated compute,
the ability to route traffic over private networks, and integrations with
services in your VPC or VNet.
Available on the [Enterprise plan](https://cube.dev/pricing) with the
[Dedicated Infrastructure][ref-dedicated-infra] add-on.
## Single-tenant Cube cluster
With Dedicated Infrastructure, Cube provisions and operates a **single-tenant**
cluster for you in a Cube-managed account on AWS, GCP, or Azure. *Single-tenant*
means the cluster — VPC/VNet, compute, storage, and the Cube data plane that
runs your deployments — is dedicated entirely to your organization and not
shared with any other customer. The cluster lives in a Cube
[Region][cube-region], can be peered or PrivateLink/PSC-connected to your own
networks, and can optionally expose Cube's APIs to your network so that no Cube
traffic ever crosses the public internet.
## Bring Your Own Cloud (BYOC)
On the Enterprise plan, Cube is also available as **Bring Your Own Cloud**: all
components that interact with your private data are deployed inside your own
AWS, Azure, or GCP account and managed remotely by the Cube Control Plane via
the Cube Operator. This keeps all data plane resources within your boundary
while preserving the managed-service experience.
[Contact us](https://cube.dev/contact) for details.
## Backend and frontend connectivity
There are two distinct directions in which Cube exchanges traffic with your
network, and each has its own connectivity story:
* **Backend connectivity** — traffic that flows **from Cube into your network**.
Cube uses these connections to query the things it needs to function:
databases and warehouses, auth providers (e.g. an internal OIDC issuer),
upstream BI APIs that Semantic Layer Sync targets, and any other service the
Cube data plane has to reach. PrivateLink, Private Link, Private Service
Connect, and Peering on the provider pages below all configure backend
connectivity.
* **Frontend connectivity** — traffic that flows **from your network into
Cube**. Anything that needs to query Cube falls in this bucket: the Cube UI
running in employee browsers, application servers, BI tools, embedded
analytics clients, and Semantic Layer Sync-generated configs. Frontend
connectivity is currently documented for AWS in
[Private API Connectivity on AWS][aws-private-api-connectivity], and
equivalent patterns are available on Azure and GCP on request.
Most enterprise deployments end up using both: a backend
PrivateLink/PSC/peering into the customer's data network, plus a frontend
private API endpoint so the Cube UI and BI tools talk to Cube over the same
private fabric.
## Choose a provider
Dedicated Infrastructure, BYOC, and private connectivity on AWS.
Dedicated Infrastructure and BYOC on GCP.
Dedicated Infrastructure, BYOC, and private connectivity on Azure.
[ref-dedicated-infra]: /admin/deployment/infrastructure#dedicated-infrastructure
[aws-private-api-connectivity]: /admin/deployment/dedicated/aws/private-api-connectivity
[cube-region]: /admin/deployment/infrastructure#understanding-cube-cloud-region
# Bring-Your-Own Pre-aggregation Storage
Source: https://docs.cube.dev/admin/deployment/dedicated/pre-aggregation-storage
Supply your own object storage bucket as the backend for Cube Store pre-aggregated data so that all data at rest stays within your infrastructure.
On the Enterprise plan, Dedicated Infrastructure customers can supply their
own object storage bucket to be used as the underlying storage for Cube Store
pre-aggregated data. This lets you keep all data at rest fully within your
own infrastructure while still leveraging the managed compute and operations
of Dedicated Infrastructure.
Available on the [Enterprise plan](https://cube.dev/pricing) with Dedicated
Infrastructure. [Contact us](https://cube.dev/contact) to enable this option
for your tenant.
## AWS — S3
To activate this option on AWS:
1. Create an S3 bucket in the same region as your Cube Region.
2. Generate a new AWS Access Key with full access to that bucket.
3. Request activation from your Customer Success Manager and share the
following:
* **AWS Access Key ID**
* **AWS Secret Access Key**
* **S3 Bucket ARN**
## GCP — Cloud Storage
To activate this option on GCP:
1. Create a Cloud Storage bucket in the same region as your Cube Region.
2. Create a service account with full access to that bucket and generate a
JSON service-account key.
3. Request activation from your Customer Success Manager and share the
following:
* **GCS Bucket Name**
* **Service-account JSON key** (transferred securely)
## Azure — Blob Storage
To activate this option on Azure:
1. Create a Storage Account and Blob container in the same region as your
Cube Region.
2. Create a SAS token (or service principal) with full read/write/delete
access to the container.
3. Request activation from your Customer Success Manager and share the
following:
* **Storage Account name**
* **Container name**
* **Access credentials** (SAS token or service-principal details)
## Supported Regions
Bring-Your-Own Pre-aggregation Storage is available wherever Dedicated
Infrastructure is supported on the corresponding cloud — see the per-provider
[Supported Regions](#supported-regions) sections in the connectivity docs for
the exact list. Government and China regions are not supported.
# Deployment types
Source: https://docs.cube.dev/admin/deployment/deployment-types
Cube provides three deployment types: Shared, Dedicated, and Multi-cluster.
* [Shared](#shared) — designed for development use cases. Runs on compute
shared with other deployments within the selected region.
* [Dedicated](#dedicated) — designed for production workloads and
high-availability. Runs on compute dedicated to your deployment.
* [Multi-cluster](#multi-cluster) — designed for demanding production workloads,
high-scalability, high-availability, and advanced multi-tenancy configurations.
Runs on multiple Dedicated deployments.
## Shared
Available for free, no credit card required. Your free trial is limited to 2
Shared deployments and only 1,000 queries per day. Upgrade to
[any paid plan](https://cube.dev/pricing) to unlock all features.
Shared deployments run on compute shared with other deployments within the
selected region, which keeps the cost low but means resources aren't reserved
for you exclusively.
If your account uses [single-tenant infrastructure][ref-dedicated-infra],
Shared deployments are only shared with your other deployments on that
infrastructure — never with other customers. Your environment remains fully
isolated at the infrastructure level.
Shared deployments are designed for development use cases only. This makes
it easy to get started with Cube quickly, and also allows you to build and
query pre-aggregations on-demand.
Shared deployments don't have dedicated [refresh workers][ref-refresh-worker]
and, consequently, they do not refresh pre-aggregations on schedule.
Shared deployments do not provide high-availability nor do they guarantee
fast response times. Shared deployments also [auto-suspend][ref-auto-sus]
after 30 minutes of inactivity, which can cause the first request after the
deployment wakes up to take additional time to process. They also have
[limits][ref-limits] on the maximum number of queries per day and the maximum
number of Cube Store Workers. We **strongly** advise not using a Shared
deployment in a production environment, it is for testing and learning about
Cube only and will not deliver a production-level experience for your users.
You can try a Shared deployment by
[signing up for Cube](https://cubecloud.dev/auth/signup) to try it free
(no credit card required).
## Dedicated
Available on [all paid plans](https://cube.dev/pricing).
Dedicated deployments run on compute dedicated exclusively to your deployment,
giving you predictable performance and full control over capacity.
Dedicated deployments are designed to support high-availability production
workloads. It consists of several key components, including starting with 2 Cube
API instances, 1 Cube Refresh Worker and 2 Cube Store Routers - all of which run
on compute dedicated to your deployment. The deployment can automatically scale
to meet the needs of your workload by adding more components as necessary;
check the page on [scalability][ref-scalability] to learn more.
## Multi-cluster
Multi-cluster deployments are designed for demanding production workloads,
high-scalability, high-availability, and large [multi-tenancy][ref-multitenancy]
configurations, e.g., with more than 100 tenants.
Available on [Premium and above plans](https://cube.dev/pricing).
It provides you with two options:
* Scale the number of [Dedicated](#dedicated) deployments serving your
workload, allowing to route requests over up to 10 Dedicated deployments and
up to 100 API instances.
* Optionally, scale the number of Cube Store routers, allowing for increased
Cube Store querying performance.
Each Dedicated deployment is billed separately, and all Dedicated deployments
can use auto-scaling to match demand.
### Configuring Multi-cluster
To switch your deployment to Multi-cluster, navigate to
**Settings → General**, select it under **Type**, and confirm
with **✓**.
To set the number of Dedicated deployments within your Multi-cluster
deployment, navigate to **Settings → Configuration** and edit
**Number of clusters**.
### Routing traffic between Dedicated deployments
Cube routes requests between multiple Dedicated deployments within a
Multi-cluster deployment based on [`context_to_app_id`][ref-ctx-to-app-id].
In most cases, it should return an identifier that does not change over time
for each tenant.
The following implementation will make sure that all requests from a
particular tenant are always routed to the same Dedicated deployment. This
approach ensures that only one Dedicated deployment keeps compiled data model
cache for each tenant and serves its requests. It allows to reduce the
footprint of the compiled data model cache on individual Dedicated deployments.
```python title="Python" theme={"dark"}
from cube import config
@config('context_to_app_id')
def context_to_app_id(ctx: dict) -> str:
return f"CUBE_APP_{ctx['securityContext']['tenant_id']}"
```
```js title="JavaScript" theme={"dark"}
module.exports = {
contextToAppId: ({ securityContext }) => {
return `CUBE_APP_${securityContext.tenant_id}`
}
}
```
If your implementation of `context_to_app_id` returns identifiers that change
over time for each tenant, requests from one tenant would likely hit multiple
Dedicated deployments and you would not have the benefit of reduced memory
footprint. Also you might see 502 or timeout errors in case of different
deployment nodes would return different `context_to_app_id` results for the
same request.
## Switching between deployment types
To switch a deployment's type, go to the deployment's **Settings** screen
and select from the available options.
[ref-ctx-to-app-id]: /reference/configuration/config#context_to_app_id
[ref-limits]: /admin/deployment/limits#resources
[ref-scalability]: /admin/deployment/scalability
[ref-multitenancy]: /embedding/multitenancy
[ref-auto-sus]: /admin/deployment/auto-suspension
[ref-refresh-worker]: /cube-core/architecture#refresh-worker
[ref-dedicated-infra]: /admin/deployment/infrastructure#dedicated-infrastructure
# Encryption keys
Source: https://docs.cube.dev/admin/deployment/encryption-keys
The Encryption Keys page in Cube Cloud allows to manage data-at-rest encryption in Cube Store.
Available on the [Enterprise plan](https://cube.dev/pricing).
Also requires the M [Cube Store Worker tier](/admin/account-billing/pricing#cube-store-worker-tiers).
Navigate to **Settings → Encryption Keys** in your Cube Cloud deployment
to [provide](#add-a-key), [rotate](#rotate-a-key), or [drop](#drop-a-key)
your own customer-managed keys (CMK) for Cube Store.
## Customer-managed keys for Cube Store
On the **Encryption Keys** page, you can see all previously provided keys:
### Add a key
To add an encryption key, click **Create** to open a modal window.
Provide the key name and the key value: an 256-bit AES encryption key, encoded
in [standard Base64][link-base64] in its canonical representation.
**Once the first encryption key is added, Cube Store will assume that data-at-rest
encryption is enabled.** After that, querying unencrypted pre-aggregation partitions
will yield the following error: `Invalid Parquet file in encrypted mode. File (or
at least the Parquet footer) is not encrypted`.
It may take a few minutes for any changes to encryption keys to take effect.
After the refresh worker builds or rebuilds pre-aggregation partitions with
respect to their [refresh strategy][ref-pre-aggs-refresh-strategy] or after they
are [built manually][ref-pre-aggs-build-manually], their data will be encrypted.
**For encryption, the most recently added encryption key is used.** For decryption,
all previously provided keys can be used, if there are still any pre-aggregation
partitions encrypted with those keys.
### Rotate a key
To rotate an encryption key, you have to [add a new key](#add-a-key) and then
rebuild pre-aggregation partitions using this key, either by the means of the
refresh worker, or manually.
You can check which encryption key is used by any pre-aggregation partition by
querying `system.tables` in Cube Store via [SQL Runner][ref-sql-runner]:
Only newly built or rebuilt pre-aggregation partitions will be encrypted using the
newly added encryption key. Previously built partitions will still be encrypted
using previously provided keys. If you [drop a key](#drop-a-key) before these
partitions are rebuilt, querying them will yield an error.
If you're using [incremental pre-aggregations][ref-pre-aggs-incremental], the
refresh worker will likely only rebuild some of their partitions. You have to [rebuild
them manually][ref-pre-aggs-build-manually] to ensure that the new encryption key
is used.
### Drop a key
To drop an encryption key, click **Delete** next to it.
[ref-cube-store-encryption]: /docs/pre-aggregations/running-in-production#data-at-rest-encryption
[link-base64]: https://datatracker.ietf.org/doc/html/rfc4648#section-4
[ref-pre-aggs-refresh-strategy]: /docs/pre-aggregations/using-pre-aggregations#refresh-strategy
[ref-pre-aggs-build-manually]: /admin/monitoring/pre-aggregations
[ref-pre-aggs-incremental]: /reference/data-modeling/pre-aggregations#incremental
[ref-sql-runner]: /docs/data-modeling/sql-runner
# Deployment environments
Source: https://docs.cube.dev/admin/deployment/environments
Outlines production, staging, and per-developer environments and how they differ inside one Cube deployment.
Every Cube Cloud deployment provides a number of environments:
* A single [production environment](#production-environment).
* Multiple [staging environments](#staging-environments).
* Per-user [development environments](#development-environments).
## Production environment
This is the main environment. It runs the data model from the *main branch*.
### Availability
The production environment is *always available* unless [suspended][ref-suspend].
### Resources and costs
Depending on the [deployment type][ref-deployment-types], the production environment
either runs on a Shared deployment or on a Dedicated deployment with [multiple API
instances][ref-api-instance-scalability], incurring [relevant
costs][ref-pricing-deployment-tiers].
### Cube version
Production environments run a [Cube version][ref-version] from the selected
[update channel][ref-version-channel].
## Staging environments
Staging environments are activated automatically for specific source code branches
when a branch is switched to in the Cube Cloud UI. Any [development mode][ref-dev-mode]
changes must be committed to a branch to be available in this environment.
### Availability
By default, they are only active and accessible while viewed by at least one user.
When no users are viewing the branch, the environment becomes inactive and inaccessible.
If you'd like to update this setting for a specific branch (e.g., for testing purposes),
go to **Settings → Staging Environments** and check the toggle next to it:
* **Toggle off** (default). A staging environment is only active when viewed by users.
When no users are viewing the branch, queries to this environment will fail.
* **Toggle on**. A staging environment remains *always active* and accessible,
regardless of user activity.
The same toggle is available in the [CLI][ref-cli] and the [Control Plane
API][ref-control-plane-api-staging], so it can be scripted, e.g. from CI:
```bash theme={"dark"}
cube data-model enable-branch DEPLOYMENT_ID BRANCH_NAME
cube data-model disable-branch DEPLOYMENT_ID BRANCH_NAME
```
### Resources and costs
Staging environments run on Shared deployments, incurring [relevant
costs][ref-pricing-deployment-tiers]. However, they automatically suspend after 10 minutes
of inactivity, regardless of the [availability](#availability) setting, so you are only
charged for the time when staging environments are being used.
### Cube version
Staging environments always run the *most up-to-date* Cube version.
## Development environments
A development environment is activated automatically when a user enters the [development
mode][ref-dev-mode] on a specific branch of the source code. It updates automatically
when a user saves changes to the data model.
Only one development environment is allocated per user.
### Availability
A development environment is only active and accessible while viewed by a user in the
development mode. Otherwise, queries to this environment will fail.
### Resources and costs
Development environments run on Shared deployments, incurring [relevant
costs][ref-pricing-deployment-tiers]. However, they automatically suspend after 10 minutes
of inactivity, so you are only charged for the time when development environments are
being used.
### Cube version
Development environments always run the *most up-to-date* Cube version.
## API endpoints
Each environment provides its own set of API endpoints.
You can access them on the [**Overview** page][ref-overview] of your deployment
or by navigating to the **Integrations** page and [clicking **API
credentials**][ref-credentials].
[ref-dev-mode]: /docs/data-modeling/dev-mode
[ref-deployment-types]: /admin/deployment/deployment-types
[ref-api-instance-scalability]: /admin/deployment/scalability#auto-scaling-of-api-instances
[ref-pricing-deployment-tiers]: /admin/account-billing/pricing
[ref-suspend]: /admin/deployment/auto-suspension
[ref-overview]: /admin/connect-to-data/visualization-tools
[ref-credentials]: /admin/connect-to-data/visualization-tools
[ref-version]: /admin/deployment#cube-version
[ref-version-channel]: /admin/deployment#update-channels
[ref-cli]: /reference/cli
[ref-control-plane-api-staging]: /reference/control-plane-api#buildapiv1deploymentsdeployment_idbranchesstaging-environment
# Overview
Source: https://docs.cube.dev/admin/deployment/index
Introduces Cube deployments and how to browse, create, and operate the environments that run your semantic layer.
Deployments are top-level entities in Cube Cloud. They include the source
code, configuration, allocated resources, API endpoints and essentially
do all the heavy lifting, running workloads and fulfilling requests.
## List of deployments
Each account in Cube Cloud can have multiple deployments. Once you've
logged in, you can access the list of deployments by clicking on
the **Cube Cloud** logo in the top left corner.
With more than one deployment, an admin can pin one as the account-wide default from its
row's **⋯** menu (**Set as default** / **Remove as default**). This is what a user
lands on when they open Cube without a deployment in the URL and haven't set their own
[default deployment](/docs/preferences#default-deployment) or previously switched deployments.
## Creating a new deployment
Creating a new deployment is an essential prerequisite to running a Cube
project in Cube Cloud.
After you've created your [Cube Cloud account][cube-cloud-signup],
click **+ Create Deployment** in the top-right corner of your
list of deployments to jump into the wizard. Watch the following video for
guidance on the rest of the steps:
All deployments within a Cube Cloud account should have uniqie names.
## Demo deployments
After you've created your [Cube Cloud account][cube-cloud-signup] and
proceeded to create a new deployment, you would be prompted to set up a
connection to your [data source][ref-data-source].
If, for any reason, you're not ready or unable to connect your staging or
production data source at the moment, you can click **Create a demo
deployment** in the yellow box:
Shortly, Cube Cloud will create and pre-configure a new deployment in your
account. You will see it as **Demo deployment** in the list of
deployments. This new deployment will be:
* Configured to use [DuckDB][ref-duckdb] as the data source.
* Provided with a sample [data model][ref-data-model] backed with CSV files
in a public S3 bucket.
* Provided with an example of [dynamic data modeling][ref-dynamic-data-models]
and programmatic [configuration options][ref-config-options].
Watch the following video to for a step-by-step walkthrough:
Exploring the demo deployment is a great way to understand how Cube works.
Use the [data model editor][ref-data-model] to review the data model and
make sure to run a few queries in [Playground][ref-playground].
## Deployment overview
The **Overview** page of each deployment provides a high-level
summary of its components and state:
* API endpoints with URLs and connection instructions.
* Allocated resources in line with the [deployment type][ref-deployment-types].
* Activity log of the most recent events.
Additionally, the bar under the **Cube Cloud** logo displays
user-specific state of the deployment:
* Branch selector displays the source code branch that is currently
selected and being viewed by the user.
* **Enter Development Mode** button indicates whether a user has
entered the [development mode][ref-dev-mode].
## Deployment settings
You can manage various settings of a deployment, such as [environment
variables](#environment-variables), by navigating to the **Settings** page of the
deployment.
### Environment variables
You can configure environment variables by navigating to **Settings → Environment
variables** where you can add, edit, or remove environment variables for your
deployment.
For convenience, environment variables are split into two lists: ones that relate to
the [data source][ref-data-sources] configuration and all other variables.
Browsing environment variables is reflected in the [Audit Log][ref-audit-log].
## Cube version
Each [environment][ref-environments] within a Cube Cloud deployment runs a specific
version of Cube that depends on the environment type and the selected [update
channel](#update-channels).
### Current version
You can check the current version of Cube that your deployment (environment) runs by
navigating to **Overview → Resources & Logs → Cube API**:
### Update channels
Cube Cloud provides two update channels:
* **LTS channel** for infrequent [long-term support][ref-lts] releases with
critical fixes.
* **Regular channel** for regular releases, which occur more frequently and
contain both fixes and new features.
You can view or change the update channel by navigating to **Settings →
General → Cube version**:
You can select a specific version in the drop-down. Only versions that have been used by
a deployment during the last 6 months are available in the list.
There's an option to **Upgrade automatically to new patch versions**, which is
available for both update channels. When it's turned on, an upgrade will occur if a new
[patch version][link-semver] is available in the selected channel. To trigger an upgrade,
deploy a new build to the [production environment][ref-environments-prod] or change its
settings, e.g., by updating environment variables of the Cube Cloud deployment.
Generally, it's recommended to use the *regular channel* to get the latest updates.
You can *pin the version* to a specific one by turning the **Upgrade automatically...**
toggle off.
If you choose to pin the version, make sure to upgrade periodically, especially when the
next [minor version][ref-distribution-versions] is released. If you use the *LTS channel*,
make sure to upgrade to the latest [LTS version][ref-lts] every few months.
## Resource consumption
Cube Cloud deployments only consume resources when they are needed to run workloads:
* For the [production environment][ref-environments-prod], resources are always consumed
unless a deployment is [suspended][ref-auto-sus].
* For any other [environment][ref-environments], a Shared deployment is allocated
while it's active. After a period of inactivity, the Shared deployment is deallocated.
* For pre-aggregations, Cube Store workers are allocated while there's some
activity related to pre-aggregations, e.g., API endpoints are serving
requests to pre-aggregations, pre-aggregations are being built, etc.
After a period of inactivity, Cube Store workers are deallocated.
Please refer to [total cost examples][ref-total-cost] to learn more about
resource consumption in different scenarios.
[cube-cloud-signup]: https://cubecloud.dev/auth/signup
[ref-deployment-types]: /admin/deployment/deployment-types
[ref-dev-mode]: /docs/data-modeling/dev-mode
[ref-auto-sus]: /admin/deployment/auto-suspension
[ref-total-cost]: /admin/account-billing/pricing#total-cost-examples
[ref-data-source]: /admin/connect-to-data/data-sources
[ref-duckdb]: /admin/connect-to-data/data-sources/duckdb
[ref-data-model]: /docs/data-modeling/overview
[ref-dynamic-data-models]: /docs/data-modeling/dynamic
[ref-config-options]: /admin/connect-to-data#configuration-options
[ref-data-model]: /docs/data-modeling/data-model-ide
[ref-playground]: /docs/explore-analyze/playground
[ref-environments]: /admin/deployment/environments
[ref-environments-prod]: /admin/deployment/environments#production-environment
[ref-lts]: /admin/account-billing/distribution#long-term-support
[link-semver]: https://semver.org
[ref-data-sources]: /admin/connect-to-data/data-sources
[ref-audit-log]: /admin/monitoring/audit-log
[ref-distribution-versions]: /admin/account-billing/distribution#versions
# Infrastructure Options
Source: https://docs.cube.dev/admin/deployment/infrastructure
Cube Cloud provides four infrastructure options to host your Cube deployments:
* [Multi-tenant infrastructure](#shared-infrastructure) - your
deployments share compute resources and network with other customers.
Data in-motion and data at-rest are both on the Cube Cloud side.
* [Single-tenant infrastructure](#dedicated-infrastructure) - your
deployments reside in a dedicated VPC inside a Cube Cloud account and do not
share resources with anyone else. Data in-motion and data at-rest are both on
the Cube Cloud side.
* [Single-tenant infrastructure with CSPS](#dedicated-infrastructure-with-csps) -
same as single-tenant infrastructure, but data at-rest is stored in a customer-supplied
object store.
* [Bring Your Own Cloud (BYOC)](#byoc) - Cube Cloud data plane is fully hosted
in your cloud account.
## Multi-tenant infrastructure
This is the most common deployment option that is the easiest to get started with.
In this scenario, everything is deployed on the Cube Cloud infrastructure in
one of our **multi-tenant** VPCs. Cube Cloud Control Plane takes care of creating,
scaling, and monitoring your Cube Deployments, as well as managing Cube Store
and persisting pre-aggregated data. This option requires the least effort to
set up.
Please note that some Enterprise features, such as VPC peering or PrivateLink are
not available on the multi-tenant infrastructure. There's also a possibility of
resource contention ("noisy neighbor") problem.
## Single-tenant infrastructure
It is similar to the previous option, but each customer gets a
[Dedicated][ref-dedicated-vpc] VPC within one of Cube Cloud's own cloud
accounts that hosts only that customer's deployments. This option is great for
most of the typical Enterprise use-cases as it provides a higher level of
performance, as well as additional security and isolation.
Available as an add-on on the [Enterprise plan](https://cube.dev/pricing).
## Single-tenant infrastructure with CSPS
Cube Cloud offers a **customer-supplied pre-aggregation storage (CSPS)** that
allows moving all data at rest to the customer
infrastructure. In this scenario, all Cube components reside on the Cube Cloud
side. However, Cube Store uses a customer-provided object store for reading and
persisting pre-aggregated data. This provides additional peace of mind when
processing highly critical business or personal information.
Available on the [Enterprise plan](https://cube.dev/pricing) with the
[Single-tenant infrastructure](#dedicated-infrastructure) add-on.
## BYOC
With [Bring Your Own Cloud](/admin/deployment/dedicated) (BYOC) all the components interacting with private data are deployed on the customer infrastructure
on a platform of choice (AWS/Azure/GCP) and managed by the Cube Cloud Control Plane via the Cube Cloud Operator.
Available as an add-on on the [Enterprise plan](https://cube.dev/pricing).
## Understanding "Cube Cloud Region"
Throughout Cube documentation and when interacting with Cube staff, you may encounter the term **"Cube Cloud Region"** or simply **"Region"**. Understanding this term is crucial for properly configuring and managing your Cube deployments.
### What is a Cube Cloud Region?
A **Cube Cloud Region** refers to a specific cloud infrastructure instance used to host your Cube deployments. While it includes a geographical location component, it encompasses much more than just a physical data center location.
Each Cube Cloud Region is identified by a unique identifier that contains several components:
* **Cloud provider** (AWS, GCP, or Azure)
* **Geographical region** (e.g., us-east-1, eu-west-1)
* **Infrastructure type** (multi-tenant, single-tenant, or BYOC)
* **Tenant identifier** (for single-tenant infrastructure)
* **Environment** (e.g., prod, staging)
For example, a region identifier might look like:
* `aws-us-east-1-shared` for multi-tenant infrastructure in AWS US East
* `aws-us-east-1-t-12345-prod` for single-tenant infrastructure with tenant ID 12345
* `gcp-europe-west1-t-12345-byoc` for a BYOC deployment in GCP Europe
Each region also has a human-readable display name that's visible in the Cube Cloud UI. For single-tenant infrastructure regions, these display names typically include the customer name and/or environment name (e.g., "Acme Corp Production (N. Virginia)" or "Acme Corp Staging (Iowa)") to help distinguish between different infrastructure instances.
### Not to be confused with...
The term "Cube Cloud Region" should **not** be confused with:
* **Cloud provider regions alone** (like AWS `us-east-1`) - A Cube Cloud Region includes but is not limited to the underlying cloud provider region
* **Geographical regions** - While geography is a component, the Cube Cloud Region encompasses infrastructure type and tenant isolation as well
* **Availability zones** - These are subdivisions within cloud provider regions and are handled transparently by Cube Cloud
### Why this matters
Understanding your Cube Cloud Region is important for:
1. **API endpoints**: Your deployment's API endpoints include the region identifier (e.g., `..cubecloudapp.dev`)
2. **Network configuration**: When setting up VPC peering, PrivateLink, or custom domains, you'll need the exact region identifier
3. **Support requests**: Providing the correct region identifier helps Cube support team quickly locate and assist with your deployment
4. **Infrastructure planning**: Different region types offer different capabilities (e.g., PrivateLink is only available in single-tenant and BYOC regions)
### Finding your region identifier
You can find your Cube Cloud Region identifier in:
* The Cube Cloud UI deployment settings
* API endpoint URLs provided in the deployment overview
* Communication from Cube Cloud support when your infrastructure is provisioned
When in doubt, contact Cube Cloud support with your deployment ID, and they can provide the exact region identifier for your infrastructure.
[ref-dedicated-vpc]: /admin/deployment/dedicated
# Limits and quotas
Source: https://docs.cube.dev/admin/deployment/limits
Cube Cloud implements limits on resource usage on account and deployment levels to ensure the best experience for all Cube Cloud users.
Applies to [all plans](https://cube.dev/pricing).
## Limit types
Each limit can be of one of the following types:
* *Hard limit* either has a threshold that can't be exceeded or prevents further
use of a resource when a threshhold is hit until a cool-down period passes
(e.g., until the next day starts) or the limit is increased (e.g., when a Cube
Cloud account upgrades to another tier).
* *Soft limit* may allow further use of a resource after a threshold is hit.
## Resources
The following resources are subject to limits, depending on [deployment
types][ref-deployment-types] and [product tiers][ref-pricing]:
| Resource | Free Tier | Starter | Premium | Enterprise |
| ---------------------------------------------------------------------------------- | :-------: | :--------------------------------------------: | :--------------------------------------------: | :---------------------------------------------------: |
| Number of deployments | 2 | Unlimited | Unlimited | Unlimited |
| Number of API instances per deployment | 1 | 10 | 10 | [Contact us][cube-contact-us] |
| Number of Cube Store workers per deployment | 2 | 16 | 16 | [Contact us][cube-contact-us] |
| Queries per day for each [Shared deployment][ref-dev-instance] | 1,000 | 10,000 | Unlimited | Unlimited |
| Queries per day for each [Dedicated deployment][ref-prod-cluster] | — | 50,000 | Unlimited | Unlimited |
| [Query History][ref-query-history] — retention period | 1 day | Depends on the [tier][ref-query-history-tiers] | Depends on the [tier][ref-query-history-tiers] | Depends on the [tier][ref-query-history-tiers] |
| [Query History][ref-query-history] — queries processed per day for each deployment | 1,000 | Depends on the [tier][ref-query-history-tiers] | Depends on the [tier][ref-query-history-tiers] | Depends on the [tier][ref-query-history-tiers] |
| [Audit Log][ref-audit-log] — retention period | — | — | — | 30 days |
| [Audit Log][ref-audit-log] — events collected | — | — | — | 10,000 |
| [Monitoring Integrations][ref-monitoring-integrations] — exported data | — | — | — | Available as an [add-on][ref-monitoring-integrations] |
### Number of deployments
This is a hard limit. Consider upgrading to [another tier][ref-pricing].
### Number of API instances
This is a hard limit. Consider upgrading to [another tier][ref-pricing].
### Number of Cube Store workers
This is a hard limit. Consider upgrading to [another tier][ref-pricing].
An option to use 32 Cube Store workers is subject to availability in select
regions. Please [contact support][cube-contact-us] for more details.
### Queries per day
This is a hard limit. Usage is calculated per Cube Cloud account, i.e., in total
for all deployments within an account.
When a threshold is hit, further queries will not be processed. In that case,
consider upgrading a Shared deployment to a Dedicated deployment.
Alternatively, consider upgrading to [another tier][ref-pricing].
### Queries processed by Query History per day
This is a soft limit. Usage is calculated per Cube Cloud account, i.e., in total
for all deployments within an account.
When a threshold is hit, query processing will be stopped. Please [contact
support][cube-contact-us] for further assistance.
### Data retained by Query History
This is a soft limit. Usage is calculated per Cube Cloud account, i.e., in total
for all deployments within an account.
### Data collected and retained by Audit Log
This is a hard limit. Usage is calculated per Cube Cloud account, i.e., in total
for all deployments within an account. Please [contact
support][cube-contact-us] for further assistance.
### Data exported via Monitoring Integrations
This is a hard limit. Usage is calculated per Cube Cloud deployment. [Contact
support][cube-contact-us] for higher quotas.
## Quotas
The [REST (JSON)][ref-rest-api] and [GraphQL][ref-gql-api] APIs both have a standard
quota of 100 requests per second per deployment; this can also go higher
for short bursts of traffic. These limits can be raised on request,
[contact support][cube-contact-us] for more details.
When the quota is exceeded, the API will return a `429 Too Many Requests`
response.
[ref-rest-api]: /reference/core-data-apis/rest-api
[ref-gql-api]: /reference/core-data-apis/graphql-api
[ref-deployment-types]: /admin/deployment/deployment-types
[ref-pricing]: /admin/account-billing/pricing
[ref-query-history]: /admin/monitoring/query-history
[ref-monitoring-integrations]: /admin/monitoring/monitoring-integrations
[ref-dev-instance]: /admin/deployment/deployment-types#shared
[ref-prod-cluster]: /admin/deployment/deployment-types#dedicated
[cube-contact-us]: https://cube.dev/contact
[ref-query-history-tiers]: /admin/account-billing/pricing#query-history-tiers
[ref-audit-log]: /admin/monitoring/audit-log
# AWS
Source: https://docs.cube.dev/admin/deployment/oidc/aws
Configure AWS IAM trust for Cube's OIDC issuer and use it for Athena, Redshift, S3 export buckets, Cube Store CSPS, and Bedrock.
This guide walks through configuring AWS to trust Cube's OIDC issuer and
shows the trust policies for the most common targets — Athena, Redshift, an
S3 export bucket, Cube Store CSPS, and Bedrock for bring-your-own LLM.
If you haven't enabled OIDC for your tenant yet, start with the
[OIDC overview][ref-oidc-overview].
Available on the [Enterprise plan](https://cube.dev/pricing).
## Prerequisites
* The Cube tenant has OIDC enabled and an `AWS` token config exists under
**Admin → OIDC**.
* IAM access to your AWS account sufficient to register an IAM OIDC provider
and create / update IAM roles.
* Your tenant slug — the leftmost label of your tenant's console URL.
Throughout this guide it's referenced as `` (and the full
issuer URL as `https://.cubecloud.dev`). Substitute your
actual slug everywhere it appears.
The trust policies, env vars, and CLI commands in this guide use angle-bracket
placeholders — ``, ``, ``,
etc. **Replace each placeholder with your real value** before
copying. AWS will accept these strings literally and the federation call will
fail with a confusing error.
## Step 1: Register Cube as an OIDC provider in AWS
This is a **one-time** setup per AWS account. Once registered, every IAM role
in this account can be configured to trust deployments in your Cube tenant.
```bash theme={"dark"}
aws iam create-open-id-connect-provider \
--url https://.cubecloud.dev \
--client-id-list sts.amazonaws.com
```
AWS no longer requires a TLS thumbprint for HTTPS OIDC providers. The
`--thumbprint-list` parameter is accepted for compatibility but ignored —
AWS validates the issuer's certificate chain against its own trust store.
The command returns the provider ARN, which looks like:
```
arn:aws:iam:::oidc-provider/.cubecloud.dev
```
You'll reference this ARN as the `Federated` principal in every trust policy
below. From AWS's point of view, this provider *is* your Cube tenant.
## Step 2: Set the deployment identity
Add `AWS_ROLE_ARN` to your deployment's environment variables under
**Settings → Environment variables**. This is the IAM role Cube assumes by
default for every AWS SDK call inside the deployment — drivers, export
bucket I/O, custom code in `cube.py` / `cube.js`. You can either grant this
role direct access to your data, or use it as the entry point for further
`AssumeRole` hops.
```dotenv theme={"dark"}
AWS_ROLE_ARN=arn:aws:iam:::role/cube-deployment-
```
`sts:AssumeRoleWithWebIdentity` does **not** accept `sts:ExternalId`. Trust
policies for OIDC-federated roles can only condition on the standard OIDC
claims (`aud`, `sub`, `iss`). If you copy a trust policy from a non-federated
role assumption (cross-account access keys, for example) and it includes
`sts:ExternalId`, remove it — STS will reject the federation call. Use the
`sub` claim to pin the role to a specific deployment or component instead.
## Step 3: Build the trust policy
Every IAM role you want Cube to assume needs a trust policy with this shape:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam:::oidc-provider/.cubecloud.dev"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
".cubecloud.dev:aud": "sts.amazonaws.com"
},
"StringLike": {
".cubecloud.dev:sub": "cube:deployment::component:cube_api"
}
}
}
]
}
```
The `aud` condition pins the audience to AWS STS — exactly what Cube's `aws`
token config emits. The `sub` condition is what scopes the trust to a
specific deployment (and optionally a specific Cube component). Patterns:
| Trust scope | `sub` pattern |
| ----------------------------------------------------------------------- | ------------------------------------------------------ |
| One specific deployment, any component | `cube:deployment::component:*` |
| One deployment, only the Cube API and refresh worker | `cube:deployment::component:cube_api` |
| One deployment, only Cube Store | `cube:deployment::component:cube_store` |
| Every deployment in the tenant, only Cube Store (e.g. tenant-wide CSPS) | `cube:deployment:*:component:cube_store` |
| Every deployment, every component | `cube:deployment:*:component:*` |
Cube's default `sub` claim is `cube:deployment:`. To match
the `:component:` patterns in the table above (or to add
`:region:`), open your AWS token config in **Admin → OIDC** and
paste one of these templates into the **Subject Claim Format** field:
* `cube:deployment:{deployment_id}:component:{component}` — for the patterns
in the table above.
* `cube:deployment:{deployment_id}:component:{component}:region:{region}` —
to additionally pin a [Cube Cloud region][ref-cube-cloud-region], useful
when you have dedicated regions per environment.
See [the subject editor section][ref-sub-editor] for the full syntax.
Update your AWS trust policy first, then change the **Subject Claim Format**
on the token config — otherwise existing tokens won't match the trust policy
and `AssumeRoleWithWebIdentity` will start failing.
## Athena
Configure an IAM role with permissions to query Athena and read query
results from your S3 results bucket.
Trust policy — substitute your AWS account ID, tenant slug, and
deployment ID for the angle-bracket placeholders:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam:::oidc-provider/.cubecloud.dev"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
".cubecloud.dev:aud": "sts.amazonaws.com"
},
"StringLike": {
".cubecloud.dev:sub": "cube:deployment::component:cube_api"
}
}
}
]
}
```
Athena needs permission to start queries, read Glue metadata, and
read / write the S3 results bucket. The example below scopes the S3
permissions to a single bucket; tighten the resource list as needed.
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"athena:StartQueryExecution",
"athena:GetQueryExecution",
"athena:GetQueryResults",
"athena:StopQueryExecution",
"athena:ListQueryExecutions",
"athena:GetWorkGroup",
"athena:ListWorkGroups"
],
"Resource": "*"
},
{
"Effect": "Allow",
"Action": [
"glue:GetDatabase",
"glue:GetDatabases",
"glue:GetTable",
"glue:GetTables",
"glue:GetPartitions"
],
"Resource": "*"
},
{
"Effect": "Allow",
"Action": [
"s3:GetBucketLocation",
"s3:GetObject",
"s3:ListBucket",
"s3:PutObject",
"s3:DeleteObject",
"s3:AbortMultipartUpload",
"s3:ListMultipartUploadParts"
],
"Resource": [
"arn:aws:s3:::my-athena-results",
"arn:aws:s3:::my-athena-results/*",
"arn:aws:s3:::my-athena-data",
"arn:aws:s3:::my-athena-data/*"
]
}
]
}
```
Set the deployment-level Athena env vars. With `AWS_ROLE_ARN` in place,
the Athena driver automatically assumes the role via OIDC federation —
no static credentials needed.
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=athena
AWS_ROLE_ARN=arn:aws:iam:::role/cube-deployment-
CUBEJS_AWS_REGION=us-east-1
CUBEJS_AWS_S3_OUTPUT_LOCATION=s3://my-athena-results/queries/
```
If Athena lives in a different role / account from the deployment's
default identity, set [`CUBEJS_AWS_ATHENA_ASSUME_ROLE_ARN`][ref-athena-assume-role]
in addition to `AWS_ROLE_ARN`. The Athena driver uses the deployment
identity to perform a second `AssumeRole` hop into the Athena role.
## Redshift
Configure an IAM role that can obtain temporary Redshift database
credentials. The Redshift driver uses the deployment's federated identity
to call `redshift:GetClusterCredentialsWithIAM` — no database password is
stored anywhere.
Use the same trust policy shape as [Athena](#athena) — federated
principal, `aud` pinned to `sts.amazonaws.com`, and `sub` pinned to your
deployment.
The role needs to fetch temporary credentials and describe the cluster:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"redshift:GetClusterCredentialsWithIAM",
"redshift:DescribeClusters"
],
"Resource": "*"
}
]
}
```
Tighten `Resource` to your cluster and database ARNs in production.
Set the Redshift driver env vars for [IAM
authentication][ref-redshift-driver-iam] and omit `CUBEJS_DB_USER` /
`CUBEJS_DB_PASS`. With `AWS_ROLE_ARN` in place, the driver assumes the
role via OIDC federation and exchanges it for temporary database
credentials:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=redshift
AWS_ROLE_ARN=arn:aws:iam:::role/cube-deployment-
CUBEJS_DB_HOST=my-cluster.xxx.us-east-1.redshift.amazonaws.com
CUBEJS_DB_NAME=my_database
CUBEJS_DB_SSL=true
CUBEJS_DB_REDSHIFT_AWS_REGION=us-east-1
CUBEJS_DB_REDSHIFT_CLUSTER_IDENTIFIER=my-cluster
```
## S3 export bucket
If your data source uses an [export bucket][ref-export-bucket] for
pre-aggregation unloads (Snowflake, Redshift, Athena, BigQuery, …), Cube
needs `s3:PutObject` / `s3:GetObject` / `s3:ListBucket` on the bucket. The
deployment's default identity is the simplest place to put this.
Add an inline statement to the policy attached to your deployment's
`AWS_ROLE_ARN`:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"s3:GetObject",
"s3:PutObject",
"s3:DeleteObject",
"s3:ListBucket",
"s3:GetBucketLocation",
"s3:AbortMultipartUpload",
"s3:ListMultipartUploadParts"
],
"Resource": [
"arn:aws:s3:::my-export-bucket",
"arn:aws:s3:::my-export-bucket/*"
]
}
]
}
```
Set the export bucket env vars on the deployment — leave the AWS access
key vars empty so the SDK falls back to the OIDC-derived credentials:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=s3
CUBEJS_DB_EXPORT_BUCKET=my-export-bucket
CUBEJS_DB_EXPORT_BUCKET_AWS_REGION=us-east-1
```
See the [export bucket reference][ref-export-bucket] for the full set
of variables.
OIDC only covers Cube's **read** side of the export bucket. The data
warehouse itself (Snowflake, Redshift, Athena, BigQuery, …) runs the
`UNLOAD` that writes objects to the bucket, and the warehouse cannot
federate with Cube's OIDC issuer. You still need to provide **separate
credentials for the `UNLOAD`** so the warehouse can write to S3 — typically
an AWS access key pair or a warehouse-side storage integration / IAM role
— via the standard export bucket env vars (e.g.
`CUBEJS_DB_EXPORT_BUCKET_AWS_KEY` and `CUBEJS_DB_EXPORT_BUCKET_AWS_SECRET`,
or the driver-specific storage-integration variables). OIDC then handles
Cube's download of the unloaded objects from the bucket.
## Cube Store CSPS bucket
Cube Store CSPS lets you store pre-aggregations in your own S3 bucket.
Cube Store gets a separate OIDC token whose `sub` claim ends in
`component:cube_store`, so the trust policy can be locked down to that
component — even if the same role were ever shared with the rest of the
deployment, only Cube Store would be able to assume it.
Because every Cube Store worker emits a `sub` of the form
`cube:deployment::component:cube_store`, the trust policy's
`StringLike` condition controls how broadly the role is shared:
* `cube:deployment:*:component:cube_store` — **one role + one bucket for the
whole tenant.** Every deployment in the tenant writes pre-aggregations to
the same bucket, isolated only by Cube Store's own per-deployment path
prefix. Easiest to operate and the most common setup.
* `cube:deployment::component:cube_store` — **per-deployment
isolation.** Pin the role to one deployment so its pre-aggregations live
in a dedicated bucket that no other deployment can read or write.
The example below shows the tenant-wide pattern; swap `*` for a specific
deployment ID if you want isolation.
Trust policy — note the `StringLike` condition pinning the component
and using `*` so every deployment in the tenant can assume the role:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam:::oidc-provider/.cubecloud.dev"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
".cubecloud.dev:aud": "sts.amazonaws.com"
},
"StringLike": {
".cubecloud.dev:sub": "cube:deployment:*:component:cube_store"
}
}
}
]
}
```
Allow the role to read, write, and list objects in your CSPS bucket.
Cube Store needs all of the actions below — including the multipart
upload primitives — to handle large pre-aggregation partitions:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"s3:GetObject",
"s3:PutObject",
"s3:DeleteObject",
"s3:ListBucket",
"s3:GetBucketLocation",
"s3:AbortMultipartUpload",
"s3:ListMultipartUploadParts"
],
"Resource": [
"arn:aws:s3:::my-csps-bucket",
"arn:aws:s3:::my-csps-bucket/*"
]
}
]
}
```
For each deployment that should use this bucket, go to **Settings →
Pre-Aggregation Storage** on the deployment and:
* Toggle **Enable CSPS** on.
* **Storage Provider**: Amazon S3.
* **S3 Bucket**: `my-csps-bucket`.
* **S3 Region**: e.g. `us-east-1`.
* **IAM Role ARN**: `arn:aws:iam:::role/cube-cubestore-`.
Click **Test Connection** to verify Cube Store can assume the role and
access the bucket, then **Apply**. Cube Store starts writing
pre-aggregations to your bucket on the next refresh. With the
tenant-wide trust policy above, every deployment in the tenant points
at the same role — no need to provision a new IAM role per deployment.
## Bedrock for bring-your-own LLM
Bring-your-own LLM lets the AI engineer service call Bedrock through your
own AWS account. The AI engineer's `sub` claim ends in
`component:ai_engineer`, so the trust policy uses a `StringLike` match on
`cube:deployment:*:component:ai_engineer` to grant access tenant-wide
(every deployment's AI engineer assumes the same role and writes against
the same Bedrock account).
Trust policy — note the `StringLike` condition pinning the component:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam:::oidc-provider/.cubecloud.dev"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
".cubecloud.dev:aud": "sts.amazonaws.com"
},
"StringLike": {
".cubecloud.dev:sub": "cube:deployment:*:component:ai_engineer"
}
}
}
]
}
```
Grant access only to the foundation models and inference profiles you
intend to use:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"bedrock:InvokeModel",
"bedrock:InvokeModelWithResponseStream",
"bedrock:Converse",
"bedrock:ConverseStream"
],
"Resource": [
"arn:aws:bedrock:*::foundation-model/anthropic.*",
"arn:aws:bedrock:*::inference-profile/*"
]
}
]
}
```
Cross-region inference profiles need access to the foundation model in
every region the profile routes through, hence the `*` region in the
foundation-model ARN.
Under **Admin → AI → Models**, click **Add Model** and configure:
* **Name** — a human-readable label for this BYOM entry.
* **Model Type** — `LLM`.
* **Provider** — `AWS Bedrock`.
* **Model** — the Claude (or other) model you want the AI engineer to
call.
* **Region** — the Bedrock region (e.g. `us-east-1`).
* **Inference Profile ID** — optional; leave blank to call the
foundation model directly, or set to a cross-region inference
profile ID.
* **Assume Role ARN** — the role you created above.
* **Use OIDC workload identity** — toggle on. With this on, the AI
engineer authenticates via the deployment's OIDC token instead of
static AWS credentials.
## Monitoring integrations (CloudWatch, S3)
[Monitoring integrations][ref-monitoring-integrations] can ship deployment logs
to AWS without a static access key pair. The Vector agent that performs the
export reads the deployment's OIDC token and assumes an IAM role, exactly like
the targets above — so the `[sinks.*.auth]` block disappears from `vector.toml`
entirely.
The credentials are resolved by the AWS SDK's default chain **for the whole
Vector agent**, not per sink, so one role serves every AWS sink you configure —
[`aws_cloudwatch_logs`][ref-cloudwatch], [`aws_s3`][ref-s3], and the `aws_s3`
sink used by [Query History export][ref-query-history-export]. Grant that single
role whatever those sinks need.
Vector authenticates as the deployment's **`cube_api`** component, so the trust
policy is the standard one from [Step 3](#step-3-build-the-trust-policy) with
`sub` set to `cube:deployment::component:cube_api`.
That `sub` uses the `:component:` form, which is **not** Cube's default — the
default claim is `cube:deployment:` with no `component` segment,
and a trust policy written against the pattern above will never match it. Set
**Subject Claim Format** on your `AWS` token config to
`cube:deployment:{deployment_id}:component:{component}` first; see
[Step 3](#step-3-build-the-trust-policy) for the full syntax and the ordering
caveat when changing it on a config that is already in use.
Create a role with that trust policy and attach a permissions policy that
can write to your log group:
```json theme={"dark"}
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": [
"logs:CreateLogStream",
"logs:PutLogEvents",
"logs:DescribeLogStreams"
],
"Resource": [
"arn:aws:logs:::log-group:",
"arn:aws:logs:::log-group::*"
]
},
{
"Effect": "Allow",
"Action": ["logs:DescribeLogGroups"],
"Resource": "*"
}
]
}
```
Add `logs:CreateLogGroup` to the first statement if you set
`create_missing_group = true` in the sink.
If the same agent also writes to S3 (the [S3][ref-s3] or
[Query History export][ref-query-history-export] sinks), add the bucket
permissions those need to this same role: `s3:PutObject` on
`arn:aws:s3:::/*`, plus `s3:ListBucket` on `arn:aws:s3:::`
if you enable the sink's healthcheck. Both S3 examples in the docs ship with
`[sinks.aws_s3.healthcheck] enabled = false`, so `s3:PutObject` alone is
enough as written — but Vector's `aws_s3` healthcheck issues a `HeadBucket`,
which AWS authorizes as `s3:ListBucket` on the bucket ARN (not `/*`).
Set these on the deployment under **Settings → Environment variables**:
```dotenv theme={"dark"}
CUBE_CLOUD_MONITORING_AWS_ROLE_ARN=arn:aws:iam:::role/
```
Setting the role ARN is what enables keyless authentication. The region for
the credential exchange is taken from the sink's own `region`; set
`CUBE_CLOUD_MONITORING_AWS_REGION` only to override it.
Then drop the `auth` block from your `vector.toml` — see
[Integration with Amazon CloudWatch][ref-cloudwatch] for the sink
configuration.
**`logs:DescribeLogGroups` must be scoped to `"*"`.** Vector's healthcheck calls
it, and it is an account-level list operation that AWS evaluates against an
*empty* resource. Scoping it to your log group ARN — the natural
least-privilege instinct — denies it, and the sink is marked unhealthy at
startup with:
```
not authorized to perform: logs:DescribeLogGroups on resource:
arn:aws:logs:::log-group::log-stream:
```
Keep it in a separate statement from the log-group-scoped write actions, as
shown above.
This role is separate from the deployment's default `AWS_ROLE_ARN`. Because
both are assumed by the same `cube_api` identity, a trust policy written for
one will also admit the other — scope each role's **permissions** policy
tightly rather than relying on the trust policy to separate them.
## Verifying the setup
The fastest way to confirm the trust policy is wired up correctly is the
**Test connection** button on the relevant settings page (data source
wizard, CSPS settings, BYO LLM provider). Behind the scenes, this issues a
real Cube OIDC token, runs `AssumeRoleWithWebIdentity` against AWS STS, and
returns a precise error if the trust policy rejects it.
If the test fails:
| Symptom | Likely cause |
| --------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `Not authorized to perform sts:AssumeRoleWithWebIdentity` | The trust policy's `sub` condition doesn't match the deployment / component. Compare the `sub` in the error with what your `StringLike` allows. |
| `Incorrect token audience` | `aud` condition is wrong, or you copied a trust policy from a different tenant. The condition key must be `.cubecloud.dev:aud`. |
| `OpenIDConnectProvider not found` | The OIDC provider hasn't been registered in this AWS account yet, or its URL doesn't match your tenant's domain. Re-run `aws iam create-open-id-connect-provider` with the correct URL. |
| `An error occurred ... InvalidIdentityToken` | Trying to use `sts:ExternalId`. Remove it — federated assume-role doesn't accept it. |
`AssumeRoleWithWebIdentity` events show up in CloudTrail with the deployment
subject in the `userIdentity.webIdFederationData.federatedProvider` and
`...attributes` fields — useful for auditing which deployments are
authenticating against which roles.
[ref-oidc-overview]: /admin/deployment/oidc
[ref-sub-editor]: /admin/deployment/oidc#subject-claim-format
[ref-cube-cloud-region]: /admin/deployment/infrastructure#what-is-a-cube-cloud-region
[ref-export-bucket]: /admin/connect-to-data#export-bucket
[ref-monitoring-integrations]: /admin/monitoring/monitoring-integrations
[ref-cloudwatch]: /admin/monitoring/monitoring-integrations/cloudwatch
[ref-s3]: /admin/monitoring/monitoring-integrations/s3
[ref-query-history-export]: /admin/monitoring/query-history-export
[ref-athena-assume-role]: /admin/connect-to-data/data-sources/aws-athena#environment-variables
[ref-redshift-driver-iam]: /admin/connect-to-data/data-sources/aws-redshift#iam-authentication
# Azure
Source: https://docs.cube.dev/admin/deployment/oidc/azure
Configure an Azure AD app registration with federated credentials that trust Cube's OIDC issuer, then use it for Azure SQL and Blob Storage.
This guide walks through configuring Azure to trust Cube's OIDC issuer using
**federated credentials** on an Azure AD app registration, and shows the
setup for the most common targets — Azure SQL and an Azure Blob Storage
export bucket.
If you haven't enabled OIDC for your tenant yet, start with the
[OIDC overview][ref-oidc-overview].
Available on the [Enterprise plan](https://cube.dev/pricing).
## Prerequisites
* The Cube tenant has OIDC enabled and an `Azure` token config exists under
**Admin → OIDC**.
* Permissions in your Azure AD tenant sufficient to create an app
registration and add federated credentials, plus an Azure subscription
where you can grant the app role assignments on resources.
* Your Cube tenant slug — the leftmost label of your tenant's console URL.
Throughout this guide it's referenced as `` (and the full
issuer URL as `https://.cubecloud.dev`). Substitute your
actual slug everywhere it appears.
Commands and federated-credential snippets in this guide use angle-bracket
placeholders — ``, ``, etc.
**Replace each placeholder with your real value** before running. Azure
will accept these strings literally and the federation call will fail with
a confusing error.
## How Azure federation works
Azure doesn't have a generic OIDC provider registration the way AWS does.
Instead, you configure each app registration (or user-assigned managed
identity) with a **federated credential** that pins the issuer URL,
expected `aud`, and exact `sub` value. When Cube presents a JWT that
matches all three, Azure AD swaps it for an access token scoped to the
app registration.
```mermaid theme={"dark"}
sequenceDiagram
participant Cube as Cube Deployment
participant AAD as Azure AD (login.microsoftonline.com)
participant Resource as Azure Resource (SQL / Blob / OpenAI)
Cube->>AAD: client_assertion = OIDC JWT grant_type = client_credentials audience = api://AzureADTokenExchange
AAD->>AAD: Match federated credential by issuer + sub + aud Validate JWT against Cube's JWKS
AAD->>Cube: Access token for the app registration
Cube->>Resource: Authenticated request
```
Azure caps you at **20 federated credentials per app registration**. For
tenants with many deployments, see [Scaling past 20
credentials](#scaling-past-20-federated-credentials) below.
## Step 1: Create or pick an app registration
You can use an existing app registration or create a new one — Cube doesn't
have any opinion as long as it has the federated credential and the role
assignments.
```bash theme={"dark"}
az ad app create --display-name "Cube Cloud Deployment"
```
Note the **Application (client) ID** — this is your `AZURE_CLIENT_ID`. The
**Directory (tenant) ID** is your `AZURE_TENANT_ID`. Both are visible in
the Azure portal under **Microsoft Entra ID → App registrations → Cube
Cloud Deployment → Overview**.
## Step 2: Add a federated credential
The federated credential is what binds the app registration to a specific
Cube subject. Each credential matches **exactly one** issuer + subject +
audience triple — Azure doesn't support wildcards or pattern matching here.
```bash theme={"dark"}
az ad app federated-credential create \
--id \
--parameters '{
"name": "cube--deployment-",
"issuer": "https://.cubecloud.dev",
"subject": "cube:deployment::component:cube_api",
"audiences": ["api://AzureADTokenExchange"]
}'
```
| Field | Value |
| ----------- | ----------------------------------------------------------------------------------------------------- |
| `issuer` | Your Cube tenant's URL: `https://.cubecloud.dev`. |
| `subject` | The exact `sub` claim Cube emits — typically `cube:deployment::component:`. |
| `audiences` | Always `["api://AzureADTokenExchange"]` — this is the standard Azure AD token-exchange audience. |
| `name` | A human-readable label. Pick something that lets you find this credential later. |
Each Cube component you want this app to authenticate as needs its own
federated credential. So if a deployment runs both Cube API and Cube Store
against the same Azure resource, you create two credentials — one with
`subject` ending in `:component:cube_api` and one ending in
`:component:cube_store`.
Cube's default `sub` claim is `cube:deployment:`. To match
the `:component:` examples in this guide (or to add
`:region:`), open your Azure token config in **Admin → OIDC** and
paste one of these templates into the **Subject Claim Format** field:
* `cube:deployment:{deployment_id}:component:{component}` — for the
`cube:deployment::component:` examples below.
* `cube:deployment:{deployment_id}:component:{component}:region:{region}` —
to additionally pin a [Cube Cloud region][ref-cube-cloud-region].
See [the subject editor section][ref-sub-editor] for the full syntax.
Azure pins the federated credential's `subject` field literally — changing
the format means recreating every federated credential that references it.
Create the new federated credential first, then change the **Subject Claim
Format** on the token config.
## Step 3: Set the deployment identity
Add two env vars to your deployment under **Settings → Environment
variables**:
```dotenv theme={"dark"}
AZURE_TENANT_ID=00000000-0000-0000-0000-000000000000
AZURE_CLIENT_ID=11111111-1111-1111-1111-111111111111
```
* **`AZURE_TENANT_ID`** — the Microsoft Entra ID (Azure AD) tenant where
your app registration lives.
* **`AZURE_CLIENT_ID`** — the Application (client) ID of the app
registration.
## Step 4: Assign roles on Azure resources
Grant the app registration the standard Azure RBAC roles it needs, the
same way you'd grant any service principal access to a resource. Examples
follow per target.
## Azure SQL
See [Step 2](#step-2-add-a-federated-credential) — pin subject to
`cube:deployment::component:cube_api`.
Connect to your Azure SQL database as an Entra-authenticated admin and
create a user backed by the app registration:
```sql theme={"dark"}
CREATE USER [Cube Cloud Deployment] FROM EXTERNAL PROVIDER;
ALTER ROLE db_datareader ADD MEMBER [Cube Cloud Deployment];
GRANT EXECUTE ON SCHEMA::dbo TO [Cube Cloud Deployment];
```
The user name must match the app registration's display name.
Set the MSSQL driver and identity env vars on the deployment:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=mssql
CUBEJS_DB_HOST=my-server.database.windows.net
CUBEJS_DB_NAME=my-db
CUBEJS_DB_PORT=1433
AZURE_TENANT_ID=00000000-0000-0000-0000-000000000000
AZURE_CLIENT_ID=11111111-1111-1111-1111-111111111111
```
The MSSQL driver uses the `azure-active-directory-access-token` auth
type with the federated token Cube provides. No username, password, or
client secret needed.
## Azure Blob Storage export bucket
If your data source uses an [export bucket][ref-export-bucket] for
pre-aggregation unloads, grant the app registration **Storage Blob Data
Contributor** on the storage account.
Assign **Storage Blob Data Contributor** on the storage account scope:
```bash theme={"dark"}
az role assignment create \
--assignee \
--role "Storage Blob Data Contributor" \
--scope "/subscriptions//resourceGroups//providers/Microsoft.Storage/storageAccounts/"
```
For tighter scoping, narrow to a specific container with
`.../blobServices/default/containers/` instead of the whole
storage account.
Point the export bucket env vars at your container:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=azure
CUBEJS_DB_EXPORT_BUCKET=https://.blob.core.windows.net/
AZURE_TENANT_ID=00000000-0000-0000-0000-000000000000
AZURE_CLIENT_ID=11111111-1111-1111-1111-111111111111
```
The Azure storage client picks up the same federated identity. See the
[export bucket reference][ref-export-bucket] for the full set of
variables.
OIDC only covers Cube's **read** side of the export bucket. The data
warehouse itself (Snowflake on Azure, Synapse, …) runs the `UNLOAD` that
writes objects to Blob Storage, and the warehouse cannot federate with
Cube's OIDC issuer. You still need to provide **separate credentials for
the unload** so the warehouse can write to the container — typically a
storage account key, SAS token, or a warehouse-side storage integration —
via the standard export bucket env vars (e.g.
`CUBEJS_DB_EXPORT_BUCKET_AZURE_KEY`, or the driver-specific
storage-integration variables). OIDC then handles Cube's download of the
unloaded objects from the bucket.
## Azure Blob Storage CSPS bucket
Cube Store CSPS lets you store pre-aggregations in your own Azure Blob
Storage container. Cube Store gets a separate OIDC token whose `sub` claim
ends in `component:cube_store`, so the federated credential can be locked
down to that component — even if the same app registration were shared with
the rest of the deployment, only Cube Store would match.
Every Cube Store worker emits a `sub` of the form
`cube:deployment::component:cube_store`. Unlike AWS IAM —
which matches `sub` with a `StringLike` wildcard, so one role can cover the
whole tenant — **Azure pins the federated credential's `subject` literally**
and has no tenant-wide `*`. So CSPS on Azure is **per deployment**: each
deployment that writes pre-aggregations to the container needs its own
federated credential carrying that deployment's exact `sub`. If you expect
many deployments, mind the 20-credential limit and see [Scaling past 20
federated credentials](#scaling-past-20-federated-credentials).
Add a federated credential on the app registration, pinning the subject
to the deployment's Cube Store component:
```bash theme={"dark"}
az ad app federated-credential create \
--id \
--parameters '{
"name": "cube--deployment--cube-store",
"issuer": "https://.cubecloud.dev",
"subject": "cube:deployment::component:cube_store",
"audiences": ["api://AzureADTokenExchange"]
}'
```
Make sure the Azure token config's **Subject Claim Format** includes the
`:component:` segment
(`cube:deployment:{deployment_id}:component:{component}`) — see
[Step 2](#step-2-add-a-federated-credential) — otherwise the Cube Store
`sub` won't match this credential.
Grant the app registration **Storage Blob Data Contributor** so Cube
Store can read, write, and list pre-aggregation blobs. Scope it to the
container for tightest isolation:
```bash theme={"dark"}
az role assignment create \
--assignee \
--role "Storage Blob Data Contributor" \
--scope "/subscriptions//resourceGroups//providers/Microsoft.Storage/storageAccounts//blobServices/default/containers/"
```
Use the storage account scope (drop the
`/blobServices/default/containers/` suffix) if you'd rather
one assignment cover multiple containers.
For each deployment that should use this container, go to **Settings →
Pre-Aggregation Storage** on the deployment and:
* Toggle **Enable CSPS** on.
* **Storage Provider**: Azure Blob Storage.
* **Storage Account**: ``.
* **Container**: ``.
* **Client ID**: the app registration's Application (client) ID.
* **Tenant ID**: the Microsoft Entra ID (Azure AD) directory ID.
Click **Test Connection** to verify Cube Store can exchange its OIDC
token with Azure AD and access the container, then **Apply**. Cube Store
starts writing pre-aggregations to your container on the next refresh.
## Scaling past 20 federated credentials
A single app registration accepts at most **20 federated credentials**.
Three patterns cover most growth scenarios:
* **One app registration per deployment.** Each deployment gets its own
app + role assignments. Clean isolation, but you have to provision a new
app every time you add a deployment.
* **Multiple app registrations behind one set of role assignments.** Group
deployments by access pattern (e.g. read-only vs read-write); each group
gets its own app, and the role assignments target the same resources.
* **User-assigned managed identities.** Each managed identity has its own
20-credential limit, and you can attach many of them to your tenant.
Useful when you want the resource permissions managed in Azure
alongside other infrastructure rather than as RBAC on app registrations.
If you expect more than \~10 deployments in a single tenant, plan for one
of these patterns up front — splitting later is a recreation, not a
migration.
## Verifying the setup
The fastest way to confirm the federated credential is wired up correctly
is the **Test connection** button on the relevant settings page (data
source wizard, BYO LLM provider). Behind the scenes, Cube issues a real
OIDC token, exchanges it with Azure AD, and returns a precise error if
anything is misconfigured.
If the test fails:
| Symptom | Likely cause |
| ----------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `AADSTS70021: No matching federated identity record found` | The federated credential's `issuer`, `subject`, and `audiences` triple doesn't match the JWT exactly. Compare with the `iss` / `sub` / `aud` of a token from the **Test connection** error response. |
| `AADSTS700213: No matching federated identity record found ... subject` | The subject claim differs by even a single character. Azure doesn't support wildcards — you need one credential per exact subject. |
| `AADSTS50034: The user account does not exist` | You're using the wrong `AZURE_TENANT_ID`. Double-check the directory ID on the app registration's Overview blade. |
| `AADSTS500011: The resource principal ... was not found` | The role assignment hasn't been created on the target resource, or it's still propagating. Wait a minute and retry; if it persists, re-check the `--assignee` and `--scope` values. |
Federation events show up in **Microsoft Entra ID → Sign-in logs** under
the **Service principal sign-ins** tab — filter by the app registration to
see which deployments are authenticating, when, and which resources they
hit.
[ref-oidc-overview]: /admin/deployment/oidc
[ref-sub-editor]: /admin/deployment/oidc#subject-claim-format
[ref-cube-cloud-region]: /admin/deployment/infrastructure#what-is-a-cube-cloud-region
[ref-export-bucket]: /admin/connect-to-data#export-bucket
# GCP
Source: https://docs.cube.dev/admin/deployment/oidc/gcp
Configure GCP Workload Identity Federation to trust Cube's OIDC issuer and use it for BigQuery and GCS export buckets.
This guide walks through configuring GCP to trust Cube's OIDC issuer using
Workload Identity Federation (WIF) and shows the setup for the most common
targets — BigQuery and a GCS export bucket.
If you haven't enabled OIDC for your tenant yet, start with the
[OIDC overview][ref-oidc-overview].
Available on the [Enterprise plan](https://cube.dev/pricing).
## Prerequisites
* The Cube tenant has OIDC enabled and a `GCP` token config exists under
**Admin → OIDC**.
* IAM access to your GCP project sufficient to create Workload Identity
Pools, providers, and service accounts.
* Your tenant slug — the leftmost label of your tenant's console URL.
Throughout this guide it's referenced as `` (and the full
issuer URL as `https://.cubecloud.dev`). Substitute your
actual slug everywhere it appears.
Commands and config snippets in this guide use angle-bracket placeholders —
``, ``, ``, etc. **Replace each
placeholder with your real value** before running. GCP will accept these
strings literally and the federation call will fail with a confusing error.
## How GCP federation works
Cube doesn't talk to GCP STS directly. Instead, it writes a small
**credential configuration JSON** to disk that points the GCP SDK at the
OIDC token file. The Google client library handles the two-step exchange
internally:
```mermaid theme={"dark"}
sequenceDiagram
participant Cube as Cube Deployment
participant STS as Google STS
participant IAM as IAM Credentials API
participant BQ as BigQuery / GCS
Cube->>STS: Exchange OIDC JWT for federated access token (audience = WIF pool provider URI)
STS->>Cube: Federated access token
Cube->>IAM: generateAccessToken on target service account (impersonation)
IAM->>Cube: Service account access token (1h)
Cube->>BQ: Authenticated request
```
You can also skip the impersonation step and grant permissions directly to
the federated principal — see [Direct federation](#direct-federation) at the
end of this page.
## Step 1: Create a Workload Identity Pool and provider
Run these once per GCP project. The pool is a container; the provider is
what actually trusts your Cube issuer.
```bash theme={"dark"}
gcloud iam workload-identity-pools create cube-pool \
--location=global \
--display-name="Cube workload identity"
gcloud iam workload-identity-pools providers create-oidc cube \
--location=global \
--workload-identity-pool=cube-pool \
--issuer-uri="https://.cubecloud.dev" \
--attribute-mapping="google.subject=assertion.sub,attribute.issuer=assertion.iss" \
--attribute-condition="assertion.sub.startsWith('cube:deployment:')"
```
What each option does:
| Option | Purpose |
| ----------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `--issuer-uri` | Cube tenant URL. GCP fetches `${issuer-uri}/.well-known/openid-configuration` and JWKS from here. |
| `--attribute-mapping` | Maps the JWT `sub` to GCP's `google.subject`. The mapped subject is what IAM bindings reference when granting impersonation rights. |
| `--attribute-condition` | An optional CEL expression that GCP evaluates on every token exchange. The condition above accepts only deployment-scoped tokens. |
Note the provider's full resource URI — you'll need it shortly:
```
//iam.googleapis.com/projects/PROJECT_NUMBER/locations/global/workloadIdentityPools/cube-pool/providers/cube
```
`PROJECT_NUMBER` is the numeric project number, not the project ID. You can
fetch it with `gcloud projects describe PROJECT_ID --format='value(projectNumber)'`.
## Step 2: Set the deployment identity
Add two env vars to your deployment under **Settings → Environment variables**:
```dotenv theme={"dark"}
GCP_POOL_AUDIENCE=//iam.googleapis.com/projects//locations/global/workloadIdentityPools/cube-pool/providers/cube
GCP_SERVICE_ACCOUNT_EMAIL=cube-deployment@my-project.iam.gserviceaccount.com
```
* **`GCP_POOL_AUDIENCE`** — the full resource URI of the WIF provider you
created in Step 1. This becomes the `aud` claim on the GCP token Cube
mints.
The GCP token config under **Admin → OIDC** must carry this **exact same**
provider resource path as its audience (unlike AWS/Azure, GCP has no global
audience, so it is not auto-filled). If the token config's audience and
`GCP_POOL_AUDIENCE` disagree, the STS exchange fails with
`invalid_grant: The audience in ID Token ... does not match the expected
audience`.
* **`GCP_SERVICE_ACCOUNT_EMAIL`** — the service account that Cube
impersonates after federation succeeds. Cube assumes this service account
by default for every GCP SDK call inside the deployment.
If you want to skip impersonation entirely and have Cube call GCP services
as the federated principal directly, leave `GCP_SERVICE_ACCOUNT_EMAIL` unset.
See [Direct federation](#direct-federation) below.
## Step 3: Build the IAM bindings
There are two distinct IAM bindings to set:
1. **Workload Identity User** on the impersonated service account — lets
the federated principal call `generateAccessToken` on it.
2. **Resource access** (BigQuery, GCS, etc.) on the impersonated service
account itself — what the SA is actually allowed to do once Cube is
running as it.
The Workload Identity User binding's `--member` controls which Cube
deployments / components can impersonate the SA. Patterns mirror the AWS
`sub` patterns:
| Trust scope | `--member` |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| One specific deployment, any component | `principalSet://iam.googleapis.com/projects/123/locations/global/workloadIdentityPools/cube-pool/subject/cube:deployment::component:cube_api` |
| Every component of every deployment in the tenant | `principalSet://iam.googleapis.com/projects/123/locations/global/workloadIdentityPools/cube-pool/*` |
| Only Cube Store, across every deployment | Use a CEL `--attribute-condition` on the provider that constrains `sub` to end in `:component:cube_store`, then bind the whole pool. |
The `principal://` (singular) form pins to one exact subject; `principalSet://`
matches a set.
Cube's default `sub` claim is `cube:deployment:`. To match
the `:component:` patterns in the table above (or to add
`:region:`), open your GCP token config in **Admin → OIDC** and
paste one of these templates into the **Subject Claim Format** field:
* `cube:deployment:{deployment_id}:component:{component}` — for the
`principalSet://...subject/cube:deployment::component:`
patterns above.
* `cube:deployment:{deployment_id}:component:{component}:region:{region}` —
to additionally pin a [Cube Cloud region][ref-cube-cloud-region].
See [the subject editor section][ref-sub-editor] for the full syntax.
Update the GCP IAM binding (or the WIF provider's CEL `--attribute-condition`)
first, then change the **Subject Claim Format** on the token config —
otherwise existing tokens won't match the binding and impersonation will
fail.
## BigQuery
Provision a regular GCP service account that will hold the BigQuery
permissions Cube assumes:
```bash theme={"dark"}
gcloud iam service-accounts create cube-deployment \
--display-name="Cube deployment"
```
Note the email — `cube-deployment@PROJECT_ID.iam.gserviceaccount.com` —
this is what you'll set as `GCP_SERVICE_ACCOUNT_EMAIL`.
Bind the federated principal to the service account so the deployment's
OIDC token can impersonate it:
```bash theme={"dark"}
PROJECT_NUMBER=""
DEPLOYMENT_ID=""
gcloud iam service-accounts add-iam-policy-binding \
cube-deployment@my-project.iam.gserviceaccount.com \
--role=roles/iam.workloadIdentityUser \
--member="principal://iam.googleapis.com/projects/${PROJECT_NUMBER}/locations/global/workloadIdentityPools/cube-pool/subject/cube:deployment:${DEPLOYMENT_ID}:component:cube_api"
```
Standard IAM, as if the service account were any normal workload:
```bash theme={"dark"}
gcloud projects add-iam-policy-binding my-project \
--member="serviceAccount:cube-deployment@my-project.iam.gserviceaccount.com" \
--role=roles/bigquery.dataViewer
gcloud projects add-iam-policy-binding my-project \
--member="serviceAccount:cube-deployment@my-project.iam.gserviceaccount.com" \
--role=roles/bigquery.jobUser
```
Tighten these to specific datasets / tables in production. The minimum
set of roles Cube needs to run BigQuery queries is
`roles/bigquery.dataViewer` plus `roles/bigquery.jobUser` on the project
the queries run in.
Set the BigQuery driver and identity env vars on the deployment:
```dotenv theme={"dark"}
CUBEJS_DB_TYPE=bigquery
CUBEJS_DB_BQ_PROJECT_ID=my-project
GCP_POOL_AUDIENCE=//iam.googleapis.com/projects//locations/global/workloadIdentityPools/cube-pool/providers/cube
GCP_SERVICE_ACCOUNT_EMAIL=cube-deployment@my-project.iam.gserviceaccount.com
```
The BigQuery driver follows the GCP default credential chain, picks up
the credential config Cube generates, and runs as
`cube-deployment@my-project.iam.gserviceaccount.com`. No service account
JSON key is ever used.
## GCS export bucket
If your data source uses an [export bucket][ref-export-bucket] for
pre-aggregation unloads (BigQuery, Snowflake on GCP, etc.), grant the
deployment's service account read / write access to the bucket.
Add a bucket-scoped IAM binding for the deployment's service account:
```bash theme={"dark"}
gcloud storage buckets add-iam-policy-binding gs://my-export-bucket \
--member="serviceAccount:cube-deployment@my-project.iam.gserviceaccount.com" \
--role=roles/storage.objectAdmin
```
`objectAdmin` covers reads, writes, and deletes within the bucket. If
you only need writes (e.g. you have a separate process cleaning up old
exports), `roles/storage.objectCreator` is enough.
Point the export bucket env vars at your bucket:
```dotenv theme={"dark"}
CUBEJS_DB_EXPORT_BUCKET_TYPE=gcs
CUBEJS_DB_EXPORT_BUCKET=my-export-bucket
```
The GCS client inside Cube picks up the same default identity. See the
[export bucket reference][ref-export-bucket] for the full set of
variables.
OIDC only covers Cube's **read** side of the export bucket. The data
warehouse itself (BigQuery, Snowflake on GCP, …) runs the `UNLOAD` /
`EXPORT DATA` that writes objects to the bucket, and the warehouse cannot
federate with Cube's OIDC issuer. You still need to provide **separate
credentials for the unload** so the warehouse can write to GCS — typically
an HMAC key pair or a warehouse-side service-account integration — via the
standard export bucket env vars (e.g.
`CUBEJS_DB_EXPORT_GCS_CREDENTIALS`, or the driver-specific
storage-integration variables). OIDC then handles Cube's download of the
unloaded objects from the bucket.
## Direct federation
If you'd rather skip the service account impersonation hop, grant
permissions directly to the federated principal and leave
`GCP_SERVICE_ACCOUNT_EMAIL` unset on the deployment. Cube generates a
credential config that performs only the OIDC-to-federated-token exchange,
and the resulting token is what your code authenticates with.
```bash theme={"dark"}
PROJECT_NUMBER=""
DEPLOYMENT_ID=""
gcloud projects add-iam-policy-binding my-project \
--member="principal://iam.googleapis.com/projects/${PROJECT_NUMBER}/locations/global/workloadIdentityPools/cube-pool/subject/cube:deployment:${DEPLOYMENT_ID}:component:cube_api" \
--role=roles/bigquery.dataViewer
```
Direct federation is simpler — fewer moving parts, and the principal
identity in audit logs is the Cube subject itself rather than an
intermediate service account. The trade-off is that some GCP services
(notably anything that requires `iam.serviceAccountTokenCreator`) only
accept service-account principals, so you may need the impersonation path
for those.
## Cube Store CSPS bucket
Cube Store CSPS lets you store pre-aggregations in your own GCS bucket. Cube
Store gets a separate OIDC token whose `sub` claim ends in
`component:cube_store`, so the IAM bindings can be locked down to that
component — even if the same service account were ever shared with the rest
of the deployment, only Cube Store would be able to impersonate it.
Every Cube Store worker emits a `sub` of the form
`cube:deployment::component:cube_store`. As with the
[deployment identity bindings](#step-3-build-the-iam-bindings), how broadly
you share access is controlled by the `principalSet` member you bind:
* `…/workloadIdentityPools/cube-pool/*` — paired with a provider
`--attribute-condition` that constrains `sub` to end in
`:component:cube_store`, this gives you **one service account + one bucket
for the whole tenant.** Every deployment writes pre-aggregations to the
same bucket, isolated by Cube Store's own per-deployment path prefix.
Easiest to operate.
* `…/cube-pool/subject/cube:deployment::component:cube_store`
— **per-deployment isolation.** Pin access to a single deployment's Cube
Store so its pre-aggregations live in a dedicated bucket no other
deployment can touch.
Cube Store can authenticate either by impersonating a service account or by
federating to the bucket directly — pick one.
**Service account impersonation** (works with every GCS feature). Give a
service account object access on the bucket, then let the Cube Store
principal impersonate it:
```bash theme={"dark"}
PROJECT_NUMBER=""
# 1. The SA Cube Store impersonates — grant it object access on the bucket.
gcloud storage buckets add-iam-policy-binding gs://my-csps-bucket \
--member="serviceAccount:cube-cubestore@my-project.iam.gserviceaccount.com" \
--role=roles/storage.objectAdmin
# 2. Let Cube Store's federated principal impersonate that SA. Swap the
# trailing `*` for a single `subject/...:component:cube_store` to pin
# one deployment.
gcloud iam service-accounts add-iam-policy-binding \
cube-cubestore@my-project.iam.gserviceaccount.com \
--role=roles/iam.workloadIdentityUser \
--member="principalSet://iam.googleapis.com/projects/${PROJECT_NUMBER}/locations/global/workloadIdentityPools/cube-pool/*"
```
**Direct federation** (no impersonation hop). Grant the Cube Store
principal object access on the bucket directly, and leave the service
account blank in the UI:
```bash theme={"dark"}
PROJECT_NUMBER=""
DEPLOYMENT_ID=""
gcloud storage buckets add-iam-policy-binding gs://my-csps-bucket \
--member="principalSet://iam.googleapis.com/projects/${PROJECT_NUMBER}/locations/global/workloadIdentityPools/cube-pool/subject/cube:deployment:${DEPLOYMENT_ID}:component:cube_store" \
--role=roles/storage.objectAdmin
```
`roles/storage.objectAdmin` covers the reads, writes, lists, and deletes
Cube Store performs as it builds and evicts pre-aggregation partitions.
Cube Store only emits the `:component:cube_store` subject if your GCP
token config uses a subject claim format that includes the `:component:`
segment. In **Admin → OIDC**, set the GCP token config's **Subject Claim
Format** to `cube:deployment:{deployment_id}:component:{component}` (see
[Step 3](#step-3-build-the-iam-bindings)). Update the IAM binding or the
provider's CEL `--attribute-condition` before changing the format.
For each deployment that should use this bucket, go to **Settings →
Pre-Aggregation Storage** on the deployment and:
* Toggle **Enable CSPS** on.
* **Storage Provider**: Google Cloud Storage.
* **GCS Bucket**: `my-csps-bucket`.
* **Service Account Email** (optional): `cube-cubestore@my-project.iam.gserviceaccount.com` — or leave blank if you granted the Cube Store principal bucket access directly (direct federation).
* **Workload Identity Provider**: the provider resource name (without the
`//iam.googleapis.com/` prefix),
`projects//locations/global/workloadIdentityPools/cube-pool/providers/cube`.
Click **Test Connection** to verify Cube Store can federate — and
impersonate the service account, if one is set — and read/write the
bucket, then **Apply**. Cube Store starts writing pre-aggregations to
your bucket on the next refresh.
## Verifying the setup
The fastest way to confirm WIF is wired up correctly is the **Test
connection** button on the relevant settings page (data source wizard,
CSPS settings). Behind the scenes, Cube issues a real OIDC token, performs
the GCP STS exchange, optionally impersonates the service account, and
returns a precise error if anything is misconfigured.
If the test fails:
| Symptom | Likely cause |
| ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `Permission iam.serviceAccounts.getAccessToken denied` | The Workload Identity User binding on the service account is missing or its `--member` doesn't match the deployment's `sub`. Double-check the principal URI. |
| `INVALID_ARGUMENT: Invalid value for audience` | `GCP_POOL_AUDIENCE` doesn't match the WIF provider's URI. Re-run `gcloud iam workload-identity-pools providers describe` and copy the value verbatim. |
| `The token issuer ... does not match the configured issuer` | The provider was created with a different `--issuer-uri` than your tenant URL, or your tenant slug has changed. Re-create the provider with the correct URL. |
| `attribute condition ... evaluated to false` | The CEL `--attribute-condition` on the provider rejected the token. Inspect the `sub` Cube emits and adjust the condition. |
Federation events show up in **Cloud Audit Logs** under
`sts.googleapis.com` (the token exchange) and `iamcredentials.googleapis.com`
(the SA impersonation). The Cube subject is included in both, so you can
trace which deployment authenticated against which service account at any
point in time.
[ref-oidc-overview]: /admin/deployment/oidc
[ref-sub-editor]: /admin/deployment/oidc#subject-claim-format
[ref-cube-cloud-region]: /admin/deployment/infrastructure#what-is-a-cube-cloud-region
[ref-export-bucket]: /admin/connect-to-data#export-bucket
# OIDC workload identity
Source: https://docs.cube.dev/admin/deployment/oidc/index
Cube's keyless authentication mechanism. Federate your cloud accounts and external services to a Cube-issued OIDC token instead of provisioning long-lived secrets.
OIDC workload identity is Cube's keyless authentication mechanism. Cube acts
as an OpenID Connect (OIDC) identity provider for every deployment in your
account: it mints short-lived, signed JWTs that you federate with the systems
your deployment talks to. The system on the other side — AWS, GCP, Azure,
Snowflake, your own Bedrock-hosted LLM — verifies the token directly against
Cube's published JWKS and grants access, eliminating the need for long-lived
access keys, manual rotation, or shared secrets passed through Cube.
Available on the [Enterprise plan](https://cube.dev/pricing).
You can use OIDC workload identity to authenticate to:
* **Data sources** — AWS Athena, Redshift, BigQuery, Snowflake, and any other
driver that supports federated credentials.
* **Export buckets** — S3 and GCS buckets used for `EXPORT_BUCKET` pre-aggregation
unloads. OIDC covers Cube's download of the unloaded objects; the warehouse's
`UNLOAD` write still needs its own credentials configured on the client side
— see the per-cloud guides for details.
* **Cube Store CSPS** — a per-deployment S3, GCS, or Azure Blob Storage
bucket that holds your Cube Store pre-aggregations (Customer-Supplied
Pre-aggregation Storage).
* **Bring-your-own LLM providers** — AWS Bedrock, Google Vertex AI, and Azure
OpenAI behind your own cloud account.
* **Other cloud services** — any cloud service that supports OIDC federation
that you'd like to access from the dynamic part of your data model
(`cube.py` / `cube.js`, dynamic schema generators, `checkAuth`, custom
pre-aggregation strategies, etc.) — works with the same token.
The same mechanism powers all of these — the only thing that changes between
integrations is the trust policy on the receiving side and which IAM role,
service account, or app registration Cube assumes.
## How it works
If you're new to OpenID Connect, the [OpenID Connect Core 1.0
specification][link-oidc-spec] is the canonical reference; the
[OAuth.com OIDC primer][link-oidc-primer] is a friendlier read. Cube's
implementation is a standard OIDC provider — anything that works with
GitHub Actions OIDC, Vercel OIDC, or AWS IAM OIDC providers works the
same way here.
Each Cube tenant hosts a per-tenant OIDC issuer at its own domain, e.g.
`https://.cubecloud.dev`. The issuer publishes the standard OIDC discovery
endpoints, and Cube signs every token with a KMS-backed RS256 key.
```mermaid theme={"dark"}
sequenceDiagram
participant Cube as Cube Deployment
participant Issuer as Cube OIDC Issuer (your tenant URL)
participant STS as Cloud STS (AWS / GCP / Azure)
participant Resource as Customer Resource (Athena / S3 / BigQuery / ...)
Cube->>Issuer: Request signed JWT for audience X
Issuer->>Cube: JWT (sub, aud, iss, exp = +1h)
Cube->>STS: Exchange JWT for temporary credentials
STS->>Issuer: Fetch JWKS, verify signature, evaluate trust policy
STS->>Cube: Temporary credentials (1h)
Cube->>Resource: Access with temporary credentials
```
The deployment never sees an access key or service account JSON. The cloud
SDKs inside Cube — AWS SDK, GCP `google-auth-library`, Azure SDK — pick up
the token automatically from a path supplied by Cube and exchange it with
the cloud's own STS for short-lived credentials.
## Endpoints served by Cube
Each tenant's issuer URL is `https://.cubecloud.dev`. Cube serves
the standard OIDC endpoints from that domain:
| Endpoint | Purpose |
| ----------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| `/.well-known/openid-configuration` | OIDC discovery document. Cloud providers fetch this to learn the issuer, supported algorithms, and JWKS location. |
| `/.well-known/jwks.json` | JSON Web Key Set. Contains the public keys cloud providers use to verify token signatures. |
Both endpoints are unauthenticated, because they're consumed by AWS STS,
Google STS, Azure AD, and any other relying party that validates
Cube-issued tokens.
## Token configs
A **token config** is a tenant-level entry that tells Cube which audience to
mint tokens for and how to format the `sub` claim each token carries. Each
token config produces one audience-scoped JWT that is delivered to all
deployments in the tenant.
| Field | What it controls |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Audience type** | One of `AWS`, `GCP`, `Azure`, or `Custom`. `AWS` and `Azure` pre-fill the audience value. |
| **Audience** | The JWT `aud` claim value. Pre-set for `AWS` (`sts.amazonaws.com`) and `Azure` (`api://AzureADTokenExchange`). For `GCP` you supply the **Workload Identity Federation provider resource path** (`//iam.googleapis.com/projects/.../providers/...`) — there is no global GCP audience. User-supplied for `Custom`. |
| **Name** | A short, filesystem-safe slug that becomes part of the token file name (`cube_cloud_token_`). |
| **Subject Claim Format** | Template for the JWT `sub` claim. Defaults to `cube:deployment:{deployment_id}` if left empty. |
| **Custom Claims** | Optional extra JWT claims emitted verbatim in every minted token — see [Custom claims](#custom-claims) below. |
| **Target Env Var** | Custom audience type only: env var Cube populates with the token file path in every execution context — see [Custom token configs](#custom-token-configs) below. |
You can have one config per well-known audience type (AWS, GCP, Azure) plus
any number of custom configs for tools like [Snowflake][ref-oidc-snowflake]
or Databricks. A deployment that integrates with several clouds
simultaneously gets a separate token file per audience — the Cube runtime
picks the right one based on which SDK is making the call.
### Subject claim format
The `sub` claim is what the receiving cloud provider's trust policy matches
on to decide which deployment or component is allowed to assume a role.
Cube's default template is `cube:deployment:{deployment_id}` — sufficient if
you only need to identify the deployment.
To scope further — by component (Cube API vs. Cube Store vs. the AI Engineer)
or by Cube Cloud region — override the template with one of the **Common
templates** the dialog offers:
* `cube:deployment:{deployment_id}` (default)
* `cube:deployment:{deployment_id}:component:{component}`
* `cube:deployment:{deployment_id}:component:{component}:region:{region}`
Three placeholders are supported anywhere in the template:
| Placeholder | What it resolves to |
| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `{deployment_id}` | The deployment's numeric ID. Substitutes to the literal `global` for tenant-scoped tokens (e.g., AI Engineer running in the control plane). |
| `{component}` | The component requesting the token: `cube_api`, `cube_store`, or `ai_engineer`. Substitutes to `global` when no component context is set. |
| `{region}` | The deployment's [Cube Cloud region][ref-cube-cloud-region]. For tenant-scoped tokens minted in the control plane, substitutes to the literal `control-plane`. |
The component values Cube emits are `cube_api` (the Cube API, refresh
worker, dev-mode pods, Cold Start Workers, etc.), `cube_store` (Cube Store),
and `ai_engineer` (the AI Engineer service, used for BYO LLM auth).
Outside of placeholders, the template may contain only the characters
`[A-Za-z0-9:_-]`. Anything else is rejected at save time.
Changing the subject claim format on a token config that's already in use
**will break customer trust policies that pin the old `sub` shape**. Update
the trust policy in your cloud account first, then change the format here.
Subject claim format changes can take **up to one hour to fully propagate**.
Already-minted tokens are cached at several layers — the credentials broker
sidecar and the cloud SDKs' own STS-credential caches — and remain valid
until their TTL expires (up to 1h). Plan trust-policy rollouts so the new
`sub` shape is accepted before the old one stops being honored, and expect
a window where both formats are in flight.
### Custom claims
Some token consumers authorize through claims beyond `iss`/`sub`/`aud`. The
canonical example is [Snowflake External OAuth][ref-oidc-snowflake], which
grants session roles exclusively through the `scp` claim — a token without
it authenticates but fails role authorization. The **Custom Claims** field
on a token config adds such claims to every token minted for that audience.
Each claim is a name plus a value. Values are emitted as strings; to emit
an array claim, comma-separate the values (a trailing comma forces a
one-element array). For example, claim `scp` with value `session:role-any`
covers the Snowflake case.
Custom claims are deliberately constrained:
* The registered JWT claims (`iss`, `sub`, `aud`, `exp`, `iat`, `nbf`,
`jti`) and anything under the `cube:` namespace cannot be overridden —
those carry the token's identity semantics and stay authoritative.
* Values are strings or arrays of strings only, with bounded lengths.
Like subject-format changes, custom-claim changes take up to one hour to
fully propagate as cached tokens expire.
### Custom token configs
For `Custom` audiences, the driver or SDK on the other side reads the token
from a file. Where that file lives depends on **where the deployment is
running** — deployed pods, a dev-mode worker, and a test-connection probe
each keep it in a different place — so the path must never be hard-coded.
The **Target Env Var** field solves this: name the env var the consumer
reads (e.g. `CUBEJS_DB_SNOWFLAKE_OAUTH_TOKEN_PATH` for the Snowflake
driver), and Cube sets it to the correct token-file path in every context.
Leave `CUBEJS_DB_USER` / `CUBEJS_DB_PASS` (or the equivalent secret) unset
and don't write the path by hand. Well-known audiences (AWS/GCP/Azure)
don't need this — their SDKs read fixed, standard env vars that Cube
already populates.
The example below — a custom config for the Snowflake driver — combines
both fields: a `scp` custom claim (which Snowflake requires for role
authorization) and a Target Env Var of `CUBEJS_DB_SNOWFLAKE_OAUTH_TOKEN_PATH`.
The [Snowflake data source page][ref-oidc-snowflake] walks through the full setup.
## Trust mechanism
The receiving system trusts Cube tokens by referencing two things:
1. **The issuer URL** — `https://.cubecloud.dev`. The cloud
provider fetches your tenant's JWKS from this URL and uses it to validate
every token signature. As long as Cube holds the matching private key,
tokens issued for your tenant verify; tokens issued by any other Cube
tenant do not.
2. **The subject claim** — `sub`. The trust policy on the receiving side
matches the `sub` exactly (or with wildcards) against the format you
configured above to decide which deployments / components are allowed
in.
## Token lifecycle
| Property | Value |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Algorithm** | RS256 |
| **TTL** | 1 hour |
| **Refresh** | Automatic — Cube re-mints each token before it expires and atomically replaces the token file. The cloud SDK on the other side picks up the new token without reconnecting. |
| **Signing keys** | Per-tenant, KMS-backed. Rotated regularly; the JWKS endpoint always serves the current and previous public keys so in-flight tokens remain valid across rotations. |
| **Revocation** | TTL-based. Emergency revocation is performed by rotating the signing key, which invalidates all tokens within the next TTL window. |
## Enabling OIDC for your tenant
Workload identity is configured at the **tenant** level under
**Admin → OIDC**. From there, you turn the issuer on and create one token
config per integration target.
In **Admin → OIDC**, flip the **OIDC token issuance is active for all
deployments** switch to on. This activates the discovery and JWKS
endpoints on your tenant URL and unlocks the deployment-level OIDC
settings everywhere else in the UI. Click **OpenID Configuration** or
**JWKS Endpoint** to grab the URLs you'll hand to your cloud provider
when configuring federation.
Click **Add Config** and pick the audience type. For `AWS` and `Azure`,
the audience auto-fills with the standard value. For `GCP`, enter the
**Workload Identity Federation provider resource path** (the same
`//iam.googleapis.com/projects/.../providers/...` value you use for
`GCP_POOL_AUDIENCE`) — GCP has no global audience. For `Custom`, enter the
audience your target system expects. Optionally set the **Subject Claim
Format** — see [Token configs](#token-configs) for the available templates.
Follow the cloud-specific guides below to register Cube as an OIDC
provider on the system you want Cube to access, create a role / service
account / app registration, and write a trust policy that pins the
issuer URL and the subject pattern.
Athena, S3 export buckets, Bedrock, and any other AWS-IAM-protected service.
BigQuery, GCS export buckets, and any other GCP service via Workload Identity Federation.
Azure SQL, Blob Storage export buckets, and any other Azure service via federated credentials.
For each deployment, set the cloud-specific "who to become" env vars
under **Settings → Environment variables**. These point Cube at the
role or service account it should assume on your side:
| Cloud | Env vars |
| ----- | ------------------------------------------------ |
| AWS | `AWS_ROLE_ARN` |
| GCP | `GCP_POOL_AUDIENCE`, `GCP_SERVICE_ACCOUNT_EMAIL` |
| Azure | `AZURE_TENANT_ID`, `AZURE_CLIENT_ID` |
These are the **default identity** for the deployment — see below.
## Default identity for a deployment
When you set `AWS_ROLE_ARN` (or the GCP / Azure equivalents) on a deployment,
that role becomes the default cloud identity for everything that runs inside
that deployment:
* **Drivers** that follow the AWS / GCP / Azure default credential chain
(Athena, Redshift, BigQuery, RDS Postgres / MySQL with IAM auth, Azure SQL,
and more) automatically authenticate as this role / service account.
* **Pre-aggregation export bucket I/O** uses the same identity to write to
S3 / GCS / Blob Storage.
* **Code in your data model** — `dataSource` factories, `driverFactory`,
`extendContext`, `checkAuth`, dynamic schema generators — runs with the same
default identity. If you instantiate an AWS / GCP / Azure SDK client with no
explicit credentials, it picks up the deployment's OIDC token automatically.
If you need a driver to talk to a different role or service account than the
deployment default — for example, the deployment's default role can read S3
but a specific Athena workgroup lives in a different account — set the
**driver-specific assume-role env var** in addition to the default. The
driver then uses the default identity to perform a second
`AssumeRole` / impersonation hop and connect as the downstream principal.
| Driver | Driver-specific assume-role env var |
| ----------------- | ------------------------------------------------------------------------------------------------ |
| AWS Athena | `CUBEJS_AWS_ATHENA_ASSUME_ROLE_ARN` |
| AWS Redshift | `CUBEJS_DB_AWS_ROLE_ARN` |
| Other AWS drivers | Driver-specific — see the driver page in [Connect to data](/admin/connect-to-data/data-sources). |
The chain is always:
```
Cube OIDC token → default deployment identity → (optional) driver-specific assumed role
```
Most setups don't need the second hop — granting permissions directly to the
deployment's default role is simpler and easier to reason about.
## Cube Store CSPS
Cube Store can store its pre-aggregations in a **per-deployment** S3, GCS, or
Azure Blob Storage bucket that you own — Customer-Supplied Pre-aggregation
Storage (CSPS). When you enable CSPS for a deployment, Cube Store
authenticates to your bucket with the same OIDC machinery as the rest of the
deployment. The trust policy (AWS IAM role / Azure federated credential) pins
the `sub` claim to `cube:deployment::component:cube_store`,
giving Cube Store its own identity that's distinct from the Cube API.
CSPS is configured under **Settings → Pre-Aggregation Storage** on the
deployment. See
[AWS](/admin/deployment/oidc/aws#cube-store-csps-bucket),
[GCP](/admin/deployment/oidc/gcp#cube-store-csps-bucket), and
[Azure](/admin/deployment/oidc/azure#azure-blob-storage-csps-bucket) for the
trust policy shape on each cloud.
## Audit trail
Every token Cube issues is recorded in the tenant audit log with the
deployment ID, component, audience, and TTL. On the receiving side, the
cloud's own audit trail shows the federation event:
* **AWS CloudTrail** — `AssumeRoleWithWebIdentity` events with the deployment
subject.
* **GCP Cloud Audit Logs** — Workload Identity Federation token-exchange and
service-account-impersonation events.
* **Azure Activity Log** — Federated-credential sign-in events on the app
registration.
Together these give you an end-to-end record of which deployment accessed
which resource and when, without ever issuing a long-lived credential.
[ref-oidc-aws]: /admin/deployment/oidc/aws
[ref-oidc-gcp]: /admin/deployment/oidc/gcp
[ref-oidc-azure]: /admin/deployment/oidc/azure
[ref-oidc-snowflake]: /admin/connect-to-data/data-sources/snowflake#oidc-workload-identity
[ref-env-vars]: /admin/deployment#environment-variables
[ref-data-sources]: /admin/connect-to-data/data-sources
[ref-sub-editor]: /admin/deployment/oidc#subject-claim-format
[ref-cube-cloud-region]: /admin/deployment/infrastructure#what-is-a-cube-cloud-region
[link-oidc-spec]: https://openid.net/specs/openid-connect-core-1_0.html
[link-oidc-primer]: https://www.oauth.com/oauth2-servers/openid-connect/
# Scalability
Source: https://docs.cube.dev/admin/deployment/scalability
Tune Cube Cloud throughput and pre-aggregation capacity by scaling API instances and Cube Store workers.
Cube Cloud also allows adding additional infrastructure to your deployment to
increase scalability and performance beyond what is available with each
Dedicated deployment.
## Auto-scaling of API instances
With a Dedicated deployment, 2 Cube API Instances are included. That said, it
is very common to use more, and [additional API instances][ref-limits] can be
added to your deployment to increase the throughput of your queries. A rough
estimate is that 1 Cube API Instance is needed for every 5-10
requests-per-second served. Cube API Instances can also auto-scale as needed.
To change how many Cube API instances are available in the Dedicated deployment,
go to the deployment’s **Settings** screen, and open
the **Configuration** tab. From this screen, you can set the minimum and
maximum number of Cube API instances for a deployment:
## Sizing Cube Store workers
Cube Store Workers are used to build and persist pre-aggregations. Each Worker
has a **maximum of 150GB** of storage; [additional Cube Store
workers][ref-limits] can be added to your deployment to both increase storage
space and improve pre-aggregation performance. A **minimum of 2** Cube Store
Workers is required for pre-aggregations; this can be adjusted. For a rough
estimate, it will take approximately 2 Cube Store Workers per 4 GB of
pre-aggregated data per day.
Idle workers will automatically hibernate after 10 minutes of inactivity, and
will not consume CCUs until they are resumed. Workers are resumed automatically
when Cube receives a query that should be accelerated by a pre-aggregation, or
when a scheduled refresh is triggered.
To change the number of Cube Store Workers in a deployment, go to the
deployment’s **Settings** screen, and open the **Configuration**
tab. From this screen, you can set the number of Cube Store Workers from the
dropdown:
[ref-limits]: /admin/deployment/limits#resources
# Deployment warm-up
Source: https://docs.cube.dev/admin/deployment/warm-up
Covers optional pre-launch warm-up for data model compilation and pre-aggregations on Dedicated deployments.
Deployment warm-up improves querying performance by executing time-consuming
tasks after a Cube Cloud deployment is spun up but before it's exposed to
requests from users.
Available on [Starter and above plans](https://cube.dev/pricing).
Deployment warm-up is only available for [Dedicated][ref-prod-cluster]
and [Multi-cluster][ref-prod-multi-cluster] deployments.
There are two warm-up options:
* [Data model warm-up](#data-model-warm-up) — compiles the data model and
populates data model compilation cache ahead of time.
* [Pre-aggregation warm-up](#pre-aggregation-warm-up) — builds and refreshes
pre-aggregations ahead of time.
## Data model warm-up
By default, an API instance compiles the [data model][ref-data-model] and
stores results in the data model compilation cache when the first request hits
that API instance. For [multi-tenant][ref-multitenancy] configurations, a
request from a particular tenant would only trigger the data model compilation
for that tenant.
Depending on the complexity of the data model (i.e., the number of cubes,
views, and their members) and the use of [dynamic data models][ref-dynamic-data-model],
its compilation might take from a few milliseconds to a few seconds.
Data-model warm-up will compile the data model before a Cube Cloud deployment
is exposed to requests from users, making sure that they aren't affected by
the data model compilation time.
If a data model warm-up takes more than 15 minutes, the deployment will be
considered unhealthy and rolled back.
### Configuring data model warm-up
To configure data model warm-up, navigate to **Settings → Configuration**
and enable **Warm-up data model before deploying API**:
## Pre-aggregation warm-up
By default, the refresh worker takes care of [pre-aggregation][ref-pre-aggs]
refresh and runs pre-aggregation builds. When first requests hit API instances,
it's possible that pre-aggregations would still be being built and users would
need to wait or the completion of that process.
Depending on the volume of data and the configuration of pre-aggregations,
their builds might take from seconds to minutes.
Pre-aggregation warm-up will refresh pre-aggregations before a Cube Cloud
deployment is exposed to requests from users, making sure they aren't affected
by the pre-aggregation build time.
If a pre-aggregation warm-up takes more than 24 hours, it will be cancelled.
However, the deployment will still succeed.
### Configuring pre-aggregation warm-up
To configure pre-aggregation warm-up, navigate to **Settings → Configuration**
and enable **Warm-up pre-aggregations before deploying API**:
[ref-prod-cluster]: /admin/deployment/deployment-types#dedicated
[ref-prod-multi-cluster]: /admin/deployment/deployment-types#multi-cluster
[ref-data-model]: /docs/data-modeling/overview
[ref-dynamic-data-model]: /docs/data-modeling/dynamic
[ref-multitenancy]: /embedding/multitenancy
[ref-pre-aggs]: /docs/pre-aggregations#pre-aggregations
# Alerts
Source: https://docs.cube.dev/admin/monitoring/alerts
Set up email alerts for API outages, database timeouts, pre-aggregation failures, and build completions in Cube.
Alerts notify you by email when something happens in your account: an API goes
down, a database stops responding in time, a pre-aggregation build fails, or a
build finishes.
Available on [Premium and above plans](https://cube.dev/pricing).
## Manage alerts
Click **Alerts** in the sidebar to see every alert configured on the account. Click
**New alert** to add one, or use the edit and delete icons on a row to change or remove
an existing alert.
On plans below [Enterprise](https://cube.dev/pricing), only account administrators can
manage alerts. On the Enterprise plan, access also follows the `AlertsCreate`,
`AlertsRead`, `AlertsUpdate`, and `AlertsDelete` actions. Administrators always have
access, and the built-in Developer and AIBI Developer roles carry all four. The same four
actions also govern [budgets](/admin/account-billing/budgets).
These four actions are not among the ones you can pick when you build a [custom
role][ref-custom-roles]. To grant them, assign a built-in Developer or AIBI Developer
role.
## Event types
Each alert watches a single event type:
| Event type | What it detects | Resolved email |
| ---------------------------------- | ------------------------------------------------------------------- | -------------- |
| **API outages** | The API stops responding | Yes |
| **Database response timeouts** | The database takes too long to answer | Yes |
| **Pre-aggregation build failures** | A [pre-aggregation](/admin/monitoring/pre-aggregations) build fails | Yes |
| **Build completed** | A build reaches a terminal state | No |
| **All** | Every event type above | Per event type |
The first three are conditions: Cube emails you when the condition starts, then emails
you again with a `Resolved —` subject prefix when it clears. While a condition persists,
Cube does not re-send the same alert for a while — 1 hour for API outages and database
response timeouts, 3 hours for pre-aggregation build failures.
**Build completed** fires when a build finishes, whether it succeeded or failed. The
subject line reads `Build finished with status: `. It is a point-in-time event,
so it has no resolved email.
## Deployments
An alert applies either to **All** deployments in the account, or to a **Specific** set
you pick. Choosing **Specific** requires at least one deployment.
## Recipients
Under **Send alerts to**, pick either **All users on this account** or **Specific users**.
Under **Also send to**, add any number of custom email addresses; these are additive, and
a custom address on its own is a valid set of recipients.
**All users on this account** only reaches users who have signed in at least once. Users
who were invited but never signed in do not receive alerts.
## Delivery
Alerts are delivered by email only — there is no Slack, webhook, or PagerDuty delivery.
[Scheduled refresh notifications](/docs/explore-analyze/notifications), which cover
dashboard refresh outcomes, are a separate feature with its own delivery channels.
To alert from your own observability stack instead, export telemetry with [monitoring
integrations](/admin/monitoring/monitoring-integrations).
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
# Audit Log
Source: https://docs.cube.dev/admin/monitoring/audit-log
Learn how to turn on Audit Log in Cube Cloud, review security events across deployments, and download them for compliance reviews.
Audit Log collects, stores, and displays security-related events within a Cube Cloud
account, across all deployments. You can use it to maintain and review a historical
record of activity for compliance purposes.
Available on the [Enterprise plan](https://cube.dev/pricing).
Read below about [collected events](#event-types). Also, see how you can [enable](#configuration)
Audit Log, [view events](#viewing-events), and [download](#downloading-events) them.
## Configuration
To enable Audit Log, navigate to the **Team & Security** page, open the
**Audit Log** tab, and turn the **Enable Audit Log** switch on.
## Viewing events
### Table view
Audit Log displays events in a table with the following information for each event:
* Event timestamp in your local time zone.
* User email, if applicable to a particular event.
* Event name (see [event types](#event-types) for the full list).
* Deployment name, if applicable to a particular event.
### Extended view
You can click on any event to view extended information:
* IP address from which an event was initiated.
* Event-specific attributes.
#### Sanitization
Audit Log uses heuristics to detect sensitive values (e.g., passwords and tokens)
and sanitize them, i.e., replace them with `[HIDDEN]` in event-specific
attributes. You can see that a password was sanitized on the screenshot above.
If you'd like your custom environment variable to be sanitized, include one of
the following substrings in its name: `PASS`, `SECRET`, `TOKEN`, `KEY`.
### Filters
You can also customize the table view with filters on the top and in the right sidebar:
* Text input for full-text search.
* Date range filter.
* Filters to narrow the view down to a specific user, event, or deployment.
## Downloading events
You can use the **↓ CSV** button to download a CSV file with all events in the
current view:
## Exporting events via API
You can also export audit log events as a CSV file programmatically using the
[`/api/v1/audit-logs/export`][ref-audit-log-api] endpoint of the
[Control Plane API][ref-control-plane-api]. This is useful for integrating
with external log aggregation or SIEM tools.
## Event types
Audit Log collects the following types of events:
| Category | Event type |
| ---------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Users | `Created user` `Deleted user` `Invited user` `Updated user profile` `Requested password reset` `Changed password` |
| SCIM provisioning | `Created user via SCIM` `Updated user via SCIM` `Deleted user via SCIM` `Created group via SCIM` `Updated group via SCIM` `Deleted group via SCIM` |
| [Custom roles](/admin/users-and-permissions/custom-roles) | `Created role` `Deleted role` `Updated role` `Assigned role to user` `Assigned roles to user` `Removed role from user` |
| Authentication | `Logged in` `Logged out` `Redirected to SAML provider for login` `Logged in via SAML` |
| Account and [single sign-on](/admin/sso) | `Updated account settings` `Updated account authentication configuration` `Updated account SAML configuration` `Uploaded identity provider metadata file` |
| [Deployments](/admin/deployment) | `Created deployment` `Deleted deployment` `Updated deployment configuration` `Enabled SQL API for deployment` `Generated data model` |
| Branches and [dev mode](/docs/data-modeling/dev-mode) | `Created Git branch` `Deleted Git branch` `Entered dev mode` `Exited dev mode` `Pulled changes to data model` `Committed and merged changes to data model` `Merged changes to data model` `Reverted changes to data model in branch` |
| [Continuous deployment](/admin/deployment/continuous-deployment) | `Requested connection of GitHub account` `Connected deployment to GitHub repository` `Generated SSH key for deployment` `Connected deployment to Git upstream` `Generated Git credentials for Cube Cloud repository` `Generated webhook token for Git repository` `Received a Git hook` `Started uploading files` `Finished uploading files` `Triggered new build` |
| [Budgets](/admin/account-billing/budgets) | `Created budget` `Deleted budget` `Updated budget` |
[ref-control-plane-api]: /reference/control-plane-api
[ref-audit-log-api]: /reference/control-plane-api#%2Fapi%2Fv1%2Faudit-logs%2Fexport
# Chats History
Source: https://docs.cube.dev/admin/monitoring/chats-history
Admins can access users chat history through Admin → Chats History.
It's possible to search specific chats by the first message content or chat UUID.
You can also filter the list by user type, creation date, and source
(**Web**, **MCP**, **Slack**, **Scheduled**, or **API**).
# Integration with Amazon CloudWatch
Source: https://docs.cube.dev/admin/monitoring/monitoring-integrations/cloudwatch
Export Cube Cloud deployment logs to Amazon CloudWatch using Vector and an aws_cloudwatch_logs sink configuration.
[Amazon CloudWatch](https://aws.amazon.com/cloudwatch/) is an application
monitoring service that collects and visualizes logs, metrics, and event data.
This guide demonstrates how to set up Cube Cloud to export logs to Amazon
CloudWatch.
## Configuration
First, enable [monitoring integrations][ref-monitoring-integrations] in Cube
Cloud.
### Exporting logs
To export logs to Amazon CloudWatch, start by creating [a log group and a log
stream](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/Working-with-log-groups-and-streams.html)
for Cube Cloud logs.
Then, configure the [`aws_cloudwatch_logs`](https://vector.dev/docs/reference/configuration/sinks/aws_cloudwatch_logs/)
sink in your [`vector.toml` configuration file][ref-monitoring-integrations-conf].
Example configuration:
```toml theme={"dark"}
[sinks.aws_cloudwatch_logs]
type = "aws_cloudwatch_logs"
inputs = [
"cubejs-server",
"refresh-scheduler",
"warmup-job",
"cubestore"
]
region = "us-east-1"
group_name = "your-group-name"
stream_name = "your-stream-name"
create_missing_group = true
create_missing_stream = true
[sinks.aws_cloudwatch_logs.auth]
access_key_id = "$CUBE_CLOUD_MONITORING_AWS_ACCESS_KEY_ID"
secret_access_key = "$CUBE_CLOUD_MONITORING_AWS_SECRET_ACCESS_KEY"
[sinks.aws_cloudwatch_logs.encoding]
codec = "json"
```
Commit the configuration for Vector, it should take effect in a minute. Then,
navigate to Amazon CloudWatch and watch the logs coming.
## Keyless authentication
Instead of the access key pair above, Vector can authenticate to CloudWatch with
the deployment's [OIDC][ref-oidc] identity — no static credentials stored in
Cube, and nothing to rotate.
The credentials apply to the **whole Vector agent**, not just this sink, so the
same role also covers the [S3][ref-s3] and [Query History export][ref-query-history-export]
sinks if you use them.
Requires [OIDC][ref-oidc] enabled for your tenant with an `AWS` token config.
[Register Cube as an OIDC provider][ref-oidc-aws-provider] in your AWS
account — a one-time step per account — then follow
[Monitoring integrations][ref-oidc-aws-cloudwatch] to create a role that
trusts your deployment and can write to the log group.
Two things to carry over from that section: the `sub` claim needs the
non-default `:component:` **Subject Claim Format**, and
`logs:DescribeLogGroups` must be scoped to `"*"` or Vector's healthcheck
fails at startup.
Under **Settings → Environment variables**:
```dotenv theme={"dark"}
CUBE_CLOUD_MONITORING_AWS_ROLE_ARN=arn:aws:iam:::role/
```
The role ARN is the only required variable — setting it is what turns
keyless authentication on.
The credential exchange also needs a region, which Cube takes from your
sink's `region`. Set `CUBE_CLOUD_MONITORING_AWS_REGION` only to override
that, for example when the STS region should differ from the sink's.
Omit `[sinks.*.auth]` entirely so Vector falls back to the AWS SDK's default
credential chain, which picks up the OIDC token Cube mounts for it:
```toml theme={"dark"}
[sinks.aws_cloudwatch_logs]
type = "aws_cloudwatch_logs"
# "warmup-job" is omitted deliberately — see the warning below.
inputs = [
"cubejs-server",
"refresh-scheduler",
"cubestore"
]
region = "us-east-1"
group_name = "your-group-name"
stream_name = "your-stream-name"
# Add create_missing_group = true only if you also grant
# logs:CreateLogGroup to the role.
create_missing_stream = true
[sinks.aws_cloudwatch_logs.encoding]
codec = "json"
```
Leaving an `auth` block in place keeps the static keys in use — the two are
mutually exclusive.
Keyless authentication does not cover the **`warmup-job`** input. The
[pre-aggregation warm-up][ref-preagg-warmup] runs as a Kubernetes Job that does
not receive the OIDC token, so its logs will not reach CloudWatch on this path
while the other inputs succeed. If you need warm-up logs in CloudWatch, keep
using the access key pair.
[ref-monitoring-integrations]: /admin/monitoring/monitoring-integrations
[ref-monitoring-integrations-conf]: /admin/monitoring/monitoring-integrations#configuration
[ref-oidc]: /admin/deployment/oidc
[ref-oidc-aws-cloudwatch]: /admin/deployment/oidc/aws#monitoring-integrations-cloudwatch-s3
[ref-oidc-aws-provider]: /admin/deployment/oidc/aws#step-1-register-cube-as-an-oidc-provider-in-aws
[ref-s3]: /admin/monitoring/monitoring-integrations/s3
[ref-query-history-export]: /admin/monitoring/query-history-export
[ref-preagg-warmup]: /admin/deployment/warm-up
# Integration with Datadog
Source: https://docs.cube.dev/admin/monitoring/monitoring-integrations/datadog
Datadog is a popular fully managed observability service. This guide demonstrates how to set up Cube Cloud to export logs and metrics to Datadog.
[Datadog][datadog] is a popular fully managed observability service. This guide
demonstrates how to set up Cube Cloud to export logs and metrics to Datadog.
## Configuration
First, enable [monitoring integrations][ref-monitoring-integrations] in Cube
Cloud.
### Exporting logs
To export logs to Datadog, go to **Organization Settings → API Keys**
obtain an API key:
Then, configure the [`datadog_logs`][vector-datadog-logs] sink in your
[`vector.toml` configuration file][ref-monitoring-integrations-conf].
Example configuration:
```toml theme={"dark"}
[sinks.datadog_logs]
type = "datadog_logs"
inputs = [
"cubejs-server",
"refresh-scheduler",
"warmup-job",
"cubestore"
]
default_api_key = "$CUBE_CLOUD_MONITORING_DATADOG_API_KEY"
site = "datadoghq.eu"
compression = "gzip"
healthcheck = false
```
Note that Datadog accounts belong to specific [sites][datadog-docs-sites]
throughout the world. Use the `site` option to configure the sink appropriately.
When miscofigured, Vector agent outputs the following error:
`Client request was forbidden`.
Commit the configuration for Vector, it should take effect in a minute. Then,
navigate to **Logs** in Datadog and watch the logs coming:
### Exporting metrics
To export metrics to Datadog, use the same API key from
**Organization Settings → API Keys** as configured for logs.
Then, configure the [`datadog_metrics`][vector-datadog-metrics] sink in your
[`vector.toml` configuration file][ref-monitoring-integrations-conf].
Example configuration:
```toml theme={"dark"}
[sinks.datadog_metrics]
type = "datadog_metrics"
inputs = [
"metrics"
]
default_api_key = "$CUBE_CLOUD_MONITORING_DATADOG_API_KEY"
site = "datadoghq.eu"
```
Again, upon commit the configuration for Vector should take effect in a minute. Then,
navigate to **Metrics → Summary** in Datadog and explore the available
metrics. Cube metrics are prefixed with `cube_`, such as `cube_cpu_usage_ratio`,
`cube_memory_usage_ratio`, and `cube_requests_total`.
[datadog]: https://www.datadoghq.com
[datadog-docs-sites]: https://docs.datadoghq.com/getting_started/site/
[vector-datadog-logs]: https://vector.dev/docs/reference/configuration/sinks/datadog_logs/
[vector-datadog-metrics]: https://vector.dev/docs/reference/configuration/sinks/datadog_metrics/
[ref-monitoring-integrations]: /admin/monitoring/monitoring-integrations
[ref-monitoring-integrations-conf]: /admin/monitoring/monitoring-integrations#configuration
# Integration with Google Cloud Storage
Source: https://docs.cube.dev/admin/monitoring/monitoring-integrations/gcs
Google Cloud Storage is a popular object storage system. This guide demonstrates how to set up Cube to export logs to Google Cloud Storage.
[Google Cloud Storage](https://cloud.google.com/storage) is a popular object
storage system. This guide demonstrates how to set up Cube to export logs to
Google Cloud Storage.
## Configuration
First, enable [monitoring integrations][ref-monitoring-integrations] in Cube.
### Exporting logs
To export logs to Google Cloud Storage, start by creating a bucket for Cube logs
and a service account that can write to it.
Then, put the service account key in an environment variable under **Settings →
Environment variables**, base64-encoded. The name must start with
`CUBE_CLOUD_MONITORING_` — only variables with that prefix are available to the
Vector agent:
```bash theme={"dark"}
CUBE_CLOUD_MONITORING_GCS_CREDENTIALS=eyJ0eXBlIjogInNlcnZpY2VfYWNjb3VudCIsIC4uLn0=
```
Finally, configure the [`gcp_cloud_storage`](https://vector.dev/docs/reference/configuration/sinks/gcp_cloud_storage/)
sink in your [`vector.toml` configuration file][ref-monitoring-integrations-conf].
Example configuration:
```toml theme={"dark"}
[sinks.gcp-cloud-storage]
type = "gcp_cloud_storage"
inputs = [
"cubejs-server",
"refresh-scheduler",
"warmup-job",
"cubestore"
]
bucket = "your-gcs-bucket-name"
compression = "gzip"
credentials_base64 = "$CUBE_CLOUD_MONITORING_GCS_CREDENTIALS"
[sinks.gcp-cloud-storage.encoding]
codec = "json"
[sinks.gcp-cloud-storage.healthcheck]
enabled = false
```
Commit the configuration for Vector, it should take effect in a minute. Then,
navigate to your bucket and watch the logs coming.
### Authentication
Authenticate with the `credentials_base64` option, referencing the environment
variable that holds the base64-encoded service account key, as in the example above.
This is the only supported way to give the sink its credentials.
`credentials_base64` is specific to Cube — it does not appear in [Vector's own sink
reference][vector-docs-sinks-gcs]. Cube decodes the key and provides it to the Vector
agent as a file named after the sink.
The sink name becomes the name of a Kubernetes object that carries the credentials
file, so it must be lowercase alphanumeric characters and dashes only — **no
underscores**. A sink named `query_history_gcs` will not get a credentials file; name
it `query-history-gcs` instead.
### Healthcheck
The example above sets `healthcheck.enabled = false`. Vector's
`gcp_cloud_storage` healthcheck sends a `HEAD` request to the bucket root, which
requires the `storage.objects.list` permission — write access alone, such as
`roles/storage.objectCreator`, does not grant it. Without this, the sink fails to
start with a forbidden-healthcheck error even though it could write objects
successfully.
With the healthcheck disabled, the sink always starts and reports healthy, so check
the bucket itself to confirm that data is arriving.
### Exporting Query History
Add the `query-history` input to the sink to bring [Query History
export][ref-query-history-export] data to the same bucket.
Query History export additionally requires the **Monitoring Integrations Tier** of
your deployment to be set to **Medium (Up to 50 GB/mo)**. On a lower tier it fails
silently — the sink reports healthy and the bucket stays empty, with no error
logged anywhere.
[ref-monitoring-integrations]: /admin/monitoring/monitoring-integrations
[ref-monitoring-integrations-conf]: /admin/monitoring/monitoring-integrations#configuration
[ref-query-history-export]: /admin/monitoring/monitoring-integrations#query-history-export
[vector-docs-sinks-gcs]: https://vector.dev/docs/reference/configuration/sinks/gcp_cloud_storage/
# Integration with Grafana Cloud
Source: https://docs.cube.dev/admin/monitoring/monitoring-integrations/grafana-cloud
Grafana Cloud is a popular fully managed observability service. This guide demonstrates how to set up Cube Cloud to export logs to Grafana Cloud.
[Grafana Cloud][grafana] is a popular fully managed observability service. This
guide demonstrates how to set up Cube Cloud to export logs to Grafana Cloud.
## Configuration
First, enable [monitoring integrations][ref-monitoring-integrations] in Cube
Cloud.
### Exporting logs
To export logs to Grafana Cloud, go to your account and get credentials for
Loki, the logging service in Grafana Cloud:
Then, configure the [`loki`][vector-loki] sink in your [`vector.toml`
configuration file][ref-monitoring-integrations-conf].
Example configuration:
```toml theme={"dark"}
[sinks.loki]
type = "loki"
inputs = [
"cubejs-server",
"refresh-scheduler",
"warmup-job",
"cubestore"
]
endpoint = "https://logs-prod-012.grafana.net"
[sinks.loki.auth]
strategy = "basic"
user = "$CUBE_CLOUD_MONITORING_GRAFANA_CLOUD_USER"
password = "$CUBE_CLOUD_MONITORING_GRAFANA_CLOUD_PASSWORD"
[sinks.loki.encoding]
codec = "json"
[sinks.loki.cubestore]
levels = [
"trace",
"info",
"debug",
"error"
]
[sinks.loki.labels]
app = "cube-cloud"
env = "production"
```
Commit the configuration for Vector, it should take effect in a minute. Then,
navigate to **Logs** in Grafana Cloud and watch the logs coming.
### Exporting metrics
To export metrics to Grafana Cloud, go to your account and get credentials for
Prometheus, the metrics service in Grafana Cloud:
Then, configure the [`prometheus_remote_write`][vector-prometheus-rw] sink in
your [`vector.toml` configuration file][ref-monitoring-integrations-conf].
Example configuration:
```toml theme={"dark"}
[sinks.prometheus]
type = "prometheus_remote_write"
inputs = [
"metrics"
]
endpoint = "https://prometheus-prod-24-prod-eu-west-2.grafana.net/api/prom/push"
[sinks.prometheus.auth]
strategy = "basic"
user = "1033221"
password = "eyJrIjoiYTg1OTQ2OGY4Yzg3MTQxODc5OTA4NDUxMGM4NTA2ZDQ3ZjliYWZjOCIsIm4iOiJwcnciLCJpZCI6ODc1NzE5fQ=="
```
Commit the configuration for Vector, it should take effect in a minute. Then,
navigate to **Explore** in Grafana Cloud, select metrics from the drop
down:
Then, you can create visualizations and add them to a dashboard:
[grafana]: https://grafana.com
[vector-loki]: https://vector.dev/docs/reference/configuration/sinks/loki/
[vector-prometheus-rw]: https://vector.dev/docs/reference/configuration/sinks/prometheus_remote_write/
[ref-monitoring-integrations]: /admin/monitoring/monitoring-integrations
[ref-monitoring-integrations-conf]: /admin/monitoring/monitoring-integrations#configuration
# Overview
Source: https://docs.cube.dev/admin/monitoring/monitoring-integrations/index
Export Cube Cloud logs and metrics to external monitoring tools like Datadog, Grafana Cloud, and New Relic.
Cube Cloud allows exporting logs and metrics to external monitoring tools so you
can leverage your existing monitoring stack and retain logs and metrics for the
long term.
Available as an add-on on the [Enterprise plan](https://cube.dev/pricing).
Monitoring integrations suspend their work when a deployment goes to [auto-suspension][ref-autosuspend].
Monitoring integrations are only available for [production environments][ref-prod-env].
Under the hood, Cube Cloud uses [Vector][vector], an open-source tool for
collecting and delivering monitoring data. It supports a [wide range of
destinations][vector-docs-sinks], also known as *sinks*.
## Guides
Monitoring integrations work with various popular monitoring tools. Check the
following guides and configuration examples to get tool-specific instructions:
Export logs and metrics to Amazon CloudWatch.
Archive logs to an Amazon S3 bucket.
Export logs and metrics to Datadog.
Archive logs to a Google Cloud Storage bucket.
Export logs and metrics to Grafana Cloud.
Export logs and metrics to New Relic.
## Configuration
To enable monitoring integrations, navigate to **Settings → Monitoring
Integrations** and click **Enable Vector** to add a Vector agent to
your deployment.
Under **Metrics export**, you will see credentials for the
`prometheus_exporter` sink, in case you'd like to setup [metrics
export][self-sinks-for-metrics].
Additionally, create a [`vector.toml` configuration file][vector-docs-config]
next to your `cube.js` file. This file is used to keep sinks configuration. You
have to commit this file to the main branch of your deployment for Vector
configuration to take effect.
### Environment variables
You can use environment variables prefixed with `CUBE_CLOUD_MONITORING_` to
reference configuration parameters securely in the `vector.toml` file.
Example configuration for exporting logs to
[Datadog][vector-docs-sinks-datadog]:
```toml theme={"dark"}
[sinks.datadog]
type = "datadog_logs"
default_api_key = "$CUBE_CLOUD_MONITORING_DATADOG_API_KEY"
```
### Inputs for logs
Sinks accept the `inputs` option that allows to specify which components of a
Cube Cloud deployment should export their logs:
| Input name | Description |
| ------------------- | -------------------------------------------------------- |
| `cubejs-server` | Logs of API instances |
| `refresh-scheduler` | Logs of the refresh worker |
| `warmup-job` | Logs of the [pre-aggregation warm-up][ref-preagg-warmup] |
| `cubestore` | Logs of Cube Store |
| `query-history` | [Query History export](#query-history-export) |
Example configuration for exporting logs to
[Datadog][vector-docs-sinks-datadog]:
```toml theme={"dark"}
[sinks.datadog]
type = "datadog_logs"
inputs = [
"cubejs-server",
"refresh-scheduler",
"warmup-job",
"cubestore"
]
default_api_key = "da8850ce554b4f03ac50537612e48fb1"
compression = "gzip"
```
When exporting Cube Store logs using the `cubestore` input, you can filter logs
by providing an array of their severity levels via the `levels` option. If not
specified, only `error` and `info` logs will be exported.
| Level | Exported by default? |
| ------- | -------------------- |
| `error` | ✅ Yes |
| `info` | ✅ Yes |
| `debug` | ❌ No |
| `trace` | ❌ No |
If you'd like to adjust severity levels of logs from API instances and the
refresh scheduler, use the [`CUBEJS_LOG_LEVEL`](/reference/configuration/environment-variables#cubejs_log_level) environment variable.
### Fields in exported logs
Before a log record reaches a sink, the following fields are attached to it:
| Field | Description |
| ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `deployment_id` | Identifier of the [deployment][ref-deployments]. |
| `deployment_name` | Name of the deployment. |
| `hostname` | Host name of the node that the record comes from. |
| `pod_name` | Name of the pod that the record comes from. |
| `service` | Name of the component that the record comes from, matching the input name: `cubejs-server`, `refresh-scheduler`, `warmup-job`, or `cubestore`. Records from the `query-history` input use `query-history` on `datadog_logs` sinks. |
| `level` | Severity of the record, lower-cased: `error`, `warn`, `info`, or `debug`. Records with the `trace` severity are exported as `debug`, and records that carry no severity as `info`. |
| `status` | Duplicate of `level`, for tools that expect the severity in a `status` field, e.g. Datadog. |
| `origin` | Always `cube_cloud`. |
Records from the `query-history` input carry a different set of fields, see
[Query History export](#query-history-export).
#### Datadog attributes
Sinks of the `datadog_logs` type additionally get the two fields that Datadog
uses for attribution:
| Field | Description |
| ---------- | ---------------------------------------------------------------------------------------------------------- |
| `ddsource` | Name of the pod that the record comes from, or `query-history` for records from the `query-history` input. |
| `ddtags` | Comma-separated list of tags. By default, the deployment name and the host name. |
To send your own tags to Datadog, list them in the `ddtags` option of the sink.
They are appended to the default ones:
```toml theme={"dark"}
[sinks.datadog]
# type, inputs, default_api_key, etc.
ddtags = [
"env:production",
"team:analytics"
]
```
### Sinks for logs
You can use a [wide range of destinations][vector-docs-sinks] for logs,
including the following ones:
* [AWS Cloudwatch][vector-docs-sinks-cloudwatch]
* [AWS S3][vector-docs-sinks-s3], [Google Cloud Storage][vector-docs-sinks-gcs],
and [Azure Blob Storage][vector-docs-sinks-azureblob]
* [Datadog][vector-docs-sinks-datadog]
Example configuration for exporting all logs, including all Cube Store logs to
[Azure Blob Storage][vector-docs-sinks-azureblob]:
```toml theme={"dark"}
[sinks.azure]
type = "azure_blob"
container_name = "my-logs"
connection_string = "DefaultEndpointsProtocol=https;AccountName=mylogstorage;AccountKey=storageaccountkeybase64encoded;EndpointSuffix=core.windows.net"
inputs = [
"cubejs-server",
"refresh-scheduler",
"warmup-job",
"cubestore"
]
[sinks.azure.cubestore]
levels = [
"trace",
"info",
"debug",
"error"
]
```
### Inputs for metrics
Metrics are exported using the `metrics` input. Metrics will have their respective
metric names and\_types: [`gauge`][vector-docs-metrics-gauge] or
[`counter`][vector-docs-metrics-counter].
All metrics of the `counter` type reset to zero at the midnight (UTC) and increment
during the next 24 hours.
You can filter metrics by providing an array of *input names* via the `list` option.
| Input name | Metric name, type | Description |
| --------------------------- | ---------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `cpu` | `cube_cpu_usage_ratio`, `gauge` | CPU usage of a particular node in the deployment. Usually, a number in the 0—100 range. May exceed 100 if the node is under load |
| `memory` | `cube_memory_usage_ratio`, `gauge` | Memory usage of a particular node in the deployment. Usually, a number in the 0—100 range. May exceed 100 if the node is under load |
| `requests-count` | `cube_requests_total`, `counter` | Number of API requests to the deployment |
| `requests-success-count` | `cube_requests_success_total`, `counter` | Number of successful API requests to the deployment |
| `requests-errors-count` | `cube_requests_errors_total`, `counter` | Number of errorneous API requests to the deployment |
| `requests-duration` | `cube_requests_duration_ms_total`, `counter` | Total time taken to process API requests, milliseconds |
| `requests-success-duration` | `cube_requests_duration_ms_success`, `counter` | Total time taken to process successful API requests, milliseconds |
| `requests-errors-duration` | `cube_requests_duration_ms_errors`, `counter` | Total time taken to process errorneous API requests, milliseconds |
You can further filter exported metrics by providing an array of `inputs`. It applies to
metics only.
Example configuration for exporting all metrics from `cubejs-server` to
[Prometheus][vector-docs-sinks-prometheus] using the `prometheus_remote_write`
sink:
```toml theme={"dark"}
[sinks.prometheus]
type = "prometheus_remote_write"
inputs = [
"metrics"
]
endpoint = "https://prometheus.example.com:8087/api/v1/write"
[sinks.prometheus.auth]
# Strategy, credentials, etc.
[sinks.prometheus.metrics]
list = [
"cpu",
"memory",
"requests-count",
"requests-errors-count",
"requests-success-count",
"requests-duration"
]
inputs = [
"cubejs-server"
]
```
### Labels on exported metrics
Metrics are exported with the following Prometheus labels:
| Label | Metrics | Description |
| ---------------- | ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `pod_name` | `cube_cpu_usage_ratio`, `cube_memory_usage_ratio` | Name of the pod that the measurement was taken on. |
| `deployment_id` | `cube_requests_*` | Identifier of the [deployment][ref-deployments]. |
| `api_type` | `cube_requests_*` | Type of [data API][ref-apis] that served the requests (`rest`, `sql`, etc.), the same values as in [Query History][ref-query-history]. `unknown` if the API type is not known. |
| `request_source` | `cube_requests_*` | Source of the requests: `ai-engineer` for requests coming from AI features, `other` for any other source, `unknown` if the requests carry no source. |
Request metrics are reported per `api_type` and `request_source` combination, so
you can break down request counts and latency by data API:
```text theme={"dark"}
cube_requests_total{deployment_id="12345",api_type="rest",request_source="unknown"} 42
cube_requests_total{deployment_id="12345",api_type="sql",request_source="ai-engineer"} 7
```
### Sinks for metrics
Metrics are exported in the Prometheus format which is compatible with the
following sinks:
* [`prometheus_exporter`][vector-docs-sinks-prometheus-exporter] (native to
[Prometheus][prometheus], compatible with [Mimir][mimir])
* [`prometheus_remote_write`][vector-docs-sinks-prometheus] (compatible with
[Grafana Cloud][grafana-cloud])
Example configuration for exporting all metrics from `cubejs-server` to
[Prometheus][vector-docs-sinks-prometheus-exporter] using the
`prometheus_exporter` sink:
```toml theme={"dark"}
[sinks.prometheus]
type = "prometheus_exporter"
inputs = [
"metrics"
]
[sinks.prometheus.metrics]
list = [
"cpu",
"memory",
"requests-count",
"requests-errors-count",
"requests-success-count",
"requests-duration"
]
inputs = [
"cubejs-server"
]
```
Navigate to **Settings → Monitoring Integrations** to take the
credentials `prometheus_exporter` under **Metrics export**:
You can also customize the user name and password for `prometheus_exporter` by
setting `CUBE_CLOUD_MONITORING_METRICS_USER` and
`CUBE_CLOUD_MONITORING_METRICS_PASSWORD` environment variables, respectively.
## Query History export
With Query History export, you can bring [Query History][ref-query-history] data to an
external monitoring solution for further analysis, for example:
* Detect queries that do not hit pre-aggregations.
* Set up alerts for queries that exceed a certain duration.
* Attribute usage to specific users and implement chargebacks.
Query History export is part of the Monitoring Integrations add-on,
available on the [Enterprise plan](https://cube.dev/pricing).
Query History export also requires the **Monitoring Integrations Tier** of your
deployment to be set to **Medium (Up to 50 GB/mo)**. You can find it under
**Settings → Monitoring Integrations**; deployments default to **X-Small (Up to 10
GB/mo)**.
On **X-Small** or **Small**, Query History export fails silently: no events reach
any sink, including a `console` sink, and no error is logged. If the tier is not
available in your deployment settings, ask your Cube contact or the [Cube support
team][ref-support] to set it.
To configure Query History export, add the `query-history` input to the `inputs`
option of the sink configuration. Example configuration for exporting Query History data
to the standard output of the Vector agent:
```toml theme={"dark"}
[sinks.my_console]
type = "console"
inputs = [
"query-history"
]
target = "stdout"
encoding = { codec = "json" }
```
Exported data includes the following fields:
| Field | Description |
| -------------------------- | --------------------------------------------------------------------------------- |
| `trace_id` | Unique identifier of the API request. |
| `account_name` | Name of the Cube Cloud account. |
| `query_history_link` | Link to the request in [Query History][ref-query-history]. |
| `deployment_id` | Identifier of the [deployment][ref-deployments]. |
| `environment_name` | Name of the [environment][ref-environments], `NULL` for production. |
| `api_type` | Type of [data API][ref-apis] used (`rest`, `sql`, etc.), `NULL` for errors. |
| `api_query` | Query executed by the API, represented as string. |
| `security_context` | [Security context][ref-security-context] of the request, represented as a string. |
| `status` | Status of the request: `success` or `error`. |
| `error_message` | Error message, if any. |
| `start_time_unix_ms` | Start time of the execution, Unix timestamp in milliseconds. |
| `end_time_unix_ms` | End time of the execution, Unix timestamp in milliseconds. |
| `api_response_duration_ms` | Duration of the execution in milliseconds. |
| `cache_type` | [Cache type][ref-cache-type]: `no_cache`, `pre_aggregations_in_cube_store`, etc. |
Unlike other logs, Query History records don't carry the
[fields listed above](#fields-in-exported-logs), except for `origin`. On sinks of
the `datadog_logs` type, they also get `service` and `ddsource` set to
`query-history`, and `ddtags`.
See [this recipe][ref-query-history-export-recipe] for an example of analyzing data from
Query History export.
[ref-autosuspend]: /admin/deployment/auto-suspension#effects-on-experience
[self-sinks-for-metrics]: #configuration-sinks-for-metrics
[vector]: https://vector.dev/
[vector-docs-config]: https://vector.dev/docs/reference/configuration/
[vector-docs-sinks]: https://vector.dev/docs/reference/configuration/sinks/
[vector-docs-sinks-cloudwatch]: https://vector.dev/docs/reference/configuration/sinks/aws_cloudwatch_logs/
[vector-docs-sinks-s3]: https://vector.dev/docs/reference/configuration/sinks/aws_s3/
[vector-docs-sinks-azureblob]: https://vector.dev/docs/reference/configuration/sinks/azure_blob/
[vector-docs-sinks-gcs]: https://vector.dev/docs/reference/configuration/sinks/gcp_cloud_storage/
[vector-docs-sinks-datadog]: https://vector.dev/docs/reference/configuration/sinks/datadog_logs/
[vector-docs-sinks-prometheus]: https://vector.dev/docs/reference/configuration/sinks/prometheus_remote_write/
[vector-docs-sinks-prometheus-exporter]: https://vector.dev/docs/reference/configuration/sinks/prometheus_exporter/
[vector-docs-metrics-gauge]: https://vector.dev/docs/about/under-the-hood/architecture/data-model/metric/#gauge
[vector-docs-metrics-counter]: https://vector.dev/docs/about/under-the-hood/architecture/data-model/metric/#counter
[prometheus]: https://prometheus.io
[mimir]: https://grafana.com/oss/mimir/
[grafana-cloud]: https://grafana.com/products/cloud/
[ref-prod-env]: /admin/deployment/environments#production-environment
[ref-preagg-warmup]: /admin/deployment/warm-up#pre-aggregation-warm-up
[ref-query-history]: /admin/monitoring/query-history
[ref-deployments]: /admin/deployment
[ref-environments]: /admin/deployment/environments
[ref-apis]: /reference
[ref-security-context]: /docs/data-modeling/access-control/context
[ref-cache-type]: /docs/pre-aggregations#cache-type
[ref-query-history-export-recipe]: /admin/monitoring/query-history-export
[ref-support]: /admin/account-billing/support
# Integration with New Relic
Source: https://docs.cube.dev/admin/monitoring/monitoring-integrations/new-relic
Forward Cube Cloud logs and metrics to New Relic with Vector sinks targeting the logs and metrics APIs.
[New Relic](https://newrelic.com) is an all-in-one observability platform.
This guide demonstrates how to set up Cube Cloud to export logs and metrics
to New Relic.
## Configuration
First, enable [monitoring integrations][ref-monitoring-integrations] in Cube
Cloud.
### Exporting logs
To export logs to New Relic, configure the
[`new_relic`](https://vector.dev/docs/reference/configuration/sinks/new_relic/)
sink in your [`vector.toml` configuration file][ref-monitoring-integrations-conf].
Use the `logs` API endpoint.
Example configuration:
```toml theme={"dark"}
[sinks.new_relic_logs]
type = "new_relic"
api = "logs"
inputs = [
"cubejs-server",
"refresh-scheduler",
"warmup-job",
"cubestore"
]
region = "eu"
compression = "gzip"
account_id = "$CUBE_CLOUD_MONITORING_NEW_RELIC_ACCOUNT_ID"
license_key = "$CUBE_CLOUD_MONITORING_NEW_RELIC_LICENSE_KEY"
```
Commit the configuration for Vector, it should take effect in a minute. Then,
navigate to New Relic and watch the logs coming.
### Exporting metrics
To export metrics to New Relic, configure the
[`new_relic`](https://vector.dev/docs/reference/configuration/sinks/new_relic/)
sink in your [`vector.toml` configuration file][ref-monitoring-integrations-conf].
Use the `metrics` API endpoint.
Example configuration:
```toml theme={"dark"}
[sinks.new_relic_metrics]
type = "new_relic"
api = "metrics"
inputs = [
"metrics"
]
region = "eu"
compression = "gzip"
account_id = "$CUBE_CLOUD_MONITORING_NEW_RELIC_ACCOUNT_ID"
license_key = "$CUBE_CLOUD_MONITORING_NEW_RELIC_LICENSE_KEY"
```
Commit the configuration for Vector, it should take effect in a minute. Then,
navigate to New Relic and watch the metrics coming.
[ref-monitoring-integrations]: /admin/monitoring/monitoring-integrations
[ref-monitoring-integrations-conf]: /admin/monitoring/monitoring-integrations#configuration
# Integration with Amazon S3
Source: https://docs.cube.dev/admin/monitoring/monitoring-integrations/s3
Amazon S3 is a popular object storage system. This guide demonstrates how to set up Cube Cloud to export logs to Amazon S3.
[Amazon S3](https://aws.amazon.com/s3/) is a popular object storage system.
This guide demonstrates how to set up Cube Cloud to export logs to Amazon S3.
## Configuration
First, enable [monitoring integrations][ref-monitoring-integrations] in Cube
Cloud.
### Exporting logs
To export logs to Amazon S3, start by creating an S3 bucket for Cube Cloud logs.
Then, configure the [`aws_s3`](https://vector.dev/docs/reference/configuration/sinks/aws_s3/)
sink in your [`vector.toml` configuration file][ref-monitoring-integrations-conf].
Example configuration:
```toml theme={"dark"}
[sinks.aws_s3]
type = "aws_s3"
inputs = [
"cubejs-server",
"refresh-scheduler",
"warmup-job",
"cubestore"
]
bucket = "your-s3-bucket-name"
region = "us-east-1"
compression = "gzip"
[sinks.aws_s3.auth]
access_key_id = "$CUBE_CLOUD_MONITORING_AWS_ACCESS_KEY_ID"
secret_access_key = "$CUBE_CLOUD_MONITORING_AWS_SECRET_ACCESS_KEY"
[sinks.aws_s3.encoding]
codec = "json"
[sinks.aws_s3.healthcheck]
enabled = false
```
Commit the configuration for Vector, it should take effect in a minute. Then,
navigate to your S3 bucket and watch the logs coming.
The `aws_s3` sink can also authenticate with the deployment's OIDC identity
instead of an access key pair — the credentials apply to the whole Vector agent.
See [Keyless authentication][ref-keyless-auth].
## Troubleshooting
### Bucket region mismatch
You might see the following error message in Vector logs:
```text theme={"dark"}
The bucket you are attempting to access must be addressed using the specified
endpoint. Please send all future requests to this endpoint.
```
This often happens if the bucket on Amazon S3 is in a region that is different
from the one you’ve specified in the configuration for Vector.
[ref-monitoring-integrations]: /admin/monitoring/monitoring-integrations
[ref-monitoring-integrations-conf]: /admin/monitoring/monitoring-integrations#configuration
[ref-keyless-auth]: /admin/monitoring/monitoring-integrations/cloudwatch#keyless-authentication
# Performance Insights
Source: https://docs.cube.dev/admin/monitoring/performance
Use Performance Insights charts in Cube Cloud to interpret API load, queues, and resource behavior when tuning a deployment.
The **Performance** page in Cube Cloud displays charts that help
analyze the performance of your deployment and fine-tune its configuration.
It's recommended to review Performance Insights when the workload changes
or if you face any performance-related issues with your deployment.
Available on [Premium and above plans](https://cube.dev/pricing).
You can also choose a [Query History tier](/admin/account-billing/pricing#query-history-tiers).
## Charts
Charts provide insights into different aspects of your deployment.
### API instances
The **API instances** chart shows the number of API instances
that served queries to the deployment.
You can use this chart to **fine-tune the
[auto-scaling][ref-scalability-api] configuration of API instances**, e.g.,
increase the minimum and maximum number of API instances.
For example, the following chart shows a deployment with sane auto-scaling
limits that don't need adjusting. It looks like the deployment needs to
sustain just a few infrequent load bursts per day and auto-scaling to 3 API
instances does the job just fine:
The next chart shows a deployment with auto-scaling limits that definitely
need an adjustment. It looks like the load is so high that most of the time
this deployment has to use at least 4-6 API instances. So, it would be wise
to increase the minimum auto-scaling limit to 6 API instances:
When in doubt, consider using a higher minimum auto-scaling limit: when an
additional API instance starts, it needs some time to compile the data model
before it would be able to serve the requests. So, over-provisioning API
instances with a higher minimum auto-scaling limit would allow to decrease
the number of requests that had to wait for the [data model
compilation](#data-model-compilation).
Also, you can use this chart to **fine-tune the
[auto-suspension][ref-auto-sus] configuration**, e.g., by turning
auto-suspension off or increasing the auto-suspension threshold.
For example, the following chart shows a [Shared
deployment][ref-dev-instance] that is only accessed a few times
a day and automatically suspends after a short period of inactivity:
The next chart shows a misconfigured [Dedicated
deployment][ref-prod-cluster] that serves the requests throughout the whole
day but was configured to auto-suspend with a tiny threshold:
### Cache type
The **Requests by cache type** chart shows the number of API
requests that were fulfilled by specific [cache types][ref-cache-types],
e.g., pre-aggregations, in-memory cache, no cache, etc. For example, the
following chart shows a deployment that fulfills about 50% of requests by
using pre-aggregations:
The **Avg. response time by cache type** shows the difference
in the response time for requests that hit pre-aggregations, in-memory cache,
or no cache (i.e., the upstream data source). The next chart shows that
pre-aggregations usually provide sub-second response times while queries to
the data source take much longer:
You can use these charts to see if you'd like to have more queries that hit
the cache and have lower response time. In that case, **consider adding more
[pre-aggregations][ref-pre-aggregations] in Cube Store** or fine-tune the
existing ones, e.g., by **[using indexes][ref-indexes] to speed up
pre-aggregations with suboptimal query plans**.
### Data model compilation
The **Requests by data model compilation** chart shows the
number of API requests that had or had not to wait for the data model
compilation. For example, the following chart shows a deployment that
only has a tiny fraction of requests that require the data model to be
compiled:
The **Wait time for data model compilation** chart
shows the total time requests had to wait for the data model compilation.
The next chart shows that at certain points of time requests had to wait
dozens of seconds while the data model was being compiled:
You can use these charts to **fine-tune the [auto-suspension][ref-auto-sus]
configuration** (e.g., turn it off or increase the threshold so that API
instances suspend less frequently), **identify [multitenancy][ref-multitenancy]
misconfiguration** (e.g., suboptimal bucketing via
[`context_to_app_id`][ref-context-to-app-id]), or
**consider using a [Multi-cluster deployment][ref-multi-cluster]** to
distribute requests to different tenants over a number of Dedicated
deployments.
### Cube Store
The **Saturation for queries by Cube Store workers** chart
shows if Cube Store workers are overloaded with serving **queries**.
High saturation for queries prevents Cube Store workers from fulfilling
requests and results in wait time displayed at the **Wait time for
queries by Cube Store workers** chart.
For example, the following chart shows a deployment that uses 4 Cube Store
workers and almost never lets them come to saturation, resulting in no wait
time for queries:
Similarly, the **Saturation for jobs by Cube Store workers**
and **Wait time for jobs by Cube Store workers** charts show if
Cube Store Workers are overloaded with serving **jobs**, i.e., building
pre-aggregations or performing internal tasks such as data compaction.
For example, the following chart shows a misconfigured deployment that uses
8 Cube Store workers and keeps them at full saturation during prolonged
intervals, resulting in huge wait time and, in case of jobs, delayed refresh
of pre-aggregations:
The next chart shows that oversaturated Cube Store workers might yield
hours of wait time for queries and jobs:
You can use these charts to **fine-tune the [number of Cube Store
workers][ref-scalability-cube-store]** used by your deployment, e.g.,
increase it until you see that there's no saturation and no wait time
for queries and jobs.
[ref-scalability-api]: /admin/deployment/scalability#auto-scaling-of-api-instances
[ref-scalability-cube-store]: /admin/deployment/scalability#sizing-cube-store-workers
[ref-auto-sus]: /admin/deployment/auto-suspension
[ref-dev-instance]: /admin/deployment/deployment-types#shared
[ref-prod-cluster]: /admin/deployment/deployment-types#dedicated
[ref-multi-cluster]: /admin/deployment/deployment-types#multi-cluster
[ref-pre-aggregations]: /docs/pre-aggregations/using-pre-aggregations
[ref-multitenancy]: /embedding/multitenancy
[ref-context-to-app-id]: /reference/configuration/config#context_to_app_id
[ref-cache-types]: /docs/pre-aggregations#cache-type
[ref-indexes]: /docs/pre-aggregations/using-pre-aggregations#using-indexes
# Pre-Aggregations
Source: https://docs.cube.dev/admin/monitoring/pre-aggregations
Monitor pre-aggregation builds, partitions, and refresh history from the Cube Cloud Pre-Aggregations console.
The Pre-Aggregations page in Cube Cloud allows you to inspect all
[pre-aggregations][ref-caching-gs-preaggs] in your deployment at a glance. You
can see which pre-aggregations are accelerating queries, if they are [being
refreshed][ref-caching-using-preaggs-refresh], along with the last 24 hours of
build history.
## Exploring pre-aggregations
You can switch between [pre-aggregations](#pre-aggregations) and
[build history](#build-history) by switching between the
**Pre-Aggregations** and **Build History** tabs.
## Pre-aggregations
The **Pre-Aggregations** tab shows a list of all configured
[pre-aggregations][ref-caching-gs-preaggs] for the deployment, and is useful for
quickly checking details such as when the last build occurred, how long the
build took to complete, the total size of the pre-aggregated data, and how many
partitions the data is split across. To see more information about a specific
pre-aggregation, click on it.
### Partitions
The **Partitions** tab shows all partitions (as specified by
[`partitionGranularity`][ref-model-ref-preaggs-partition-granularity]) of the
pre-aggregation. In the example below, the pre-aggregation has a single
partition:
Pre-aggregations with multiple partitions show a list of all partitions, which
is helpful for checking pre-aggregation partition build times and sizes:
Ideally partitions should be of a similar size and build time. If a partition is
taking significantly longer to build than others, it may be worth investigating
why.
If the pre-aggregation requires rebuilding (usually due to the underlying data
changing in a significant way), clicking **Build All** rebuilds **all**
partitions of the pre-aggregation:
You can also choose to rebuild **specific** partitions by selecting them and
then clicking the **Build Selected** button:
### Definition
The **Definition** tab shows how the [pre-aggregation is defined in the
data model][ref-model-ref-preaggs] without having to switch to the **Data
Model** page:
### Preview
To see a sample of the data contained in the pre-aggregation, click the
**Preview** tab:
This can be used to verify that the data within the pre-aggregation is being
built correctly.
### Used By
The **Used By** tab shows recent queries that used this pre-aggregation
for acceleration:
This is especially useful for checking that the pre-aggregation is being used as
expected by queries sent to the deployment.
### Indexes
[Indexes][ref-model-ref-preaggs-index] significantly improve the performance of
configured pre-aggregations. The **Indexes** tab displays any configured
indexes for this pre-aggregation:
## Build History
The **Build History** tab shows the last 24 hours of build history for
all pre-aggregations in the deployment. From here, you can see if a
pre-aggregation build succeeded (or failed), how the build was triggered, when
the build started, and how long it took to complete.
To see more information about a specific build, click on it:
[ref-caching-gs-preaggs]: /docs/pre-aggregations/getting-started-pre-aggregations
[ref-caching-using-preaggs-refresh]: /docs/pre-aggregations/using-pre-aggregations#refresh-strategy
[ref-model-ref-preaggs]: /reference/data-modeling/pre-aggregations
[ref-model-ref-preaggs-index]: /reference/data-modeling/pre-aggregations#indexes
[ref-model-ref-preaggs-partition-granularity]: /reference/data-modeling/pre-aggregations#partition_granularity
# Query History
Source: https://docs.cube.dev/admin/monitoring/query-history
The Query History feature in Cube Cloud is a one-stop shop for all performance and diagnostic information about queries issued for a deployment.
It provides a real-time and historic view of requests to [data APIs][ref-apis] of your
Cube Cloud deployment, so you can check whether queries were accelerated with
[pre-aggregations][ref-caching-gs-preaggs], how long they took to execute, and if they
failed.
You can choose a [Query History tier](/admin/account-billing/pricing#query-history-tiers)
to fit your retention and throughput needs.
You can set the [time range](#setting-the-time-range), [explore queries](#exploring-queries)
and filter them, drill down on specific queries to [see more details](#inspecting-api-queries).
You can also use [Query History export][ref-query-history-export] to bring Query History
data to an external monitoring solution for further analysis.
## Setting the time range
By default, Cube Cloud shows you a live feed of queries made to the API and
connected [data sources][ref-conf-db].
You can navigate throughout the query history by using the date picker in the
top right corner of the page and selecting a time period:
To go back to live mode, click **▶**:
## Exploring queries
You can switch between [queries made to the API](#inspecting-api-queries) and
[queries made to connected data sources](#inspecting-database-queries) by
switching between the **API** and **Database** tabs.
### All queries and top queries
Clicking **All Queries** will show all queries in order of recency,
while **Top Queries** will show the queries with the largest total duration (hits \* avg duration) in the selected time frame. Both results are limited to 100 unique queries.
### Filtering
You can use filters to find queries by various criteria:
* duration,
* cache status,
* whether the query was accelerated with pre-aggregations,
* whether the query yielded an error,
* API type.
## Inspecting API queries
In the table, all queries are shown in their collapsed view by default.
To see an expanded view of a query, click on the **❯** button
to the left of any query.
Check the columns to see the details:
* **Query** column shows a representation of a query, similar to the
REST (JSON) API [query format][ref-query-format]. In case of the SQL API, if the
query is not coercible to a REST (JSON) API query, raw SQL is shown "as is."
* **API** column shows the API type that was used to run the query:
REST (JSON) API via HTTP transport, REST (JSON) API via WebSockets, GraphQL API, or SQL API.
* **Duration** column shows how long the query took.
* Bolt icon indicates the [cache type][ref-cache-types] that was used to
fulfill the query.
* **Time** column shows the point in time the query was run at.
To drill down on a specific query, click it to see more information.
### Query
The **Query** tab shows the raw JSON query sent to the Cube Cloud
deployment.
### Errors
If the query failed, the **Errors** tab will show you the error message
and stacktrace:
### SQL
The **SQL** tab shows the generated SQL query sent by Cube to either your
data source **or** Cube Store if the query was accelerated with a
pre-aggregation:
You can also run the SQL query directly against your data source using the [SQL
Runner][ref-workspace-sqlrunner] by clicking **▶️**:
### Pre-aggregations
The **Pre-Aggregations** tab shows the
[pre-aggregation][ref-caching-gs-preaggs] used to accelerate this query, if one
was used:
If no pre-aggregations were used for this query, you should see the following
screen:
Clicking **Accelerate** allows you to add a pre-aggregation to
accelerate similar queries in the future.
### Security context
The **Security Context** tab shows the [security context][ref-security-context]
that was used to process this query:
This is helpful in case [multitenancy][ref-multitenancy] is configured.
### Queue graph
The **Queue Graph** tab details any activity in the query queue while
processing the query. This may include other queries that were being processed
or were waiting in the queue by Cube Cloud at the same time as this query:
A large number of queries in the queue may indicate that your deployment is
under-provisioned, and you may want to consider scaling up your deployment.
### Flame graph
The **Flame Graph** tab shows a [flame graph][datadog-kb-flamegraph] of a
query's execution time across resources in the Cube Cloud deployment. This is
extremely useful for diagnosing where time is being spent in a query, and can
help identify bottlenecks in your Cube deployment or data source.
## Inspecting database queries
### Query
For Database requests, the **Query** tab shows the SQL query compiled by
Cube that is executed on the data source:
This can be useful for debugging queries that are failing or taking a long time,
as you can copy the query and run it directly against your data source.
### Errors
If the query failed, the **Errors** tab will show you the error message
and stacktrace:
Errors here generally indicate a problem with querying the data source. The
generated SQL query can be copied from the **[Query](#query-2)** tab and
run directly against your data source to debug the issue.
### Events
The **Events** tab shows all data source-related events that occurred
while the query is in the query execution queue:
[datadog-kb-flamegraph]: https://www.datadoghq.com/knowledge-center/distributed-tracing/flame-graph/
[ref-caching-gs-preaggs]: /docs/pre-aggregations/getting-started-pre-aggregations
[ref-conf-db]: /admin/connect-to-data/data-sources
[ref-workspace-sqlrunner]: /docs/data-modeling/sql-runner
[ref-query-format]: /reference/core-data-apis/rest-api/query-format
[ref-cache-types]: /docs/pre-aggregations#cache-type
[ref-security-context]: /docs/data-modeling/access-control/context
[ref-multitenancy]: /embedding/multitenancy
[ref-apis]: /reference
[ref-query-history-export]: /admin/monitoring/monitoring-integrations#query-history-export
# Analyzing data from Query History export
Source: https://docs.cube.dev/admin/monitoring/query-history-export
Walk through exporting Query History to Amazon S3 with Vector and analyzing the files with DuckDB inside Cube.
You can use [Query History export][ref-query-history-export] to bring [Query
History][ref-query-history] data to an external monitoring solution for further
analysis.
In this recipe, we will show you how to export Query History data to Amazon S3, and then
analyze it using Cube by reading the data from S3 using DuckDB.
Before you start, check that the **Monitoring Integrations Tier** of your deployment is
set to **Medium (Up to 50 GB/mo)**. On a lower tier, [Query History
export][ref-query-history-export] fails silently — no events reach any sink and no error
is logged.
## Configuration
[Vector configuration][ref-vector-configuration] for exporting Query History to Amazon S3
and also outputting it to the console of the Vector agent in your Cube Cloud deployment.
In the example below, we are using the `aws_s3` sink to export the `cube-query-history-export-demo`
bucket in Amazon S3, but you can use any other storage solution that Vector supports.
```toml theme={"dark"}
[sinks.aws_s3]
type = "aws_s3"
inputs = [
"query-history"
]
bucket = "cube-query-history-export-demo"
region = "us-east-2"
compression = "gzip"
[sinks.aws_s3.auth]
access_key_id = "$CUBE_CLOUD_MONITORING_AWS_ACCESS_KEY_ID"
secret_access_key = "$CUBE_CLOUD_MONITORING_AWS_SECRET_ACCESS_KEY"
[sinks.aws_s3.encoding]
codec = "json"
[sinks.aws_s3.healthcheck]
enabled = false
[sinks.my_console]
type = "console"
inputs = [
"query-history"
]
target = "stdout"
encoding = { codec = "json" }
```
You'd also need to set the following environment variables in the **Settings → Environment
variables** page of your Cube Cloud deployment:
```bash theme={"dark"}
CUBE_CLOUD_MONITORING_AWS_ACCESS_KEY_ID=your-access-key-id
CUBE_CLOUD_MONITORING_AWS_SECRET_ACCESS_KEY=your-secret-access-key
CUBEJS_DB_DUCKDB_S3_ACCESS_KEY_ID=your-access-key-id
CUBEJS_DB_DUCKDB_S3_SECRET_ACCESS_KEY=your-secret-access-key
CUBEJS_DB_DUCKDB_S3_REGION=us-east-2
```
The `aws_s3` sink can also authenticate with the deployment's OIDC identity
instead of an access key pair — the credentials apply to the whole Vector agent.
See [Keyless authentication][ref-keyless-auth].
This replaces the `CUBE_CLOUD_MONITORING_AWS_*` pair above, which is Vector's
**write** path only. The `CUBEJS_DB_DUCKDB_S3_*` variables are separate — they
belong to the DuckDB data source that reads the exported files back in
[Data modeling](#data-modeling), and are still required.
## Data modeling
Example data model for analyzing data from Query History export that is brought to a
bucket in Amazon S3. The data is accessed directly from S3 using DuckDB.
With this data model, you can run queries that aggregate data by dimensions such as
`status`, `environment_name`, `api_type`, etc. and also calculate metrics like
`count`, `total_duration`, or `avg_duration`:
```yaml theme={"dark"}
cubes:
- name: requests
sql: |
SELECT
*,
api_response_duration_ms / 1000 AS api_response_duration,
EPOCH_MS(start_time_unix_ms) AS start_time,
EPOCH_MS(end_time_unix_ms) AS end_time
FROM read_json_auto('s3://cube-query-history-export-demo/**/*.log.gz')
dimensions:
- name: trace_id
sql: trace_id
type: string
primary_key: true
- name: deployment_id
sql: deployment_id
type: number
- name: environment_name
sql: environment_name
type: string
- name: api_type
sql: api_type
type: string
- name: api_query
sql: api_query
type: string
- name: security_context
sql: security_context
type: string
- name: cache_type
sql: cache_type
type: string
- name: start_time
sql: start_time
type: time
- name: end_time
sql: end_time
type: time
- name: duration
sql: api_response_duration
type: number
- name: status
sql: status
type: string
- name: error_message
sql: error_message
type: string
- name: user_name
sql: "SUBSTRING(security_context::JSON ->> 'user', 3, LENGTH(security_context::JSON ->> 'user') - 4)"
type: string
segments:
- name: production_environment
sql: "{environment_name} IS NULL"
- name: errors
sql: "{status} <> 'success'"
measures:
- name: count
type: count
- name: count_non_production
description: |
Counts all non-production environments.
See for details: https://cube.dev/docs/product/administration/deployment/environments
type: count
filters:
- sql: "{environment_name} IS NOT NULL"
- name: total_duration
type: sum
sql: "{duration}"
- name: avg_duration
type: number
sql: "{total_duration} / {count}"
- name: median_duration
type: number
sql: "MEDIAN({duration})"
- name: min_duration
type: min
sql: "{duration}"
- name: max_duration
type: max
sql: "{duration}"
pre_aggregations:
- name: count_and_durations_by_status_and_start_date
measures:
- count
- min_duration
- max_duration
- total_duration
dimensions:
- status
time_dimension: start_time
granularity: hour
refresh_key:
sql: SELECT MAX(end_time) FROM {requests.sql()}
every: 10 minutes
```
## Result
Example query in Playground:
[ref-query-history-export]: /admin/monitoring/monitoring-integrations#query-history-export
[ref-query-history]: /admin/monitoring/query-history
[ref-vector-configuration]: /admin/monitoring/monitoring-integrations#configuration
[ref-keyless-auth]: /admin/monitoring/monitoring-integrations/cloudwatch#keyless-authentication
# Usage Analytics
Source: https://docs.cube.dev/admin/monitoring/usage-analytics
Use built-in Usage Analytics dashboards to understand how your deployments are used — query activity, performance, user adoption, and AI usage.
The **Usage Analytics** page provides a set of pre-built dashboards that show
how your deployments are used: how many queries run and how fast they return,
how effectively the cache serves them, how actively users engage with the
platform, and how much AI activity and token consumption your account generates.
Usage Analytics is itself a Cube application: your account's usage telemetry
is modeled as a curated set of views, and the pre-built dashboards are regular
Cube dashboards on top of them. You can use the same views to
[build your own dashboards](#build-your-own-dashboards).
Available on [Premium and above plans](https://cube.dev/pricing).
Usage Analytics is only available to account administrators.
## Pre-built dashboards
Each dashboard answers a different question about your account's usage.
| Dashboard | Description |
| -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| **Overview** | High-level health and adoption at a glance: requests, error rate, cache hit rate, response time, AI activity, and top tenants. |
| **Query Activity & Performance** | Query volume by API type, latency percentiles, execution-stage breakdown, slowest queries, errors, and activity heatmap. |
| **Users & Adoption** | Active users over time, new vs. returning users, engagement frequency, seat utilization, and dormant accounts. |
| **AI & Token Tracking** | AI messages and conversations by surface, AI adoption, token consumption with estimated cost, and latency. |
## Build your own dashboards
Usage Analytics runs the full Cube application in
[Creator Mode](/embedding/iframe/creator-mode), embedded inside Cube. Beyond
the pre-built dashboards, you get the complete authoring experience: explore
the usage data, create workbooks, and assemble your own dashboards from it.
Your dashboards query the same three curated views that power the pre-built
ones. Anything you build stays private unless you share it.
| View | What it contains |
| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| **API Requests** | Every query hitting your deployments' data APIs, with timing, caching, errors, and the user, tenant, and deployment behind each request. |
| **AI Usage** | AI messages grouped into conversations and sessions, with context, token consumption, estimated cost, and latency. |
| **Users & Adoption** | Users and their engagement: sign-ups, active users, activity frequency, seat utilization, and dormant accounts. |
The API Requests view includes the security context behind each request, so
embedded analytics customers can slice usage by tenant — for example, to find
the heaviest tenants or compare per-tenant latency and cache efficiency.
## Query programmatically
Beyond the embedded dashboards, account administrators can mint a short-lived
Cube API token to query the same usage and billing data directly — from the
[REST API](https://cube.dev/docs/product/apis-integrations/rest-api), a BI
tool, or a script — instead of only viewing it in-app.
```bash theme={"dark"}
curl -X POST https://your-account.cubecloud.dev/api/v1/usage-analytics/token \
-H "Authorization: Bearer "
```
The response includes a `token`, the `apiUrl` to send it to, and an
`expiresAt` timestamp:
```json theme={"dark"}
{
"token": "eyJhbGc...",
"apiUrl": "https://usage-analytics.cubecloud.dev/cubejs-api/v1",
"deploymentId": 331,
"expiresAt": "2026-09-03T13:00:00.000Z"
}
```
The token is scoped to your account, the same way the embedded page is, so
queries only ever see your own data. It expires after **1 hour** — issue a
fresh one for each session or scheduled refresh rather than caching it.
Requires an API key belonging to an account administrator, and Usage
Analytics enabled for the account.
## Data freshness and isolation
Usage Analytics data is available on a slight delay compared to live activity.
For real-time inspection of individual queries, use
[Query History](/admin/monitoring/query-history).
Usage Analytics content is isolated from your own deployments' content: the
views and dashboards described above don't appear in your data model, and your
end users can't access them.
# Google Workspace
Source: https://docs.cube.dev/admin/sso/google-workspace
Walks Google Workspace super admins through SAML single sign-on between Google Admin and Cube Cloud.
Cube Cloud supports authenticating users through Google Workspace, which is
useful when you want your users to access Cube Cloud using single sign on. This
guide will walk you through the steps of configuring SAML authentication in Cube
Cloud with Google Workspace. You **must** be a super administrator in your
Google Workspace to access the Admin Console and create a SAML integration.
Available on [Enterprise plan](https://cube.dev/pricing).
## Enable SAML in Cube Cloud
First, we'll enable SAML authentication in Cube Cloud. To do this, log in to
Cube Cloud and
1. Navigate to **Admin → Settings**.
2. On the **Authentication & SSO** tab, enable **SAML**.
3. For a new integration, replace any prefilled **Audience (SP Entity ID)**
value with the **Single Sign-On URL** shown directly above it.
Take note of the **Single Sign-On URL** and **Audience (SP Entity ID)**
values. You will need them when you configure the SAML integration in Google
Workspace.
## Create a SAML Integration in Google Workspace
Next, we'll create a [SAML app integration for Cube Cloud in Google
Workspace][google-docs-create-saml-app].
1. Log in to [admin.google.com](https://admin.google.com) as an administrator,
then navigate to
**Apps → Web and Mobile Apps** from the left sidebar.
2. Click **Add App**, then click **Add custom SAML app**:
3. Enter a name for your application and click **Next**. You can
optionally add a description and upload a logo for the application, but this
is not required. Click **Continue** to go to the next screen.
4. Take note of the **SSO URL**, **Entity ID** and
**Certificate** values here, as we will need them when we finalize the
SAML integration in Cube Cloud. Click **Continue** to go to the next screen.
5. Enter the following values for the **Service provider details**
section and click **Continue**.
| Name | Description |
| --------- | ---------------------------------------------------------- |
| ACS URL | Use the **Single Sign-On URL** value from Cube Cloud. |
| Entity ID | Use the **Audience (SP Entity ID)** value from Cube Cloud. |
6. On the final screen, click **Finish**.
7. From the app details page, click **User access** and ensure the app is
**ON for everyone**:
## Complete SAML configuration in Cube Cloud
In this step, we'll finalise the configuration by entering the values from our
SAML integration in Google into Cube Cloud.
1. Return to **Admin → Settings → Authentication & SSO → SAML**.
2. Confirm that **Audience (SP Entity ID)** still matches the **Single
Sign-On URL** exactly.
3. Enter the following values in the **SAML Settings** section:
| Name | Description |
| ------------------ | ---------------------------------------------------- |
| Entity ID / Issuer | Use the **Entity ID** value from Google Workspace. |
| SSO (Sign on) URL | Use the **SSO URL** value from Google Workspace. |
| Certificate | Use the **Certificate** value from Google Workspace. |
4. Enable **Auto-provision new users** if you want users to be automatically
created in Cube on their first login via this SAML provider. New users
are assigned the Viewer role by default — see
[Default role for new users](#default-role-for-new-users) to choose a
different role. Enable this if you are not using SCIM provisioning.
5. Click **Apply** to save the changes.
Existing working Google Workspace integrations with a blank Audience do not
need to change immediately. A blank Audience disables audience validation. To
enable validation, update the Google **Entity ID** and Cube **Audience (SP
Entity ID)** together, keep another authentication method enabled, and test
SAML sign-in before you disable the fallback method.
## Default role for new users
By default, users auto-provisioned via SAML receive the **Viewer** role.
To assign a different role, expand the **Advanced** section of the SAML
configuration form and pick from **Default role for new users**:
* **Developer**, **Explorer**, or **Viewer** — Cube Cloud's [default
roles][ref-roles].
* Any [custom role][ref-custom-roles] defined in your account, listed
below the divider.
The selected role applies **only when a user is first created**. Existing
users are not modified on subsequent SSO logins. It is applied **in
addition to** any roles your identity provider sends via the role
attribute (subject to the `rolesMap`).
Admin status is not assignable through this picker — Admin is controlled
separately. To grant admin permissions, update the user's role manually
under [Admin → Users][ref-manage-users].
If the selected role is later renamed or deleted, new users will fall
back to the **Viewer** role until you pick a valid role here. The Viewer
fallback applies whenever the configured default cannot be resolved —
whether that's because no default is set or the configured role no longer
exists.
## Test SAML authentication
To start using SAML authentication, use the
[single sign-on URL provided by Cube Cloud](#enable-saml-in-cube-cloud)
(typically `/sso/saml`) to log in to Cube Cloud.
[google-docs-create-saml-app]: https://support.google.com/a/answer/6087519?hl=en
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
[ref-manage-users]: /admin/users-and-permissions/manage-users
# Overview
Source: https://docs.cube.dev/admin/sso/index
Configure how your team accesses Cube Cloud using passwords, social logins, or SAML-based single sign-on.
Authentication & SSO
As an account administrator, you can manage how your team and users access Cube Cloud.
You can authenticate using email and password, a GitHub account, or a Google account.
Cube Cloud also provides single sign-on (SSO) via identity providers supporting
[SAML](#saml), e.g., Okta, Google Workspace, Azure AD, etc.
[SAML](#saml) is available on [Enterprise plan](https://cube.dev/pricing).
## Configuration
To manage authentication settings, navigate to **Admin → Settings**
of your Cube Cloud account, and switch to the **Authentication & SSO** tab.
Use the toggles in **Password**, **Google**, and **GitHub**
sections to enable or disable these authentication options.
### SAML
Use the toggle in the **SAML** section to enable or disable the authentication
via an identity provider supporting the [SAML protocol][wiki-saml].
Once it's enabled, you'll see the **SAML Settings** section directly below.
Check the following guides to get tool-specific instructions on configuration:
#### Match SAML service provider identifiers
Cube shows two service provider values in the SAML settings:
| Cube setting | Identity provider setting |
| --------------------------- | ------------------------------------ |
| **Single Sign-On URL** | ACS URL or Reply URL |
| **Audience (SP Entity ID)** | Audience, Entity ID, or SP Entity ID |
The **Audience (SP Entity ID)** value validates the audience in SAML responses.
It must exactly match the value configured in your identity provider. Leaving
the field blank disables audience validation and is supported only for
compatibility with existing configurations.
Some identity providers, including Amazon Federate, use one service provider
identifier for both the AuthnRequest issuer and the response audience. For
these providers, set **Audience (SP Entity ID)** in Cube to the **Single
Sign-On URL**, then use that same value as the identity provider's Entity ID
or Audience.
Existing working SAML integrations do not need to change their Audience. When
you change an existing configuration, keep another authentication method
enabled until you have tested SAML sign-in in a separate browser session.
[wiki-saml]: https://en.wikipedia.org/wiki/SAML_2.0
[ref-apis]: /reference
[ref-dap]: /docs/data-modeling/data-access-policies
[ref-security-context]: /docs/data-modeling/access-control/context
[ref-auth-integration]: /docs/data-modeling/access-control#authentication-integration
# SAML authentication with Microsoft Entra ID
Source: https://docs.cube.dev/admin/sso/microsoft-entra-id/saml
Step-by-step guide for configuring Microsoft Entra ID as a SAML identity provider for Cube sign-in.
With SAML (Security Assertion Markup Language) enabled, you can authenticate
users in Cube through Microsoft Entra ID (formerly Azure Active Directory),
allowing your team to access Cube using single sign-on.
Available on [Enterprise plan](https://cube.dev/pricing).
## Prerequisites
Before proceeding, ensure you have the following:
* Admin permissions in Cube.
* Sufficient permissions in Microsoft Entra to create and configure
Enterprise Applications.
## Enable SAML in Cube
First, enable SAML authentication in Cube:
1. In Cube, navigate to **Admin → Settings**.
2. On the **Authentication & SSO** tab, enable the **SAML**
toggle.
3. For a new integration, set **Audience (SP Entity ID)** to the
**Single Sign-On URL** shown directly above it.
4. Take note of the **Single Sign-On URL** and **Audience (SP Entity ID)**
values — you'll need them when configuring the Enterprise Application in
Entra.
Existing working Entra integrations with a different Audience do not need to
change. Keep the existing audience claim override unless you intentionally
update both sides of the integration.
## Create an Enterprise Application in Entra
1. Sign in to the [Microsoft Entra admin center](https://entra.microsoft.com).
2. Go to [Enterprise Applications](https://portal.azure.com/#view/Microsoft_AAD_IAM/StartboardApplicationsMenuBlade/~/AppAppsPreview)
and click **New application**.
3. Select **Create your own application**.
4. Give it a name and choose a **non-gallery application**, then click
**Create**.
## Configure SAML in Entra
1. In your new Enterprise Application, go to the **Single sign-on**
section and select **SAML**.
2. In the **Basic SAML Configuration** section, enter the following:
* **Entity ID** — Use the **Single Sign-On URL** value from Cube.
* **Reply URL** — Use the **Single Sign-On URL** value from Cube.
3. If the **Audience (SP Entity ID)** value from Cube differs from the
**Entity ID** configured in the previous step, go to **Attributes & Claims
→ Edit → Advanced settings** and set the audience claim override to the
Cube value. If the values match, no audience claim override is required.
4. Go to **SAML Certificates → Edit** and select **Sign SAML
response and assertion** for the **Signing Option**.
5. Download the **Federation Metadata XML** file — you'll need it
when completing the Cube configuration.
## Configure attribute mappings
Before returning to Cube, configure the SAML claims Entra sends during
login. Cube uses these claims to identify the user and map optional
attributes such as display name.
Create explicit SAML claims in Entra with the names Cube uses by default.
1. In your Entra Enterprise Application, go to **Single sign-on →
Attributes & Claims**.
2. Add the following claims. Leave **Namespace** blank for each claim:
* **Email** — Set **Name** to `email` and **Source attribute** to
`user.userprincipalname` or `user.mail`.
* **Display name** — Set **Name** to `name` and **Source attribute** to
`user.displayname`.
If you plan to map Cube roles based on Entra group membership (see
[Map roles by group](#map-roles-by-group) below), also add a group claim:
1. Still in **Attributes & Claims**, click **Add a group claim**.
2. Choose which groups to include (e.g. **Security groups** or **Groups
assigned to the application**) and pick a **Source attribute** for the
group name. For most setups, select **sAMAccountName** or
**Cloud-only group display names** so the assertion carries
human-readable group names that match the **IdP group name** values
you'll configure in Cube Cloud.
3. Save the claim.
Cube reads Entra's canonical groups claim URL
(`http://schemas.microsoft.com/ws/2008/06/identity/claims/groups`)
automatically, so no further attribute renaming is required on the
Entra side.
## Complete configuration in Cube
Return to the SAML configuration page in Cube and provide the identity
provider details. You can do this in one of two ways:
**Option A: Upload metadata file**
1. In the **Import IdP Metadata** section, click **Upload
Metadata File**.
2. Select the **Federation Metadata XML** file you downloaded from Entra.
This will automatically populate the **Entity ID / Issuer**,
**SSO (Sign on) URL**, and **Certificate** fields.
**Option B: Enter details manually**
If you prefer to configure the fields manually, enter the following
values from the Entra **Single sign-on** page:
* **Entity ID / Issuer** — Use the **Microsoft Entra Identifier**
value.
* **SSO (Sign on) URL** — Use the **Login URL** value.
* **Certificate** — Paste the Base64-encoded certificate from the
**SAML Certificates** section.
In both options, also configure the following setting:
* **Auto-provision new users** — When enabled, users are automatically
created in Cube on their first login via this SAML provider. Enable this
if you want to provision users only when they first access Cube and you
are not using SCIM provisioning. New users receive the Viewer role by
default; see [Default role for new users](#default-role-for-new-users)
to choose a different role.
## Default role for new users
Auto-provisioned users — both via SAML and via [SCIM][ref-scim] — receive
the **Viewer** role by default. To assign a different role, expand the
**Advanced** section of the SAML configuration form and pick from
**Default role for new users**:
* **Developer**, **Explorer**, or **Viewer** — Cube's [default
roles][ref-roles].
* Any [custom role][ref-custom-roles] defined in your account, listed
below the divider.
The selected role applies **only when a user is first created** during
provisioning. Existing users are not modified on subsequent SSO logins or
SCIM updates.
Admin status is not assignable through this picker — Admin is controlled
separately. To grant admin permissions, update the user's role manually
under [Admin → Users][ref-manage-users].
If the selected role is later renamed or deleted, new users will fall
back to the **Viewer** role until you pick a valid role here. The Viewer
fallback applies whenever the configured default cannot be resolved —
whether that's because no default is set or the configured role no longer
exists.
## Map roles by group
For finer-grained role assignment, enable **Map roles by group** in the
**Advanced Settings** section to assign Cube roles based on a user's
Entra group memberships.
To configure group-based role mapping:
1. Make sure Entra sends a group claim on the SAML assertion. See the
group-claim step in [Configure attribute
mappings](#configure-attribute-mappings).
2. In the SAML configuration form in Cube, expand **Advanced Settings**.
3. (Optional) Under **SAML attribute customization**, set the **Groups
attribute** to the simple name of the SAML attribute carrying group
memberships. Defaults to `groups`. Cube also reads Entra's canonical
groups claim URL automatically, so the default usually works
out of the box.
4. Enable the **Map roles by group** toggle.
5. Click **Add group mapping** and create one entry per group you want
to map:
* **IdP group name** — the group display name as it appears in the
assertion (case-insensitive). With **Source attribute** set to
**sAMAccountName** or **Cloud-only group display names**, this is
the human-readable group name. If you left the default (group object
ID), use the GUID instead.
* **Cube role** — pick a default or [custom role][ref-custom-roles].
For SAML SSO, group mappings are evaluated **only when a new user is
auto-provisioned** on first login. If any matching group resolves to a
Cube role, those roles are assigned to the new user **instead of** the
configured [default role](#default-role-for-new-users). The default role
is used as a fallback when no IdP group matches (or when the mapped Cube
roles no longer exist). Existing users' role assignments are never
modified by subsequent logins.
The same mapping is also applied by [SCIM][ref-scim] when group
memberships are pushed, so a single configuration drives both SAML SSO
and SCIM group sync.
## Assign users
Make sure the new Enterprise Application is assigned to the relevant
users or groups in Entra before testing.
## Test the integration
1. In the Entra **Single sign-on** section, click **Test**
to verify the SAML integration works for your Cube account.
2. Alternatively, copy the **Single Sign-On URL** from Cube,
open it in a new browser tab, and verify you are redirected to
Entra for authentication and then back to Cube.
[ext-ms-entra-id]: https://www.microsoft.com/en-us/security/business/identity-access/microsoft-entra-id
[ref-scim]: /admin/sso/microsoft-entra-id/scim
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
[ref-manage-users]: /admin/users-and-permissions/manage-users
# SCIM provisioning with Microsoft Entra ID
Source: https://docs.cube.dev/admin/sso/microsoft-entra-id/scim
Automates user and group lifecycle in Cube by connecting SCIM provisioning to Microsoft Entra ID.
With SCIM (System for Cross-domain Identity Management) enabled, you can
automate user provisioning in Cube and keep user groups synchronized
with Microsoft Entra ID (formerly Azure Active Directory).
Available on [Enterprise plan](https://cube.dev/pricing).
## Prerequisites
Before proceeding, ensure you have the following:
* Microsoft Entra SAML authentication already configured. If not, complete
the [SAML setup][ref-saml] first.
* Admin permissions in Cube.
* Sufficient permissions in Microsoft Entra to manage Enterprise Applications.
## Enable SCIM provisioning in Cube
Before configuring SCIM in Microsoft Entra, you need to enable SCIM
provisioning in Cube:
1. In Cube, navigate to **Admin → Settings**.
2. In the **SAML** section, enable **SCIM Provisioning**.
## Generate an API key in Cube
To allow Entra ID to communicate with Cube via SCIM, you'll need to
create a dedicated API key:
1. In Cube, navigate to **Settings → API Keys**.
2. Create a new API key. Give it a descriptive name such as **Entra SCIM**.
3. Copy the generated key and store it securely — you'll need it in the
next step.
## Set up provisioning in Microsoft Entra
This section assumes you already have a Cube Enterprise Application
in Microsoft Entra. If you haven't created one yet, follow the
[SAML setup guide][ref-saml] first.
1. Sign in to the [Microsoft Entra admin center](https://entra.microsoft.com).
2. Go to **Applications → Enterprise Applications** and open your
Cube application.
3. Navigate to **Manage → Provisioning**.
4. Set the **Provisioning Mode** to **Automatic**.
5. Under **Admin Credentials**, fill in the following:
* **Tenant URL** — Your Cube deployment URL with `/api/scim/v2`
appended. For example: `https://your-deployment.cubecloud.dev/api/scim/v2`
* **Secret Token** — The API key you generated in the previous step.
6. Click **Test Connection** to verify that Entra ID can reach
Cube. Proceed once the test is successful.
## Configure attribute mappings
Next, configure which user and group attributes are synchronized with
Cube:
1. In the **Mappings** section, select the object type you want to
configure — either users or groups.
2. Remove all default attribute mappings **except** the following:
* **For users**: keep `userName`, `displayName` and `active`.
* **For groups**: keep `displayName` and `members`.
3. Click **Save**.
Users provisioned via SCIM receive the **Viewer** role by default. To
choose a different default role (including [custom roles][ref-custom-roles]),
see [Default role for new users][ref-saml-default-role] on the SAML setup
page — the setting is shared between SAML and SCIM.
Admin permissions cannot be assigned through this setting. To grant admin
permissions, update the user's role manually in Cube under **Admin →
Users**.
## Map roles by SCIM group
If you have configured [Map roles by group][ref-saml-map-roles-by-group]
on the SAML setup page, the same mapping is applied when SCIM provisions
group memberships from Entra — when Entra adds a user to a synchronized
group, Cube assigns the mapped role to that user. Role assignment is
**additive**: removing a user from an Entra group does not strip the
corresponding role; adjust the user's roles in Cube manually.
No separate configuration is required on the SCIM side — once the
mapping is defined on the SAML page, it drives both SAML SSO and SCIM
group sync. Match is on the group **display name** Entra pushes (the
`displayName` attribute in the **Group** mapping, case-insensitive).
## Syncing user attributes
You can sync [user attributes][ref-user-attributes] from Microsoft Entra to
Cube via SCIM, allowing you to centralize user management in Entra.
### Create a user attribute in Cube
In Cube, navigate to **Admin → Settings → User Attributes** and
create a new attribute. Take note of the attribute **reference** name — you will
need it when configuring Entra.
### Create an Entra user attribute
1. In the [Microsoft Entra admin center](https://entra.microsoft.com), navigate
to **Applications → Enterprise Applications** and open your Cube
application.
2. Go to **Manage → Provisioning → Mappings**.
3. Select the user mapping you want to add the attribute to.
4. At the bottom of the page, select **Show advanced options**.
5. Select **Edit attribute list for customappsso**.
6. Add a new attribute with the following settings:
* **Name** — The reference of the attribute you created in Cube,
prefixed with `urn:cube:params:1.0:UserAttribute:`.
For example, for an attribute with the reference `country`, enter
`urn:cube:params:1.0:UserAttribute:country`.
* **Type** — Select the matching type (`string` or `integer`).
7. Save the changes.
### Create attribute mapping
1. After saving, click **Yes** when prompted.
2. In the **Attribute Mapping** page, click **Add New Mapping**.
3. In the **Target attribute** dropdown, select the attribute you created
in the previous step.
4. Configure the source mapping to the appropriate Entra field.
5. Click **OK**, then **Save**.
The next time the Entra application syncs, the attribute values will be
provisioned as [user attributes][ref-user-attributes] in Cube.
[ref-saml]: /admin/sso/microsoft-entra-id/saml
[ref-saml-default-role]: /admin/sso/microsoft-entra-id/saml#default-role-for-new-users
[ref-saml-map-roles-by-group]: /admin/sso/microsoft-entra-id/saml#map-roles-by-group
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
[ref-user-attributes]: /admin/users-and-permissions/user-attributes
# SAML authentication with Okta
Source: https://docs.cube.dev/admin/sso/okta/saml
Connects Okta to Cube Cloud with SAML so members authenticate through your existing Okta application.
With SAML (Security Assertion Markup Language) enabled, you can authenticate
users in Cube Cloud through Okta, allowing your team to access Cube Cloud
using single sign-on.
Available on [Enterprise plan](https://cube.dev/pricing).
## Prerequisites
Before proceeding, ensure you have the following:
* Admin permissions in Cube Cloud.
* Account administrator permissions in your Okta organization to access
the Admin Console and create SAML integrations.
## Enable SAML in Cube Cloud
First, enable SAML authentication in Cube Cloud:
1. In Cube Cloud, navigate to **Admin → Settings**.
2. On the **Authentication & SSO** tab, enable the **SAML**
toggle.
3. For a new integration, set **Audience (SP Entity ID)** to the
**Single Sign-On URL** shown directly above it. This gives the AuthnRequest
issuer and response audience one matching service provider identifier.
4. Take note of the **Single Sign-On URL** and **Audience (SP Entity ID)**
values — you'll need them when configuring the SAML integration in Okta.
Existing working Okta integrations with a different Audience do not need to
change. If you update the value, change it in Cube and Okta together and test
SAML sign-in before disabling another authentication method.
## Create a SAML integration in Okta
1. Log in to your Okta organization as an administrator, then navigate to
the Admin Console by clicking **Admin** in the top-right corner.
2. Click **Applications → Applications** from the navigation on the
left, then click **Create App Integration**.
3. Select **SAML 2.0** and click **Next**.
4. Enter a name for your application and click **Next**.
5. Enter the following values in the **SAML Settings** section:
* **Single sign on URL** — Use the **Single Sign-On URL**
value from Cube Cloud.
* **Audience URI (SP Entity ID)** — Use the **Audience (SP Entity ID)**
value from Cube Cloud. It must match exactly.
6. Click **Next** to go to the **Feedback** screen, fill in
any necessary details and click **Finish**.
## Configure attribute statements in Okta
After the application is created, configure attribute statements to map
user attributes from Okta to Cube Cloud:
1. In your SAML app integration, go to the **Sign On** tab.
2. Scroll down to the **Attribute statements** section.
3. Click **Add expression** and create the following entries:
| Name | Expression |
| ------- | ------------------------ |
| `email` | `user.profile.email` |
| `name` | `user.profile.firstName` |
4. If you plan to map Cube roles based on Okta group membership (see
[Map roles by group](#map-roles-by-group) below), also add a **Group
Attribute Statement**. Scroll to the **Group Attribute Statements**
section and add:
| Name | Filter |
| -------- | ---------------------- |
| `groups` | **Matches regex** `.*` |
Adjust the filter to scope which groups Okta sends — e.g.
**Starts with** `cube-` to limit the assertion to Cube-related groups.
The attribute name must match the **Groups attribute** value configured
in Cube Cloud (defaults to `groups`).
## Retrieve SAML details from Okta
Next, retrieve the values you'll need to complete the configuration
in Cube Cloud:
1. In your SAML app integration, go to the **Sign On** tab.
2. In the sidebar, click **View SAML setup instructions**.
3. Take note of the following values from the setup instructions page:
* **Identity Provider Single Sign-On URL**
* **Identity Provider Issuer**
* **X.509 Certificate**
## Complete configuration in Cube Cloud
Return to the SAML configuration page in Cube Cloud and provide the
identity provider details:
* **Entity ID / Issuer** — Use the **Identity Provider Issuer**
value from Okta.
* **SSO (Sign on) URL** — Use the **Identity Provider Single
Sign-On URL** value from Okta.
* **Certificate** — Paste the **X.509 Certificate** from Okta.
* **Auto-provision new users** — When enabled, users are automatically
created in Cube on their first login via this SAML provider. Enable this
if you want to provision users only when they first access Cube and you
are not using SCIM provisioning. New users receive the Viewer role by
default; see [Default role for new users](#default-role-for-new-users)
to choose a different role.
## Default role for new users
Auto-provisioned users — both via SAML and via [SCIM][ref-scim] — receive
the **Viewer** role by default. To assign a different role, expand the
**Advanced** section of the SAML configuration form and pick from
**Default role for new users**:
* **Developer**, **Explorer**, or **Viewer** — Cube's [default
roles][ref-roles].
* Any [custom role][ref-custom-roles] defined in your account, listed
below the divider.
The selected role applies **only when a user is first created** during
provisioning. Existing users are not modified on subsequent SSO logins or
SCIM updates.
Admin status is not assignable through this picker — Admin is controlled
separately. To grant admin permissions, update the user's role manually
under [Admin → Users][ref-manage-users].
If the selected role is later renamed or deleted, new users will fall
back to the **Viewer** role until you pick a valid role here. The Viewer
fallback applies whenever the configured default cannot be resolved —
whether that's because no default is set or the configured role no longer
exists.
## Map roles by group
For finer-grained role assignment, enable **Map roles by group** in the
**Advanced Settings** section to assign Cube Cloud roles based on a user's
Okta group memberships.
To configure group-based role mapping:
1. Make sure Okta sends a group attribute statement on the SAML assertion.
See step 4 of [Configure attribute statements in
Okta](#configure-attribute-statements-in-okta).
2. In the SAML configuration form in Cube Cloud, expand **Advanced
Settings**.
3. (Optional) Under **SAML attribute customization**, set the **Groups
attribute** to the name of the SAML attribute you configured in Okta.
Defaults to `groups`.
4. Enable the **Map roles by group** toggle.
5. Click **Add group mapping** and create one entry per group you want to
map:
* **IdP group name** — the Okta group **display name** exactly as it
appears in the SAML assertion (case-insensitive).
* **Cube Cloud role** — pick a default or [custom role][ref-custom-roles].
For SAML SSO, group mappings are evaluated **only when a new user is
auto-provisioned** on first login. If any matching group resolves to a
Cube Cloud role, those roles are assigned to the new user **instead of**
the configured [default role](#default-role-for-new-users). The default
role is used as a fallback when no IdP group matches (or when the mapped
Cube Cloud roles no longer exist). Existing users' role assignments are
never modified by subsequent logins.
The same mapping is also applied by [SCIM][ref-scim] when group
memberships are pushed, so a single configuration drives both SAML SSO
and SCIM group sync.
## Test SAML authentication
1. Copy the **Single Sign-On URL** from the SAML configuration page
in Cube Cloud.
2. Open a new browser tab and paste the URL into the address bar, then
press **Enter**.
3. You should be redirected to Okta to log in. After a successful login,
you should be redirected back to Cube Cloud.
[okta-docs-create-saml-app]: https://help.okta.com/en-us/Content/Topics/Apps/Apps_App_Integration_Wizard_SAML.htm
[ref-scim]: /admin/sso/okta/scim
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
[ref-manage-users]: /admin/users-and-permissions/manage-users
# SCIM provisioning with Okta
Source: https://docs.cube.dev/admin/sso/okta/scim
With SCIM (System for Cross-domain Identity Management) enabled, you can automate user provisioning in Cube Cloud and keep user groups synchronized with Okta.
Available on [Enterprise plan](https://cube.dev/pricing).
## Prerequisites
Before proceeding, ensure you have the following:
* Okta SAML authentication already configured. If not, complete
the [SAML setup][ref-saml] first.
* Admin permissions in Cube Cloud.
* Admin permissions in Okta to manage application integrations.
## Enable SCIM provisioning in Cube Cloud
Before configuring SCIM in Okta, you need to enable SCIM
provisioning in Cube Cloud:
1. In Cube, navigate to **Admin → Settings**.
2. In the **SAML** section, enable **SCIM Provisioning**.
## Generate an API key in Cube Cloud
To allow Okta to communicate with Cube Cloud via SCIM, you'll need to
create a dedicated API key:
1. In Cube Cloud, navigate to **Settings → API Keys**.
2. Create a new API key. Give it a descriptive name such as **Okta SCIM**.
3. Copy the generated key and store it securely — you'll need it in the
next step.
## Enable SCIM provisioning in Okta
This section assumes you already have a Cube Cloud SAML app integration
in Okta. If you haven't created one yet, follow the
[SAML setup guide][ref-saml] first.
1. In the Okta Admin Console, go to **Applications → Applications**
and open your Cube Cloud application.
2. On the **General** tab, click **Edit** in the
**App Settings** section.
3. Set the **Provisioning** field to **SCIM** and click
**Save**.
## Configure SCIM connection in Okta
1. Navigate to the **Provisioning** tab of your Cube Cloud application.
2. In the **Settings → Integration** section, click **Edit**.
3. Fill in the following fields:
* **SCIM connector base URL** — Your Cube Cloud deployment URL with
`/api/scim/v2` appended. For example:
`https://your-deployment.cubecloud.dev/api/scim/v2`
* **Unique identifier field for users** — `userName`
* **Supported provisioning actions** — Select **Push New Users**,
**Push Profile Updates**, and **Push Groups**.
* **Authentication Mode** — Select **HTTP Header**.
4. In the **HTTP Header** section, paste the API key you generated
earlier into the **Authorization** field.
5. Click **Test Connector Configuration** to verify that Okta can
reach Cube Cloud. Proceed once the test is successful.
6. Click **Save**.
## Configure provisioning actions
After saving the SCIM connection, configure which provisioning actions
are enabled for your application:
1. On the **Provisioning** tab, go to **Settings → To App**.
2. Click **Edit** and enable the actions you want:
* **Create Users** — Automatically create users in Cube Cloud when
they are assigned in Okta.
* **Update User Attributes** — Synchronize profile changes from Okta
to Cube Cloud.
* **Deactivate Users** — Deactivate users in Cube Cloud when they are
unassigned or deactivated in Okta.
3. Click **Save**.
## Assign users and groups
For users and groups to be provisioned in Cube Cloud, you need to assign
them to your Cube Cloud application in Okta. This is also required for
group memberships to be correctly synchronized — pushing a group alone
does not assign its members to the application.
1. In your Cube Cloud application, navigate to the **Assignments** tab.
2. Click **Assign** and choose **Assign to Groups** (or
**Assign to People** for individual users).
3. Select the groups or users you want to provision and click **Assign**,
then click **Done**.
If users were assigned to the application before SCIM provisioning was enabled,
Okta will show the following message in the **Assignments** tab:
*"User was assigned this application before Provisioning was enabled and not
provisioned in the downstream application. Click Provision User."*
To resolve this, click **Provision User** next to each affected user.
This will trigger SCIM provisioning for them without needing to remove and
re-add their assignment.
## Push groups to Cube Cloud
To synchronize groups from Okta to Cube Cloud, you need to select which
groups to push:
1. In your Cube Cloud application, navigate to the **Push Groups** tab.
2. Click **Push Groups** and choose how to find your groups — you can
search by name or rule.
3. Select the groups you want to push to Cube Cloud and click **Save**.
## Default role for SCIM-provisioned users
Users created through SCIM receive the same default role configured for
SAML auto-provisioning — there is no separate SCIM control. By default
this is the **Viewer** role; to choose a different default role (including
[custom roles][ref-custom-roles]), see
[Default role for new users][ref-saml-default-role] on the SAML setup page.
The default role only applies to users created by SCIM (`POST
/api/scim/v2/Users`). Existing users are not modified by subsequent SCIM
profile updates.
## Map roles by SCIM group
If you have configured [Map roles by group][ref-saml-map-roles-by-group]
on the SAML setup page, the same mapping is applied when SCIM pushes
group memberships from Okta — when Okta adds a user to a pushed group,
Cube Cloud assigns the mapped role to that user. Role assignment is
**additive**: removing a user from an Okta group does not strip the
corresponding role; adjust the user's roles in Cube Cloud manually.
No separate configuration is required on the SCIM side — once the
mapping is defined on the SAML page, it drives both SAML SSO and SCIM
group sync. Match is on the group **display name** Okta pushes
(case-insensitive).
[ref-saml]: /admin/sso/okta/saml
[ref-saml-default-role]: /admin/sso/okta/saml#default-role-for-new-users
[ref-saml-map-roles-by-group]: /admin/sso/okta/saml#map-roles-by-group
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
# Time zones
Source: https://docs.cube.dev/admin/time-zones
Run queries in the time zone your readers actually work in — account-wide, per user, and per dashboard.
By default, every query Cube runs buckets time in the deployment's
[default time zone](/docs/data-modeling/configuration#default-time-zone) — the
[`CUBEJS_DEFAULT_TIMEZONE`](/reference/configuration/environment-variables#cubejs_default_timezone)
environment variable, `UTC` unless you change it. That means "orders today" answers the
same question for everyone, regardless of where they sit — which is wrong by up to a day
for anyone outside that zone.
Turning on **user time zones** lets a zone be resolved per account, per user, and per
dashboard instead.
This feature is **off by default**, and turning it on **moves numbers**. While it is
off, nothing changes for anyone. Once it is on, a reader whose effective zone differs
from that default sees different daily, weekly, and monthly totals — because the days
are cut in a different place.
## What a time zone changes
The effective zone is applied to every query Cube runs on your behalf:
* **Time dimension bucketing** — which rows fall into which day, week, month, or quarter.
* **Relative dates** — `today`, `yesterday`, `this week`, `last 7 days`, and the dates
the agent resolves when you ask about "today".
* **Date range filters** — the boundaries you type are interpreted in the effective zone.
It applies to charts, dashboards, drill-downs, subtotals and totals, sparklines, period
comparison, [Analytics Chat](/docs/explore-analyze/analytics-chat), and embedded
surfaces alike, so a dashboard's charts and its agent panel always agree.
A time zone is a **display and bucketing** concern only. It never affects what data a
user can see — access control still comes from roles and the security context.
## Turn it on
Go to **Admin → Settings → Time Zones**. Three controls, in the order the decisions are
made:
| Control | What it does |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Enable user time zones** | The master switch. Off by default; while off, no surface resolves a zone at all. |
| **Tenant time zone** | The account-wide zone: everyone gets it unless they override it. Leave it as **Deployment default** to keep using each deployment's own default. |
| **Allow personal time zones** | Whether users may choose their own zone on their Preferences page. On by default once the feature is enabled. |
The last two appear only while the feature is enabled, and they apply to every user in
the account. The UI labels the middle control **Tenant time zone**; this page calls the
zone it sets the *account-wide zone*, matching how the docs scope things.
## Personal time zone
When **Allow personal time zones** is on, each user can pick their own zone under
**Preferences → Time zone** (see [Preferences](/docs/preferences#time-zone)). Only a zone
the user has explicitly chosen is ever applied — Cube never silently uses the browser's
zone, though it will offer the detected zone as a suggestion.
Turning **Allow personal time zones** off makes everyone query in the account-wide zone again,
and existing personal choices stop applying.
## Dashboard time zone
A dashboard is one artifact many people read, so its zone is a property of the dashboard
rather than of whoever opens it. Set it in the dashboard builder under
**Options → Time zone**, which offers three choices:
| Choice | Behavior |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| **Deployment default** | Inherit — follow the account-wide zone, or the deployment's own default when no account-wide zone is set. |
| **Viewer time zone** | Resolve per reader, so each viewer sees their own local day. Use this for an operational board. |
| A named zone | Pin the dashboard — "this dashboard reports in `America/New_York`", and keeps doing so after an admin changes the account-wide zone. |
The zone is stored with the **published** version, so editing a draft does not move the
numbers on the dashboard people are currently reading. Publish to apply it.
**Viewer time zone** is offered only while **Allow personal time zones** is on — without
it, a per-reader promise is one Cube would not keep.
### Reading a dashboard in another zone
A published dashboard shows the zone its numbers are bucketed in, next to its title,
along with where that zone came from — **Set by this dashboard**, **Your own time zone**,
or **Deployment default**.
Where the dashboard leaves the choice open, that control is also a dropdown: pick another
zone to look at the same dashboard in it. This is a temporary lens, not an edit — nothing
is saved, nobody else is affected, and leaving the dashboard drops it.
The zone is shown but **not changeable** when the dashboard is pinned to a named zone, or
when the account does not allow personal time zones. In both cases the zone is not the
reader's to reinterpret.
## Exploration time zone
A saved exploration carries a zone the same way, chosen from the **Time zone** control in
the Explore header. The rows mean what they mean on a dashboard: inherit, resolve per
viewer, or pin a named zone. It saves as soon as you pick it, so anyone who opens the
exploration afterwards gets that zone; readers with view-only access see the zone but
cannot change it.
## How the zone is resolved
Highest priority first.
**A dashboard or a saved exploration:**
1. A reader's temporary lens, or an embed host's `?timezone=` (see
[embedded time zones](/embedding/iframe/time-zones)).
2. The artifact's own pinned zone.
3. The reader's personal zone — only when the artifact is set to **Viewer time zone**,
and only when the account allows personal zones.
4. What the artifact inherits: the account-wide embed zone for an embed, otherwise the
account-wide zone.
5. The deployment's default time zone.
**Ad-hoc surfaces** — a new exploration, a standalone chat — resolve the reader's own
personal zone first, then the account-wide zone. Here the only reader is the person asking, so
their own zone is the right answer.
**Embedded surfaces** follow their own chain, documented in
[embedded time zones](/embedding/iframe/time-zones).
At every level, if nothing resolves, Cube sends no zone and the deployment applies its
default time zone — exactly as it did before this feature existed.
## Valid time zone values
Cube accepts [IANA time zone names][link-tzdb] such as `America/New_York` or
`Asia/Tokyo`. Bare UTC offsets like `+05:30` are **rejected** rather than accepted,
because Cube would compute them in UTC while reporting the offset back — silently wrong.
Legacy aliases are understood (`US/Eastern` resolves to `America/New_York`).
[link-tzdb]: https://en.wikipedia.org/wiki/List_of_tz_database_time_zones
# Custom roles
Source: https://docs.cube.dev/admin/users-and-permissions/custom-roles
Define fine-grained custom roles for your organization on the Enterprise plan.
Custom roles are available on the [Enterprise plan](https://cube.dev/pricing).
Cube comes with [default roles][ref-default-roles] (Admin, Developer, Explorer, Viewer) that cover common use cases. When you need finer control — for example, to let a user edit the data model on a single deployment but nothing else — you can define **custom roles** with a tailored set of permissions.
[ref-default-roles]: /admin/users-and-permissions/roles-and-permissions
## How custom roles work
Each custom role is built around three concepts:
1. **Base Role** (required) — Viewer, Explorer, or Developer. Determines the user's license tier and the inherited level of access.
2. **Global permissions** — org-wide capabilities such as billing, customization, integrations, and account-wide deployment management.
3. **Deployment permissions** — one or more policies, each targeting "All deployments" or specific deployments.
Permissions stack: a user gets the union of every role assigned to them, so a user with multiple custom roles holds the broadest set of granted permissions.
## Browsing roles
To see the list of custom roles, go to **Admin → Custom Roles** in your Cube account. Click on a role to edit it, or click **Add Role** to create a new one.
## Anatomy of a custom role
The role builder uses a two-column layout. The left column (sticky as you scroll) holds the role's identity and the **Create / Update** button. The right column (scrollable) holds the permission cards.
### Name and description
* **Name** is required and must be unique within the account. The names `Admin`, `Guest`, `Developer`, `None`, and `All` are reserved.
* **Description** is optional but recommended — it shows up alongside the role on the Custom Roles list and on user profile pages.
### Base Role
Every custom role has exactly one **Base Role**. It determines:
* The user's **license tier** — Cube infers the tier from the highest Base Role across all of the user's roles.
* The **inherited** capabilities the role grants out of the box.
| Base Role | Description |
| --------- | ----------------------------------------------- |
| Viewer | Read-only access to dashboards and chats. |
| Explorer | Viewer + create/edit workbooks, run queries. |
| Developer | Explorer + edit data model, manage deployments. |
Inheritance is hierarchical:
* **Explorer** automatically grants Viewer access.
* **Developer** automatically grants Explorer and Viewer access.
The Base Role is required. Save is disabled until one is selected.
#### Auto-bump to Developer
If you check any deployment-scoped action stronger than **Access deployment** (`DeploymentRead`) — for example, **Edit deployment** or **Edit data model** — the Base Role is automatically forced to **Developer**, and the Viewer and Explorer options are disabled with a tooltip:
> Selected actions require Developer role
Removing the elevated action re-enables the lower tiers. This mirrors the server-side rule that any deployment write or data-model write requires the Developer license tier.
### Global permissions
Global permissions are org-wide. Check any number of them on the **Global permissions** card:
| Group | Permission | What it grants |
| ------------------------------------ | --------------------------------- | ----------------------------------------------- |
| Audit & Billing | Manage audit log | Access and configure the audit log. |
| Audit & Billing | View billing | View invoices and usage. |
| Customization | Manage chart palettes | Create and edit org-wide chart palettes. |
| Customization | Manage dashboard themes | Create and edit org-wide dashboard themes. |
| Integrations / Network | Manage OAuth integrations | Configure OAuth integrations and clients. |
| Integrations / Network | Issue OAuth tokens | Mint OAuth tokens for integrations. |
| Deployment management (account-wide) | Manage deployments (account-wide) | Create and list deployments across the account. |
### Deployment permissions
Deployment permissions are scoped to specific deployments. A custom role can have any number of **deployment policies**, each represented as its own card. Click **Add deployment policy** to add another card.
Each card has two sections:
#### Scope
Choose which deployments the policy applies to:
* **All deployments** — the policy covers every deployment in the account, including ones added later.
* **Specific deployments** — pick deployments from a searchable multi-select picker.
#### Actions
Either grant **Full access** (a shortcut that enables every current and future deployment-scoped permission) or pick granular actions:
| Group | Permission | What it grants |
| ----------------------- | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| General | Access deployment | Foundation permission required to access the deployment at all. A Viewer needs this just to open it. |
| General | Edit deployment | Modify general deployment settings (name, Build & Deploy, Configuration flags, etc.). Does **not** include environment variables, data sources, or other configuration secrets — see **Manage secrets**. |
| General | Delete deployment | Remove the deployment. |
| Configuration & Secrets | Manage secrets | View and edit environment variables, manage data sources (add/edit database connections, including the test-connection step), the Model Configuration card on **Settings → Configuration**, Power BI (XMLA) settings, and restarting dev mode to apply saved environment variable changes. |
| Data Model | View data model | Read data model files and dev branches. |
| Data Model | Edit data model | Edit the data model on any branch, including the main/deploy branch. Allows committing and merging to main, force-syncing main, and starting dev mode against main. |
| Data Model | Edit data model on dev branches | Edit the data model only on non-default branches. Blocks any write that targets the main/deploy branch (including merging to main). |
| Monitoring | Access query history | View deployment query history, performance, and traces. |
| Data Export | Download data | Download query results as CSV from workbooks, Analytics Chat, and published dashboards. Granted by default to the built-in Viewer, Explorer, and Developer roles. See [Data download controls][ref-data-download-controls]. |
When **Full access** is checked, the granular checkboxes appear checked and disabled — granting Full access today also covers any deployment-scoped permissions added in the future.
**Access deployment** (`DeploymentRead`) and **Download data** (`DownloadData`) are special. They are the only deployment actions that do **not** auto-bump the Base Role to Developer, because Viewers also need them just to open a deployment and download data from it.
[ref-data-download-controls]: /admin/users-and-permissions/roles-and-permissions#restricting-data-downloads
## Walkthroughs
### Create a Viewer with read access to all deployments
A common pattern: someone who consumes dashboards and chats but never edits anything.
Go to **Admin → Custom Roles** and click **Add Role**.
Set the **Name** to, e.g., `Org Viewer`.
Select **Viewer**.
Click **Add deployment policy**, leave **Scope** on **All deployments**, and check **Access deployment** under General.
Click **Create**.
### Create an Explorer for a single deployment
For an analyst who works in one deployment.
Name the role (e.g., `Marketing Analyst`) and pick **Explorer** as the Base Role.
Add a deployment policy, switch **Scope** to **Specific deployments**, and pick the relevant deployment(s) from the picker.
Check **Access deployment** so the user can open the deployment, and add **Access query history** if they should see query performance.
Click **Create**.
### Create a Developer with Full access to specific deployments
For a data engineer who owns a subset of deployments end-to-end.
Name the role (e.g., `Sales Domain Owner`).
Add a deployment policy, switch **Scope** to **Specific deployments**, pick the relevant deployments, and check **Full access**.
The Base Role automatically switches to **Developer** and the Viewer/Explorer radios are disabled.
Click **Create**.
## Assigning roles to users
To assign a custom role to a user:
1. Navigate to **Admin → Users**.
2. Either change the role from the dropdown in the users table, or click into the user's profile page.
3. Add one or more custom roles. Multiple roles stack — the user holds the union of permissions.
See [Manage users][ref-manage-users] for more.
[ref-manage-users]: /admin/users-and-permissions/manage-users
## Validation
The role builder enforces a few rules client-side:
* **Name** is required and must be unique. Reserved names (`Admin`, `Guest`, `Developer`, `None`, `All`) are rejected.
* **Base Role** is required — Save is disabled until one is picked, and an inline error appears if you try to submit without one.
* A deployment policy with **no actions selected** is silently dropped on save (treated as a no-op).
* A deployment policy with **Specific deployments** scope but no deployments selected is also dropped on save.
## Reference: action catalog
This section lists every action a custom role can grant. The internal action names match the names shown in the [audit log][ref-audit-log].
[ref-audit-log]: /admin/monitoring/audit-log
### Base Role actions (Global)
| Internal name | Label | Description |
| ------------- | --------- | ----------------------------------------------- |
| `AIBIView` | Viewer | Read-only access to dashboards and chats. |
| `AIBIExplore` | Explorer | Viewer + create/edit workbooks, run queries. |
| `AIBIDevelop` | Developer | Explorer + edit data model, manage deployments. |
### Global actions
| Internal name | Label |
| ------------------------------ | --------------------------------- |
| `AuditLogManage` | Manage audit log |
| `BillingRead` | View billing |
| `ChartPalettesManage` | Manage chart palettes |
| `DashboardThemesManage` | Manage dashboard themes |
| `OAuthIntegrationsManage` | Manage OAuth integrations |
| `OAuthIntegrationsIssueTokens` | Issue OAuth tokens |
| `DeploymentsManage` | Manage deployments (account-wide) |
### Deployment-scoped actions
| Internal name | Label |
| ------------------------- | -------------------------------------------------------- |
| `All` | Full access (every current and future deployment action) |
| `DeploymentRead` | Access deployment |
| `DeploymentUpdate` | Edit deployment |
| `DeploymentDelete` | Delete deployment |
| `SecretsManage` | Manage secrets |
| `SchemaRead` | View data model |
| `SchemaUpdate` | Edit data model |
| `SchemaUpdateDevBranches` | Edit data model on dev branches |
| `APMRead` | Access query history |
| `DownloadData` | Download data |
# Manage users
Source: https://docs.cube.dev/admin/users-and-permissions/manage-users
Add users to your Cube account, assign roles, and control their access.
Use the **Admin → Users** page to add people to your Cube account,
change their roles, and control what they can access.
Only users with the Admin role can manage other users.
## The user list
The user list displays all users in your account along with their roles
and status. Use the search bar to filter users by name or email.
## Inviting users
To invite a new user:
1. Navigate to **Admin → Users**.
2. Click **Add User**.
3. Enter the user's email address.
4. Select a [role][ref-roles] for the user: Admin, Developer, Explorer, or
Viewer.
5. Optionally, assign one or more [custom roles][ref-custom-roles].
6. Click **Create** to send the invitation.
After the user is created, an invitation link is generated. You can copy the
link and share it with the user. The user must visit the link to set their
password and activate their account.
### Resending invitations
If a user hasn't activated their account, you can resend or copy the
invitation link from the user list.
## Managing individual users
Click on a user in the user list to access their settings page. From here,
you can:
* Update the user's name
* Change their [role][ref-roles] (Admin, Developer, Explorer, or Viewer)
* Assign or remove [custom roles][ref-custom-roles]
* Add the user to [user groups][ref-groups]
* Set [user attribute][ref-attributes] values for data access control
### Changing a user's role
To change a user's role:
1. Navigate to **Admin → Users** and click on the user.
2. Select a new role from the role dropdown.
3. Save the changes.
Alternatively, you can change a user's role directly from the user list
using the role dropdown.
Admin roles are billed at the developer rate.
## Deactivating and reactivating users
Deactivating a user revokes their access to Cube without permanently
removing their account. The user's active sessions are terminated immediately.
To deactivate a user:
1. Navigate to **Admin → Users**.
2. Click the actions menu for the user.
3. Select **Deactivate**.
You cannot deactivate yourself or the last active Admin user.
To reactivate a deactivated user, follow the same steps and select
**Activate**. The user can then log in again with their existing credentials.
## Deleting users
Deleting a user permanently removes their account from Cube.
To delete a user:
1. Navigate to **Admin → Users**.
2. Click the actions menu for the user.
3. Select **Delete**.
You cannot delete your own account. Deleting a user is irreversible.
## Provisioning users via SCIM
If your organization uses an identity provider such as [Okta][ref-okta] or
[Microsoft Entra ID][ref-entra-id], you can automate user provisioning and
deprovisioning through SCIM. See the [SSO & Identity Providers][ref-sso]
documentation for setup instructions.
Users created via SCIM — and users auto-provisioned on first SAML
login — receive the **Viewer** role by default. To assign a different
default role (including [custom roles][ref-custom-roles]), configure
**Default role for new users** in the **Advanced** section of your SAML
configuration. The setting is shared between SAML and SCIM.
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
[ref-attributes]: /admin/users-and-permissions/user-attributes
[ref-groups]: /admin/users-and-permissions/user-groups
[ref-sso]: /admin/sso
[ref-okta]: /admin/sso/okta/scim
[ref-entra-id]: /admin/sso/microsoft-entra-id/scim
# Roles & Permissions
Source: https://docs.cube.dev/admin/users-and-permissions/roles-and-permissions
_Understanding user roles and permissions in Cube._
Cube has four built-in default roles: Admin, Developer, Explorer, and Viewer. Each role has specific permissions and access levels designed to support different responsibilities within the platform.
The Enterprise tier allows creation of custom roles with a customized set of permissions tailored to your organization's specific needs.
## Permissions matrix
| Permission | Admin | Developer | Explorer | Viewer |
| --------------------------------------------- | :---: | :-------: | :------: | :----: |
| Manage users and account settings | ✅ | ❌ | ❌ | ❌ |
| Manage deployments and their settings | ✅ | ✅ | ❌ | ❌ |
| Edit semantic model | ✅ | ✅ | ❌ | ❌ |
| Execute SQL queries against data sources | ✅ | ✅ | ❌ | ❌ |
| Explore and query semantic models | ✅ | ✅ | ✅ | ❌ |
| Create and edit workbooks | ✅ | ✅ | ✅ | ❌ |
| View published dashboards | ✅ | ✅ | ✅ | ✅ |
| Export data from dashboards (CSV, PDF, PNG) | ✅ | ✅ | ✅ | ✅ |
| Access Analytics Chat | ✅ | ✅ | ✅ | ✅ |
| Query from external tools (Tableau, Power BI) | ✅ | ✅ | ✅ | ✅ |
Admin roles are billed at the developer rate.
## Restricting data downloads
By default, all roles can download query results as CSV from workbooks, Analytics Chat, and published dashboards. Two controls restrict this:
* **Account-wide** — the **Allow data downloads** switch on the **Admin → Settings** page. Turning it off hides export controls everywhere, for all users — including admins and embeds.
* **Per user** — the **Download data** action in [custom roles][ref-custom-roles]. Built-in roles grant it by default; to remove download access for a group of users, assign them a custom role without it. Admins and anonymous embed viewers bypass it.
The account-wide switch always wins; the permission applies only while downloads are allowed account-wide. Both govern the download button, not data access — query results are still rendered on screen.
[ref-custom-roles]: /admin/users-and-permissions/custom-roles
## Agent Permissions
Agents are connected to Semantic Model Deployments and inherit the permission level of the user they are operating under.
Each agent can be configured to use the *Restrict Views* feature which allows to select the views that are visible to the agent.
This feature should not be used as a security measure. Configure [access policies][ref-data-access-policies] instead.
## Typical Usage Scenarios
* Viewers: Business users who consume published dashboards, use Analytics Chat, and query data through external tools like Tableau or Power BI
* Explorers: Typically data consumers and analysts
* Developers: Usually data stewards and data engineers
* Admins: Typically assigned to data engineers managing the entire Cube instance (billed at the developer rate with additional privileges)
[ref-data-access-policies]: /docs/data-modeling/data-access-policies
# User Attributes
Source: https://docs.cube.dev/admin/users-and-permissions/user-attributes
_Secure data access with user attributes for filtering based on individual permissions._
User attributes allow you to implement row-level security by filtering data based on user-specific values. This documentation explains how to set up and use user attributes for access control.
## Creating User Attributes
1. Go to **Admin → Attributes**
2. Click to create a new attribute
3. Configure the attribute:
* Set a name
* Choose the type
* Optionally set a default value
* Optionally set a display name
## Setting User Attribute Values
User attributes can be set on a per-user basis:
1. Go to the user's page
2. Locate the attributes section
3. Set the desired attribute value (e.g., setting city to "Los Angeles")
## Implementing Row-Level Access Policy
To filter data based on user attributes, implement an access policy in your views:
```yaml theme={"dark"}
views:
- name: orders_view
access_policy:
- group: "*" # Applies to all groups
row_level:
filters:
- member: customers_city
operator: equals
values: ["{ userAttributes.city }"]
```
## Effect on Queries
When the access policy is implemented, queries will automatically be filtered based on the user's attributes. This ensures users can only access data that matches their attribute values.
# User groups
Source: https://docs.cube.dev/admin/users-and-permissions/user-groups
Organize Cube Cloud users into groups and tie those groups to data access policies instead of managing every account in isolation.
User groups allow you to organize users and manage access collectively.
Instead of assigning [user attributes][ref-user-attributes] to individual users, you can add users to groups for easier management at scale.
[Access policies][ref-dap] can be configured based on groups to control [row-level security][ref-rls].
## Creating groups
To create a user group:
1. Navigate to **Admin → User Groups**
2. Click **Create Group**
3. Enter a group name and optional description
4. Add users to the group
## Assigning roles to groups
Assigning roles to groups is currently in preview, and the user experience may still
change. Reach out to the [Cube support team](/admin/account-billing/support) to activate
this feature for your account.
Open a group and use the **Roles** section to assign [roles][ref-roles] to it. Every member
of the group gains the access that role grants, and removing the role from the group revokes
it from all members immediately.
Roles only ever add permissions — there is no precedence and no deny. A user with a Viewer
role who belongs to a group assigned Developer holds both, and is effectively a Developer.
This includes Admin: anyone who can edit a group's roles can grant administrative access
through it.
The **Effective access** section on a user's page lists the roles a user holds directly and
the roles conferred by a group, with the group named.
[ref-user-attributes]: /admin/users-and-permissions/user-attributes
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-rls]: /docs/data-modeling/access-control/row-level-security
[ref-dap]: /docs/data-modeling/data-access-policies
# Get app config
Source: https://docs.cube.dev/api-reference/app-theme/get-app-config
/api-reference/api.yaml get /api/v1/app-config
# Authentication
Source: https://docs.cube.dev/api-reference/authentication
Authenticate REST and SCIM requests with a token sent using the Bearer prefix.
How you authenticate depends on which API you call. The REST management API
(`/api/v1/…` and `/build/api/v1/…`) and the SCIM API (`/api/scim/v2/…`) all take
a token in the `Authorization` header with the `Bearer` prefix, and only HTTPS is
accepted.
## REST API (`/api/v1`)
Authenticate deployment, report, and workbook requests with a token sent using
the `Bearer` prefix:
```text theme={"dark"}
Authorization: Bearer
```
The token can be an API key generated in your Cube account settings, or an
OAuth access token from Cube's OAuth flow — both are sent the same way.
If the API key is [scoped to specific deployments](/admin/account-billing/api-keys#deployment-scope),
it only authorizes requests for those deployments. A request targeting a
deployment outside the key's scope returns `403 Forbidden`. Unscoped keys
(**All deployments**) authorize requests for every deployment.
```bash theme={"dark"}
curl --request GET \
--url 'https://.cubecloud.dev/api/v1/deployments' \
--header 'Authorization: Bearer '
```
The [data model endpoints](/api-reference/introduction#platform-api) under
`/build/api/v1/…` take the same token — only the path prefix differs.
## SCIM API (`/api/scim/v2`)
Authenticate user and group provisioning requests with a SCIM bearer token,
configured in your identity provider's SCIM integration:
```text theme={"dark"}
Authorization: Bearer
```
The SCIM API requires the `Bearer` prefix because identity providers (Microsoft
Entra ID, Okta, and others) send credentials this way per
[RFC 7644](https://datatracker.ietf.org/doc/html/rfc7644). The SCIM token is a
Cube API key, so `Api-Key ` is also accepted here.
Generate a SCIM token in your Cube account settings, then configure it in
your identity provider alongside the
[base URL](/api-reference/introduction#platform-api).
```bash theme={"dark"}
curl --request GET \
--url 'https://.cubecloud.dev/api/scim/v2/Users?count=100&startIndex=1' \
--header 'Authorization: Bearer '
```
Treat API keys and SCIM tokens like passwords — unless an API key is scoped to
specific deployments, they grant full access to your account's resources. Store
them securely and rotate them if they are exposed.
A `401 Unauthorized` response means the credential is missing, malformed, or
expired. A `403 Forbidden` response means the credential is valid but not
authorized for the target — for example, an API key scoped to other deployments.
# Abort chat stream
Source: https://docs.cube.dev/api-reference/chat/abort-chat-stream
/api-reference/chat.yaml post /chat/abort
Stops an in-progress chat stream — for example, when the user cancels a request or navigates away while the agent is still generating a response.
The abort URL is derived from the Chat API URL by replacing `/stream-chat-state` with `/abort`.
A successful cancellation is indicated solely by the `204 No Content` status — the response has no body. Note that the aborted `stream-chat-state` request itself does not fail with an HTTP error: the streaming response already returned `200 OK` when the stream opened, so treat this endpoint's `204` as the confirmation that the stream was aborted.
# Stream chat state
Source: https://docs.cube.dev/api-reference/chat/stream-chat-state
/api-reference/chat.yaml post /chat/stream-chat-state
Real-time streaming conversation with Cube's AI agents for analytics and data exploration.
The response is streamed as newline-delimited JSON (each line is a standalone JSON object
terminated by `\n`) and served with `Content-Type: application/json`. Each line represents a
chat message, a state update, or an error object. Assistant content arrives incrementally —
updates to a previously emitted message carry the same `id` with `isDelta: true`.
Provide `input` to send a new message. Omit `input` (with a `chatId`) to retrieve the current
thread state. If `chatId` is omitted a new thread is created and its id is returned on the
`__cutoff__` message.
Special sentinel messages:
- `__cutoff__` — initial marker emitted before new content; also carries the thread `chatId` and `isStreaming` under `state`.
- `__state__` — final thread snapshot emitted at the end, carrying the full `{ messages: [...] }` state.
Errors mid-stream are emitted as a standalone line `{ "error": "" }` rather than a non-200 status.
Finding the final answer: filter assistant messages where `graphPath[0] === "final"` and
`graphPath.length <= 2`; the last such message is the final answer.
Idempotency: provide a client-generated `messageId` in the format `-message` (a 13+ digit
Unix-millisecond timestamp). If a request with the same `messageId` arrives while a stream for that
message is still active on the thread, the API resubscribes the new request to the existing stream
instead of starting a new one.
# List a dashboard's embed access
Source: https://docs.cube.dev/api-reference/dashboard-embed-access/list-a-dashboards-embed-access
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/workbooks/{workbookId}/embed-access
Return the embed audience of a workbook's published dashboard: the global `allowEmbed` (signed-embedding) flag, whether it is shared with **all embed users**, and the list of individual embed **tenants** it is shared with.
Embed access is a read-only share of the workbook with a system sharing group, so this is scoped to the workbook that owns the dashboard. Requires **read** access to the workbook. Returns `404` if the workbook does not exist or belongs to a different deployment.
# Update a dashboard's embed access
Source: https://docs.cube.dev/api-reference/dashboard-embed-access/update-a-dashboards-embed-access
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/workbooks/{workbookId}/embed-access
Add, update, or remove one embed audience from a workbook's published dashboard, and/or toggle signed embedding. Returns the updated embed-access view.
Provide exactly one target — `embedTenantName` (a single embed tenant) or `allEmbedUsers: true` (every embed tenant) — together with `action`: `"read"` grants access, `"none"` removes it. Include `allowEmbed` to also flip "Allow signed embedding" in the same request; omit both targets to change only `allowEmbed`.
Requires **manage** access to the workbook. Returns `404` if the workbook or the named embed tenant does not exist, and `400` if the request specifies no change or sets `allowEmbed` on a workbook that has no published dashboard.
# Download a completed dashboard export
Source: https://docs.cube.dev/api-reference/dashboard-exports/download-a-completed-dashboard-export
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/dashboard-exports/{jobId}/download
Stream the rendered file as an attachment, named `dashboard.png` / `dashboard.pdf` unless `?filename=` supplies a different stem (the extension always follows the rendered format).
Returns `409` while the job is still rendering — poll `GET /dashboard-exports/{jobId}` first — `422` if the render failed (terminal: submit a new export rather than retrying this one), and `404` once the job has expired or if it belongs to a different user.
# Export a dashboard as PNG or PDF
Source: https://docs.cube.dev/api-reference/dashboard-exports/export-a-dashboard-as-png-or-pdf
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/dashboard-exports
Submit a render of a published dashboard. Identify it with either `dashboardId` (numeric) or `dashboardPublicId` (string) — supply exactly one; it must belong to this deployment, otherwise `404` is returned.
Returns as soon as the job is accepted, because a render waits for every widget to finish querying. Poll `GET /dashboard-exports/{jobId}` until `status` is `completed`, then fetch the file from `GET /dashboard-exports/{jobId}/download`. The job and its file are dropped at `expiresAt` (5 minutes) — download before then or submit again.
Pass `filters`, `timeGrains` and `memberSwitchers` to render the board with specific control values, exactly as a scheduled notification does; omit them to render its saved defaults.
By default the dashboard renders under the calling user's own security context, so the file shows what that user would see. Requires **manage** access to the dashboard's workbook — the same permission the console's own "Download as PNG / PDF" actions require, because a whole-dashboard render is a bulk data export.
To export on someone else's behalf, pass `renderAs`: the queries behind the image then run under that user's security context instead of yours. **Requires the caller to be a tenant administrator.** Name a console user by `userId` or `email` (`type: USER`), or an embed user by `embedTenantName` + `externalId` (`type: EMBED_USER`). A deactivated console user is not a valid subject and returns `404`; an email is matched case-insensitively.
An embed subject is provisioned if it does not exist yet, and `groups`, `userAttributes` and `securityContext` apply to **this render only** — they are carried on the render's own session and are NOT written to the user's stored context, so exporting for someone never changes what their live embedded sessions see. (Use the embed-tenant users API to change a user's stored context.) An embed subject additionally requires the dashboard's workbook to be shared with that embed tenant, otherwise `403` is returned.
Returns `403` when the workspace has data downloads restricted — including for an on-behalf-of export, which deliberately does **not** bypass that setting the way a built-in scheduled notification does — and `429` when there are already too many exports in flight for the identity being rendered as.
# Get the status of a dashboard export
Source: https://docs.cube.dev/api-reference/dashboard-exports/get-the-status-of-a-dashboard-export
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/dashboard-exports/{jobId}
Progress of a submitted render. Poll until `status` is `completed` (then download it) or `failed` (then read `error`); `pending` and `processing` both mean keep waiting.
Returns `404` once the job has expired, and for a job submitted by a different user — a job is only ever visible to whoever started it.
# Content hashes of the current data-model files (for upload diffing)
Source: https://docs.cube.dev/api-reference/data-model-uploads/content-hashes-of-the-current-data-model-files-for-upload-diffing
/api-reference/api.yaml get /build/api/v1/deployments/{deploymentId}/data-model/file-hashes
# Finish the upload transaction (prune removed files) and trigger a build
Source: https://docs.cube.dev/api-reference/data-model-uploads/finish-the-upload-transaction-prune-removed-files-and-trigger-a-build
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/data-model/upload/finish
# Start a data-model upload transaction
Source: https://docs.cube.dev/api-reference/data-model-uploads/start-a-data-model-upload-transaction
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/data-model/upload/start
# Upload a single file into an open transaction
Source: https://docs.cube.dev/api-reference/data-model-uploads/upload-a-single-file-into-an-open-transaction
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/data-model/upload/file
# Commit (and push) the active branch's pending changes
Source: https://docs.cube.dev/api-reference/data-model/commit-and-push-the-active-branchs-pending-changes
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/commit
# Compile a deployment's data model and report compilation errors
Source: https://docs.cube.dev/api-reference/data-model/compile-a-deployments-data-model-and-report-compilation-errors
/api-reference/api.yaml get /build/api/v1/deployments/{deploymentId}/data-model/validate
Asks the branch's own Cube runtime for `GET /cubejs-api/v1/meta` — the same call the console makes in dev mode — so the verdict is the one that branch's API would give. With neither parameter it validates the deploy branch (production). `branchName` validates a named branch, and `devMode=true` validates the caller's active dev-mode branch; the two are mutually exclusive. Answers 200 with `valid: false` and the compiler's errors when the model doesn't compile — a broken model is a result, not a request error.
# Create a branch (optionally entering dev mode on it)
Source: https://docs.cube.dev/api-reference/data-model/create-a-branch-optionally-entering-dev-mode-on-it
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/branches
# Create or overwrite data-model source files
Source: https://docs.cube.dev/api-reference/data-model/create-or-overwrite-data-model-source-files
/api-reference/api.yaml put /build/api/v1/deployments/{deploymentId}/data-model/files
# Delete a branch
Source: https://docs.cube.dev/api-reference/data-model/delete-a-branch
/api-reference/api.yaml delete /build/api/v1/deployments/{deploymentId}/branches
Deletes the branch and its git ref — the API-side counterpart of pruning a branch in Cube Cloud, for callers that create branches per run (a CI job syncing dbt on every pull request, say) and would otherwise accumulate them forever. Child branches are re-parented onto the deleted branch’s own parent, and anyone viewing the branch in dev mode is moved back to the default branch, so a branch that is currently open is deleted rather than refused. The deploy branch cannot be deleted (400) and an unknown branch is a 404.
`removeOnUpstream=true` additionally deletes the branch on the connected git provider (GitHub/GitLab). It defaults to `false`, so by default the branch is removed from Cube Cloud only and the ref on your own remote is left alone — a CI job cleaning up branches it created itself should opt in explicitly.
`branchName` is a query parameter rather than a path segment because branch names contain slashes (`dbt-sync/main-20260817120000-a1b2c3d4`).
# Delete data-model source files
Source: https://docs.cube.dev/api-reference/data-model/delete-data-model-source-files
/api-reference/api.yaml delete /build/api/v1/deployments/{deploymentId}/data-model/files
# Enable or disable a branch's staging environment
Source: https://docs.cube.dev/api-reference/data-model/enable-or-disable-a-branchs-staging-environment
/api-reference/api.yaml put /build/api/v1/deployments/{deploymentId}/branches/staging-environment
Enabling a branch keeps its staging environment always active and accessible at `/dev-mode/{branchName}/cubejs-api/v1`; disabled (the default) it is only active while the branch is viewed in Cube Cloud. Enabled branches are the ones listed by `GET /api/v1/deployments/{deploymentId}/environments?type=staging`. Only shared branches qualify — personal dev branches and the deploy branch are rejected.
# Enter dev mode (switch to a branch)
Source: https://docs.cube.dev/api-reference/data-model/enter-dev-mode-switch-to-a-branch
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/dev-mode
# Exit dev mode
Source: https://docs.cube.dev/api-reference/data-model/exit-dev-mode
/api-reference/api.yaml delete /build/api/v1/deployments/{deploymentId}/dev-mode
# List a deployment's branches
Source: https://docs.cube.dev/api-reference/data-model/list-a-deployments-branches
/api-reference/api.yaml get /build/api/v1/deployments/{deploymentId}/branches
# List the data-model source files for a deployment
Source: https://docs.cube.dev/api-reference/data-model/list-the-data-model-source-files-for-a-deployment
/api-reference/api.yaml get /build/api/v1/deployments/{deploymentId}/data-model/files
Returns the data-model file tree. `first`/`after` (from the shared cursor params) paginate the **top-level** tree nodes only — a single node can carry a large subtree, so this bounds the number of roots returned, not the total file count. Omit `first` to get the whole tree.
# Merge a branch into its parent branch (git-flow merge)
Source: https://docs.cube.dev/api-reference/data-model/merge-a-branch-into-its-parent-branch-git-flow-merge
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/merge
# Merge a branch straight into the deploy (default) branch
Source: https://docs.cube.dev/api-reference/data-model/merge-a-branch-straight-into-the-deploy-default-branch
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/merge-to-default
# Pull the latest for a branch from its remote and rebuild if changed
Source: https://docs.cube.dev/api-reference/data-model/pull-the-latest-for-a-branch-from-its-remote-and-rebuild-if-changed
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/pull
# Rename data-model source files
Source: https://docs.cube.dev/api-reference/data-model/rename-data-model-source-files
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/data-model/files/rename
# JSON query
Source: https://docs.cube.dev/api-reference/data/json-query
/api-reference/core-data.yaml post /v1/load
# Load Metadata
Source: https://docs.cube.dev/api-reference/data/load-metadata
/api-reference/core-data.yaml get /v1/meta
# SQL query
Source: https://docs.cube.dev/api-reference/data/sql-query
/api-reference/core-data.yaml post /v1/cubesql
Run a SQL query against the Cube [SQL API](/reference/core-data-apis/sql-api) and stream the results. The response is newline-delimited JSON: the first line carries the `schema` (column names and types), optionally `lastRefreshTime`, and optionally `usedPreAggregations` naming the pre-aggregations the result was served from; each subsequent line carries a `data` chunk with one or more result rows.
# Cancel a running dbt sync
Source: https://docs.cube.dev/api-reference/dbt-sync/cancel-a-running-dbt-sync
/api-reference/api.yaml delete /api/v1/deployments/{deploymentId}/dbt-sync/{syncJobId}
# Get the log of a dbt sync
Source: https://docs.cube.dev/api-reference/dbt-sync/get-the-log-of-a-dbt-sync
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/dbt-sync/{syncJobId}/logs
# Get the result of a completed dbt sync
Source: https://docs.cube.dev/api-reference/dbt-sync/get-the-result-of-a-completed-dbt-sync
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/dbt-sync/{syncJobId}/result
# Get the status of a dbt sync
Source: https://docs.cube.dev/api-reference/dbt-sync/get-the-status-of-a-dbt-sync
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/dbt-sync/{syncJobId}
# List dbt syncs for a deployment
Source: https://docs.cube.dev/api-reference/dbt-sync/list-dbt-syncs-for-a-deployment
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/dbt-sync
# Start a dbt sync for a deployment
Source: https://docs.cube.dev/api-reference/dbt-sync/start-a-dbt-sync-for-a-deployment
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/dbt-sync
# Create a deployment with an empty starter project and trigger its first build
Source: https://docs.cube.dev/api-reference/deployment-creation/create-a-deployment-with-an-empty-starter-project-and-trigger-its-first-build
/api-reference/api.yaml post /build/api/v1/deployments
Scaffolds an empty starter project and triggers the first build. When `creationMethod` is `github`, no starter project is scaffolded — connect a repository afterwards via POST /build/api/v1/deployments/:deploymentId/github/connect.
# Advance a deployment to a given onboarding step (e.g. after connecting a source)
Source: https://docs.cube.dev/api-reference/deployments/advance-a-deployment-to-a-given-onboarding-step-eg-after-connecting-a-source
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/creation-step/advance
# Delete deployment
Source: https://docs.cube.dev/api-reference/deployments/delete-deployment
/api-reference/api.yaml delete /api/v1/deployments/{deploymentId}
# Deployment token
Source: https://docs.cube.dev/api-reference/deployments/deployment-token
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/token
# Get deployment
Source: https://docs.cube.dev/api-reference/deployments/get-deployment
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}
# Get deployment settings
Source: https://docs.cube.dev/api-reference/deployments/get-deployment-settings
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/settings
Every setting of a deployment in one payload — the full contents of the console Settings pages (general, configuration flags, deploy, launchpad, Cube Store private storage), plus the read-only values displayed alongside them. Secrets are excluded; read environment variables through `/env-vars` instead. Writes go to `PUT /:deploymentId`, which accepts the same fields and returns this payload.
# Get deployments
Source: https://docs.cube.dev/api-reference/deployments/get-deployments
/api-reference/api.yaml get /api/v1/deployments
# Latest build/compile status for a branch (production build or dev-mode)
Source: https://docs.cube.dev/api-reference/deployments/latest-buildcompile-status-for-a-branch-production-build-or-dev-mode
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/build-status
# List available Cube versions
Source: https://docs.cube.dev/api-reference/deployments/list-available-cube-versions
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/versions
Every Cube version this deployment can be switched to — the same set the console’s version picker offers: the head of each release channel, plus the older versions the tenant has run before (its release history). Pass a listed `releaseChannelVersion` to `PUT /:deploymentId`; anything else is rejected. The container image is resolved server-side and is not part of the write surface.
# List the pods currently scheduled for a deployment
Source: https://docs.cube.dev/api-reference/deployments/list-the-pods-currently-scheduled-for-a-deployment
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/pods
# Recent runtime logs (tail) for a deployment. With no `source`, reads both the production API / worker pods and the dev-mode worker, merged chronologically; scope with `?source=production` or `?source=dev`.
Source: https://docs.cube.dev/api-reference/deployments/recent-runtime-logs-tail-for-a-deployment-with-no-`source`-reads-both-the-production-api-worker-pods-and-the-dev-mode-worker-merged-chronologically;-scope-with-`?source=production`-or-`?source=dev`
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/logs
# Reset a deployment to the first onboarding step (project)
Source: https://docs.cube.dev/api-reference/deployments/reset-a-deployment-to-the-first-onboarding-step-project
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/creation-step/reset
# Update a deployment
Source: https://docs.cube.dev/api-reference/deployments/update-a-deployment
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}
Update any subset of a deployment’s settings — everything the console’s Settings pages own — and read the full settings back. Fields you omit are left alone, and `templateVariables` is merged key-by-key rather than replaced, so setting one configuration flag never clears the others. `releaseChannelVersion` must be one of the versions `GET /:deploymentId/versions` lists for the target channel (in any of the forms it reports) — anything else is a 400, and the container image is resolved from the version rather than sent. Choosing a version that is not the channel’s latest also sets `releaseChannelVersionHold` unless you send it explicitly, so auto-upgrade does not undo the pin.
# Add embed group members
Source: https://docs.cube.dev/api-reference/embed-tenants/add-embed-group-members
/api-reference/api.yaml post /api/v1/embed-tenants/{embedTenantName}/groups/{id}/users
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Adds one or more embed users (1–1000 per request) to a group. Each member is identified by exactly one of `externalId` (the id passed to `generate-session`) or `embedUserId` (the numeric id this API returns). An `externalId` that has never signed in is provisioned as an embed user, so groups can be populated ahead of a user’s first session; an unknown `embedUserId` is a `404` instead. Memberships are added, never replaced: the user’s other groups — including the account-wide `groups` that drive data-model access — are left untouched. Idempotent: the response buckets every requested member into `addedMembers` or `unchangedMembers`.
# Create an embed group
Source: https://docs.cube.dev/api-reference/embed-tenants/create-an-embed-group
/api-reference/api.yaml post /api/v1/embed-tenants/{embedTenantName}/groups
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Creates a group inside the embed tenant. The name must be unique within the tenant and may not start with `system:` (that prefix is reserved for platform-managed groups). The group is immediately usable as a `tenantGroups` entry when generating a creator-mode embed session, and reaches the data model as `system:tenant:{embedTenantName}:{name}`. Returns `400` if the name is taken or reserved.
# Delete an embed group
Source: https://docs.cube.dev/api-reference/embed-tenants/delete-an-embed-group
/api-reference/api.yaml delete /api/v1/embed-tenants/{embedTenantName}/groups/{id}
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Deletes one of the embed tenant’s groups. Returns `400` while the group still has members — remove them first — and `404` if the group does not exist in this tenant.
# Delete embed tenant
Source: https://docs.cube.dev/api-reference/embed-tenants/delete-embed-tenant
/api-reference/api.yaml delete /api/v1/embed-tenants/{embedTenantName}
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Deletes an embed tenant and everything inside it — users, groups, memberships, attributes, policies, and content — plus the tenant's `system:tenant:*` sharing group. Idempotent: deleting a tenant that does not exist succeeds. The name must already be in canonical form (lowercase, per the `embedTenantName` rule); unlike the routes below, a name that differs only in case is rejected with `400` rather than being lowercased, so a typo can never cascade-delete a tenant the caller did not name exactly.
# Get an embed group
Source: https://docs.cube.dev/api-reference/embed-tenants/get-an-embed-group
/api-reference/api.yaml get /api/v1/embed-tenants/{embedTenantName}/groups/{id}
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Returns one of the embed tenant’s groups with its member count, or `404` if no such group exists in this tenant.
# List embed group members
Source: https://docs.cube.dev/api-reference/embed-tenants/list-embed-group-members
/api-reference/api.yaml get /api/v1/embed-tenants/{embedTenantName}/groups/{id}/users
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Lists the embed users that belong to a group, cursor-paginated, most recently added first. Each member carries both its numeric `id` and its `externalId`. Returns `404` if the group does not exist in this tenant.
# List embed groups
Source: https://docs.cube.dev/api-reference/embed-tenants/list-embed-groups
/api-reference/api.yaml get /api/v1/embed-tenants/{embedTenantName}/groups
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Lists the embed tenant's own groups (`embed_user_groups`), cursor-paginated, newest first, each with its member count. These are the groups an embed session references through `tenantGroups`; the implicit `system:tenant:*` sharing group is not included.
# List embed tenants
Source: https://docs.cube.dev/api-reference/embed-tenants/list-embed-tenants
/api-reference/api.yaml get /api/v1/embed-tenants
List the account's embed tenants, cursor-paginated. Supports `first`/`after` pagination and a case-insensitive `search` substring match on the tenant name. Requires an admin API key.
# List embed users
Source: https://docs.cube.dev/api-reference/embed-tenants/list-embed-users
/api-reference/api.yaml get /api/v1/embed-tenants/{embedTenantName}/users
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Lists the embed users belonging to one embed tenant, cursor-paginated and ordered by email. Alongside the numeric `id` this API uses, a user carries the `externalId` your integration knows it by (the id passed to `generate-session`) whenever it was recorded on the account — users provisioned before it was recorded come back without an `externalId`, though they stay addressable by it everywhere it is accepted as input. Use `search` to match a substring of the email or external id, or a complete external id, which resolves a user even when the response cannot echo it back.
# Provision an embed user
Source: https://docs.cube.dev/api-reference/embed-tenants/provision-an-embed-user
/api-reference/api.yaml post /api/v1/embed-tenants/{embedTenantName}/user
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Creates or updates one embed user from the identity your own system knows. An embed user is otherwise created the first time they generate a session, so this is how an integration puts its own directory into Cube up front: a user provisioned here is immediately addressable — listable, searchable, assignable to groups, and selectable wherever the product offers a choice of users — before they have ever opened the embed.
Idempotent, and the body is the desired state: provisioning an `externalId` that already exists updates it rather than failing. `email`, `userProfile.displayName` and `userProfile.picture` are overwritten when supplied and preserved when omitted; each group lane is REPLACED when supplied, preserved when omitted, and cleared by `[]`. The embed tenant itself is created on demand, so it need not exist yet.
Everything set here is exactly what `generate-session` would have set, and a later session for the same `externalId` re-applies whatever it carries — so provisioning changes when a user exists, never what their session grants them.
# Provision embed users in bulk
Source: https://docs.cube.dev/api-reference/embed-tenants/provision-embed-users-in-bulk
/api-reference/api.yaml post /api/v1/embed-tenants/{embedTenantName}/users
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Provisions up to 100 embed users in one request, each exactly as `POST /embed-tenants/{embedTenantName}/user` would. Use it to load a user directory into an embed tenant, and send further requests to cover a directory larger than one batch.
**This endpoint is partially successful.** It returns `200` whenever the request itself is well formed, and reports each user separately:
* `succeeded` — the users that were provisioned, in the order requested.
* `failed` — the entries that were not, each with the `externalId` and an `error` carrying the `status` and `message` the single-user endpoint would have returned (`400` for a group name that does not exist, `429` once the tenant is at its embed-user limit).
Users are applied one at a time in the order given, and earlier ones are not rolled back when a later one fails. Since provisioning is idempotent, resending the whole batch after fixing the failures is safe. A repeated `externalId` within one batch is applied once per entry, so the last one wins.
# Remove embed group members
Source: https://docs.cube.dev/api-reference/embed-tenants/remove-embed-group-members
/api-reference/api.yaml delete /api/v1/embed-tenants/{embedTenantName}/groups/{id}/users
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Removes one or more embed users (1–1000 per request) from a group, identified the same way as when adding them — exactly one of `externalId` or `embedUserId` each. Only this group’s membership is dropped; the users themselves and their other groups are left alone, and unlike adding, an unknown `externalId` is never provisioned. Idempotent: removing a user who is not a member is a no-op. Returns `204 No Content`, or `404` if the group does not exist in this tenant.
# Update an embed group
Source: https://docs.cube.dev/api-reference/embed-tenants/update-an-embed-group
/api-reference/api.yaml patch /api/v1/embed-tenants/{embedTenantName}/groups/{id}
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Updates a group’s description. The name is deliberately immutable: it is part of the security context an embed session carries (`system:tenant:{embedTenantName}:{name}`), so renaming would silently detach the data-model access policies written against it — create a new group instead. Pass an empty `description` to clear it.
# Enable or disable signed embedding for a dashboard
Source: https://docs.cube.dev/api-reference/embed/enable-or-disable-signed-embedding-for-a-dashboard
/api-reference/api.yaml patch /api/v1/embed/dashboard/{publicId}
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Sets the dashboard's "Allow signed embedding" flag (`allowEmbed`), which gates the customer-signed embedding path: with the flag off, embed sessions fetching `GET /api/v1/embed/dashboard/{publicId}` receive `403`. Trusted cube-cloud session flows (screenshoter, creator mode) and dashboards shared with all embed users are not affected by this flag.
Note that the flag alone does not make a dashboard embeddable — it must also be published and shared with the embed tenant. Returns `404` if the dashboard does not exist.
# Exchange a session for an embed token
Source: https://docs.cube.dev/api-reference/embed/exchange-a-session-for-an-embed-token
/api-reference/api.yaml post /api/v1/embed/session/token
Exchanges a one-time embed session id (created via `POST /api/v1/embed/generate-session`) for a signed, short-lived embed JWT used to authenticate the embedded analytics in the browser.
The session is **single-use**: it is consumed (deleted) on the first successful exchange, so a given `sessionId` can be redeemed only once. The returned token is signed with the tenant's embed secret, issued by `cubecloud`, and expires after 24 hours.
This endpoint is unauthenticated — it is called from the embedding client and the session id itself is the credential. Returns `401` if the session id is unknown or has already been redeemed.
# Generate an embed session
Source: https://docs.cube.dev/api-reference/embed/generate-an-embed-session
/api-reference/api.yaml post /api/v1/embed/generate-session
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Creates a one-time embed session for a deployment and returns its `sessionId`.
The session captures the embed context that will be baked into the embed token once redeemed — the target `deploymentId`, the end user's identity (`externalId` / `email` / `userProfile`), their group memberships, `userAttributes`, and an optional `securityContext`. Exchange the returned `sessionId` for a signed embed JWT via `POST /api/v1/embed/session/token` (single use).
The end user can be assigned to two independent kinds of group, which serve different purposes:
* **`groups`** — global, tenant-wide groups for **data-model access control**. Their names are passed verbatim into the Cube security context (`cubeCloud.groups`), where your data model's `access_policy` rules use them to gate cubes, views, members, and row-/column-level filters. They must already exist in the tenant.
* **`tenantGroups`** — per-embed-tenant groups for **sharing and organizing content within a single embed tenant** (e.g. sharing a workbook or dashboard with a group in Creator Mode). Requires `creatorMode: true` and `embedTenantName`; create them inline via `tenantGroupDefinitions`. In the security context they are namespaced as `system:tenant:{embedTenantName}:group:{group}`, so they never collide with a same-named global group.
See the individual request-body fields for the full contract.
`deploymentId` is required and the caller must have read access to it. Embedding must be enabled for the tenant, otherwise `403` is returned.
# Get an embeddable dashboard
Source: https://docs.cube.dev/api-reference/embed/get-an-embeddable-dashboard
/api-reference/api.yaml get /api/v1/embed/dashboard/{publicId}
Returns the dashboard identified by its `publicId` — including its reports/widgets — resolved within the caller's embed session scope, for rendering inside an embedded view.
Access is authorized against the embed user: the dashboard is returned when the user owns its workbook, has been granted workbook access (folder-based or direct), holds `EmbedDashboardRead` on the dashboard, or when it has been shared with all embed users (creator-mode flow). Returns `404` if the dashboard does not exist or is not visible to the caller.
Signed embedding additionally requires the dashboard to have "Allow signed embedding" enabled; the trusted screenshoter and creator-mode session flows are exempt. Returns `403` when embedding is not permitted for the dashboard.
# Run a python report over dashboard-filtered SQL
Source: https://docs.cube.dev/api-reference/embed/run-a-python-report-over-dashboard-filtered-sql
/api-reference/api.yaml post /api/v1/embed/dashboard/{publicId}/python-run
Re-runs the Python analysis attached to one of this dashboard's reports over `sqlQuery` — the report's own query with the dashboard's filters and time grains already applied — and returns the result **without persisting it**.
Authorized against the dashboard, exactly as `GET /api/v1/embed/dashboard/{publicId}` is. The `reportVersionId` must be one of that dashboard's published report snapshots and must belong to `reportId`; anything else returns `403`. The Python itself is read server-side from that snapshot — it is never supplied by the caller.
The analysis runs under the calling embed user's own security context, so row-level security applies to the viewer, and the stored result other viewers see is left untouched. Results are cached briefly per (report, code, SQL, user); failures are not cached. Returns `403` when Python analysis is not enabled for the account.
# Get env variables
Source: https://docs.cube.dev/api-reference/env-variables/get-env-variables
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/env-vars
Lists the deployment environment variables, cursor-paginated via `first`/`after`. Values of secret-named variables (containing PASS, SECRET, TOKEN, KEY, or CREDENTIAL) are returned as `[ENCRYPTED]`.
# Set env variables
Source: https://docs.cube.dev/api-reference/env-variables/set-env-variables
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/env-vars
Upserts deployment environment variables by name; variables not included keep their existing values. Passing the `[ENCRYPTED]` placeholder (the masked read value) is rejected — omit variables you do not intend to change. New variable names must be POSIX identifiers (`^[A-Za-z_][A-Za-z0-9_]*$`); names already stored on the deployment may be resent unchanged so a legacy name can still be updated or removed.
# Create deployment environment token
Source: https://docs.cube.dev/api-reference/environments/create-deployment-environment-token
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/environments/{environmentId}/tokens
# Create deployment environment token for meta sync
Source: https://docs.cube.dev/api-reference/environments/create-deployment-environment-token-for-meta-sync
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/environments/{environmentId}/tokens-for-meta-sync
# Get deployment environment tokens
Source: https://docs.cube.dev/api-reference/environments/get-deployment-environment-tokens
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/environments/{environmentId}/tokens
# Get deployment environments
Source: https://docs.cube.dev/api-reference/environments/get-deployment-environments
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/environments
# Create a folder
Source: https://docs.cube.dev/api-reference/folders/create-a-folder
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/folders
Create a new folder in a deployment's workspace.
Provide a `name`, and optionally a `parentId` to nest the folder inside an existing one (omit it to create the folder at the workspace root). An optional `position` controls the folder's ordering among its siblings.
Requires the **AI BI User** role (or higher). When `parentId` is set, the caller must additionally have **edit** or **manage** access to that parent folder.
Folders can be nested up to a fixed maximum depth; exceeding it returns `400`.
# Delete a folder
Source: https://docs.cube.dev/api-reference/folders/delete-a-folder
/api-reference/api.yaml delete /api/v1/deployments/{deploymentId}/folders/{folderId}
Delete a folder from a deployment's workspace. Responds with `204 No Content` on success.
The folder must be empty of sub-folders: if it still has child folders, the request is rejected with `400` — delete or move the sub-folders first.
Content directly inside the folder (workbooks, dashboards, reports) is **not** deleted. It is detached and returned to the workspace root.
Requires **manage** access to the folder. Returns `404` if the folder does not exist or belongs to a different deployment.
# List folder ancestors
Source: https://docs.cube.dev/api-reference/folders/list-folder-ancestors
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/folders/{folderId}/ancestors
Return the ancestor chain of a folder, ordered from the workspace root down to (and including) the folder itself.
Use this to render a breadcrumb path for a nested folder. For example, for `Root / Sales / Q1` requested on the `Q1` folder, the response is `[Root, Sales, Q1]`.
Requires visibility access to the target folder. Returns `404` if the folder does not exist or belongs to a different deployment.
# List folders
Source: https://docs.cube.dev/api-reference/folders/list-folders
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/folders
List the folders in a deployment's workspace.
By default this returns the folders at the workspace root. Pass the `parentId` query parameter to list the direct children of a specific folder instead.
Results are scoped to the calling user: only folders the user can see are returned. A folder is visible when the user owns it, has been granted access to it directly, or has access to something inside it (a sub-folder, workbook, or dashboard), in which case the ancestor folders are surfaced as navigation.
This endpoint is not recursive — it returns a single level of the folder tree per call. Use `parentId` to walk deeper, or `GET /folders/{folderId}/ancestors` to resolve a breadcrumb path.
Results are returned in pages using cursor-based pagination: pass `first` to set the page size and `after` (the previous response's `pageInfo.endCursor`) to fetch the next page. The legacy `data` and `count` fields are still populated for backwards compatibility but are deprecated in favour of `items` and `pageInfo`.
# Update a folder
Source: https://docs.cube.dev/api-reference/folders/update-a-folder
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/folders/{folderId}
Update a folder's metadata — its `name` and/or its `position` among its siblings. Both fields are optional; only the fields you send are changed.
This endpoint does **not** move a folder to a different parent. To re-parent a folder, use `POST /workspace/move` with `type: FOLDER`.
Requires **edit** access to the folder. Returns `404` if the folder does not exist or belongs to a different deployment.
# Connect a deployment to a GitHub repository and trigger a build
Source: https://docs.cube.dev/api-reference/github-connection/connect-a-deployment-to-a-github-repository-and-trigger-a-build
/api-reference/api.yaml post /build/api/v1/deployments/{deploymentId}/github/connect
# List branches of a repository
Source: https://docs.cube.dev/api-reference/github/list-branches-of-a-repository
/api-reference/api.yaml get /api/v1/github/repositories/{owner}/{repo}/branches
# List repositories accessible to a GitHub App installation
Source: https://docs.cube.dev/api-reference/github/list-repositories-accessible-to-a-github-app-installation
/api-reference/api.yaml get /api/v1/github/installations/{installationId}/repositories
# List the user's GitHub App installations (orgs/accounts)
Source: https://docs.cube.dev/api-reference/github/list-the-users-github-app-installations-orgsaccounts
/api-reference/api.yaml get /api/v1/github/installations
# Whether GitHub is linked, plus the link/install URLs
Source: https://docs.cube.dev/api-reference/github/whether-github-is-linked-plus-the-linkinstall-urls
/api-reference/api.yaml get /api/v1/github/status
# Introduction
Source: https://docs.cube.dev/api-reference/introduction
Cube's HTTP APIs — query your data with the Core Data API, stream agent conversations with the AI API, and manage Cube with the Platform API.
Cube exposes three HTTP API families, each for a different job and with slightly
different authentication:
| API | What it's for | Authentication |
| ----------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- |
| **Core Data API** | Query your data model over HTTP — run queries (`/v1/load`) and read metadata (`/v1/meta`) | Cube API token (a JWT) sent **directly** in the `Authorization` header — no prefix |
| **AI API** | Stream conversations with Cube AI agents (`/chat/stream-chat-state`) | Cube API key with the **`Api-Key`** prefix |
| **Platform API** | Manage Cube — deployments, data models, reports, workbooks, users, policies, and more | **`Bearer`** prefix for REST and SCIM 2.0 |
The exact header for each is shown above. The Platform and SCIM schemes are
detailed on the [Authentication](/api-reference/authentication) page; the Core
Data and AI headers are also documented on their own reference pages.
## Available endpoints
Each API family has its own base URL — copy the exact host from the Cube
interface where noted. Only HTTPS is accepted, and every request must be
authenticated (see [Authentication](/api-reference/authentication)).
### Core Data API
Base URL — your deployment's data API host (the `/cubejs-api` base path is configurable):
```text theme={"dark"}
https://{deployment}.{region}.cubecloudapp.dev/cubejs-api
```
| Endpoint | Path |
| --------------------------------------------- | ------------------ |
| [JSON query](/api-reference/data/json-query) | `POST /v1/load` |
| [SQL query](/api-reference/data/sql-query) | `POST /v1/cubesql` |
| [Metadata](/api-reference/data/load-metadata) | `GET /v1/meta` |
### AI API
Base URL — your agent's Chat API URL (copy it from **Admin → Agents → Chat API URL**):
```text theme={"dark"}
https://ai.{cloudRegion}.cubecloud.dev/api/v1/public/{accountName}/agents/{agentId}
```
| Endpoint | Path |
| ---------------------------------------------------------- | ------------------------------ |
| [Stream chat state](/api-reference/chat/stream-chat-state) | `POST /chat/stream-chat-state` |
### Platform API
Base URL — your tenant host:
```text theme={"dark"}
https://{tenant}.cubecloud.dev
```
Endpoints live under three path prefixes on that host, all taking the same token:
| Prefix | What it covers |
| ----------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| `/api/v1/…` | Management of deployments, workspace content, users, policies, and embedding |
| `/build/api/v1/…` | Data model authoring — files, uploads, dev mode, and branches. Routed to the build workers that own a deployment's data model. |
| `/api/scim/v2/…` | SCIM 2.0 user and group provisioning |
Resources by entity:
| Entity | Resource | Version |
| --------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ | -------- |
| [Deployments](/api-reference/deployments/get-deployments) | `/api/v1/deployments` | v1 |
| [Deployment Creation](/api-reference/deployment-creation/create-a-deployment-with-an-empty-starter-project-and-trigger-its-first-build) | `/build/api/v1/deployments` | v1 |
| [Environments](/api-reference/environments/get-deployment-environments) | `/api/v1/deployments/{deploymentId}/environments` | v1 |
| [Env Variables](/api-reference/env-variables/get-env-variables) | `/api/v1/deployments/{deploymentId}/env-vars` | v1 |
| [Regions](/api-reference/regions/list-regions) | `/api/v1/regions` | v1 |
| [Data Model](/api-reference/data-model/list-a-deployments-branches) | `/build/api/v1/deployments/{deploymentId}` | v1 |
| [Data Model Uploads](/api-reference/data-model-uploads/content-hashes-of-the-current-data-model-files-for-upload-diffing) | `/build/api/v1/deployments/{deploymentId}/data-model` | v1 |
| [GitHub](/api-reference/github/list-repositories-accessible-to-a-github-app-installation) | `/api/v1/github` | v1 |
| [GitHub Connection](/api-reference/github-connection/connect-a-deployment-to-a-github-repository-and-trigger-a-build) | `/build/api/v1/deployments/{deploymentId}/github/connect` | v1 |
| [dbt Sync](/api-reference/dbt-sync/list-dbt-syncs-for-a-deployment) | `/api/v1/deployments/{deploymentId}/dbt-sync` | v1 |
| [Folders](/api-reference/folders/list-folders) | `/api/v1/deployments/{deploymentId}/folders` | v1 |
| [Reports](/api-reference/reports/list-reports) | `/api/v1/deployments/{deploymentId}/reports` | v1 |
| [Workbooks](/api-reference/workbooks/get-workbooks) | `/api/v1/deployments/{deploymentId}/workbooks` | v1 |
| [Notifications](/api-reference/notifications/list-scheduled-notifications) | `/api/v1/deployments/{deploymentId}/notifications` | v1 |
| [Workspace](/api-reference/workspace/list-shared-workspace-items) | `/api/v1/deployments/{deploymentId}` | v1 |
| [Users](/api-reference/users/update-my-settings) | `/api/v1/users/me/settings` | v1 |
| [Users Admin](/api-reference/users-admin/create-user) | `/api/v1/users` | v1 |
| [User Attributes](/api-reference/user-attributes/get-user-attributes) | `/api/v1/user-attributes` | v1 |
| [User Attribute Values](/api-reference/user-attribute-values/upsert-user-attribute-value) | `/api/v1/user-attribute-values` | v1 |
| [Tenant Settings](/api-reference/tenant-settings/get-tenant-settings) | `/api/v1/tenant/settings` | v1 |
| [OAuth Integrations](/api-reference/oauth-integrations/list-oauth-integrations) | `/api/v1/oauth-integrations` | v1 |
| [User OAuth Tokens](/api-reference/user-oauth-tokens/list-user-oauth-tokens) | `/api/v1/user-oauth-tokens` | v1 |
| [OIDC Token Configs](/api-reference/oidc-token-configs/list-oidc-token-configs) | `/api/v1/oidc-token-configs` | v1 |
| [App Theme](/api-reference/app-theme/get-app-config) | `/api/v1/app-config` | v1 |
| [Embed](/api-reference/embed/get-an-embeddable-dashboard) | `/api/v1/embed` | v1 |
| [Embed Tenants](/api-reference/embed-tenants/list-embed-tenants) | `/api/v1/embed-tenants` | v1 |
| [Dashboard Embed Access](/api-reference/dashboard-embed-access/list-a-dashboards-embed-access) | `/api/v1/deployments/{deploymentId}/workbooks/{workbookId}/embed-access` | v1 |
| [Usage Analytics](/api-reference/usage-analytics/issue-a-cube-api-token-for-usage-analytics) | `/api/v1/usage-analytics/token` | v1 |
| [OpenAPI Spec](/api-reference/openapi-spec/get-the-openapi-specification) | `/api/v1/spec` | v1 |
| [Dashboard Exports](/api-reference/dashboard-exports/export-a-dashboard-as-png-or-pdf) | `/api/v1/deployments/{deploymentId}/dashboard-exports` | v1 |
| [Users (SCIM)](/api-reference/scim-users/list-users) | `/api/scim/v2/Users` | SCIM 2.0 |
| [Groups (SCIM)](/api-reference/scim-groups/list-groups) | `/api/scim/v2/Groups` | SCIM 2.0 |
## Client libraries
### Core Data API
Query the Core Data API from JavaScript or TypeScript with Cube's JS client,
[`@cubejs-client/core`](https://www.npmjs.com/package/@cubejs-client/core), which
also ships React, Vue, and Angular bindings and a WebSocket transport for
real-time updates. See the [JavaScript SDK](/reference/javascript-sdk) reference
for full usage.
```bash theme={"dark"}
npm install @cubejs-client/core
```
### Platform API
The recommended way to call the Platform API from JavaScript or TypeScript is the
official client, [`@cube-dev/platform-client`](https://www.npmjs.com/package/@cube-dev/platform-client).
It wraps the OpenAPI spec on this site with end-to-end types for every endpoint,
request, and response, plus optional [React Query](https://tanstack.com/query)
bindings.
```bash npm theme={"dark"}
npm install @cube-dev/platform-client
```
```bash yarn theme={"dark"}
yarn add @cube-dev/platform-client
```
```bash pnpm theme={"dark"}
pnpm add @cube-dev/platform-client
```
Create a client with your tenant's base URL and an auth header (see
[Authentication](/api-reference/authentication)), then call any endpoint through
the fully-typed `fetchClient`:
```ts theme={"dark"}
import { createCubePlatformClient } from '@cube-dev/platform-client';
const client = createCubePlatformClient({
baseUrl: 'https://.cubecloud.dev',
// Returned on every request — provide the Platform API auth header.
getHeaders: () => ({ Authorization: `Bearer ${process.env.CUBE_API_KEY}` }),
});
const { data, error } = await client.fetchClient.GET('/api/v1/deployments/', {
params: { query: { first: 50 } },
});
for (const deployment of data?.items ?? []) {
console.log(deployment.id, deployment.name);
}
```
#### React Query bindings
The `@cube-dev/platform-client/react-query` entry point adds a provider and typed
hooks built on `@tanstack/react-query` (a peer dependency alongside `react`):
```tsx theme={"dark"}
import { createCubePlatformClient } from '@cube-dev/platform-client';
import {
CubePlatformApiProvider,
useCubePlatformApiQuery,
} from '@cube-dev/platform-client/react-query';
const client = createCubePlatformClient({ baseUrl, getHeaders });
function App() {
return (
);
}
function Deployments() {
const { data } = useCubePlatformApiQuery('get', '/api/v1/deployments/');
return
{data?.items.map((d) =>
{d.name}
)}
;
}
```
Schema types are exported as `PlatformApiSchemas` (e.g.
`PlatformApiSchemas['Deployment']`). See the package
[`CHANGELOG`](/api-reference/changelog) for release notes and breaking changes.
# Add notification recipients
Source: https://docs.cube.dev/api-reference/notifications/add-notification-recipients
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/notifications/{id}/recipients
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Subscribes one or more recipients (1–1000 per request) to a notification. Each recipient is one of: a main console `USER` (identified by `userId` or `email`); an `EMBED_USER` (identified by `embedTenantName` + `externalId`, auto-provisioned if it does not yet exist, and optionally given `securityContext`, `userAttributes`, and `groups` that drive per-recipient rendering); or `SLACK` — **not yet supported**, a request containing a Slack recipient is rejected with `400`. Every recipient must resolve to a valid email (an embed user’s `email`, or an email-shaped `externalId`); otherwise the whole request fails with `400` before anything is written. The operation is idempotent: the response buckets each recipient into `createdRecipients`, `updatedRecipients` (an existing recipient whose embed properties changed), or `unchangedRecipients`.
# Create a scheduled notification
Source: https://docs.cube.dev/api-reference/notifications/create-a-scheduled-notification
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/notifications
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Creates a scheduled notification — a recurring run of a dashboard whose rendered result is delivered to its recipients. The cadence is set by `scheduleType` plus the relevant time fields (`minute`, `hour`, `dayOfWeek`, `dayOfMonth`) or a raw `customCron` expression when `scheduleType` is `CUSTOM`; `timezone` defaults to UTC. Identify the target dashboard with either `dashboardId` (numeric) or `dashboardPublicId` (string) — supply exactly one; it must belong to this deployment, otherwise `404` is returned. The created notification has no recipients — add them via `POST /notifications/{id}/recipients`.
# Delete a scheduled notification
Source: https://docs.cube.dev/api-reference/notifications/delete-a-scheduled-notification
/api-reference/api.yaml delete /api/v1/deployments/{deploymentId}/notifications/{id}
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Permanently deletes a scheduled notification and all of its recipients (main, embed, and Slack) in a single cascade. Returns `204 No Content` on success, or `404` if the notification does not belong to this deployment.
# Get a scheduled notification
Source: https://docs.cube.dev/api-reference/notifications/get-a-scheduled-notification
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/notifications/{id}
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Returns a single scheduled notification by id, including its cron expression, timezone, enabled state, notification format, and a human-readable schedule summary. Returns `404` if the notification does not exist or does not belong to this deployment.
# List notification recipients
Source: https://docs.cube.dev/api-reference/notifications/list-notification-recipients
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/notifications/{id}/recipients
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Returns the recipients subscribed to a notification, cursor-paginated and ordered by creation time (newest first). Recipients span three kinds, distinguished by `type`: `USER` (a main console user, with `userId` + `email`), `EMBED_USER` (an embed user, with `embedUserId`, `embedTenantName`, `externalId`, and `email`), and `SLACK` (a channel, with `channelId` + `channelName`). Returns `404` if the notification is not part of this deployment.
# List scheduled notifications
Source: https://docs.cube.dev/api-reference/notifications/list-scheduled-notifications
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/notifications
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Returns the deployment’s scheduled notifications (recurring dashboard runs), ordered by creation time (newest first) and cursor-paginated. Optionally filter to a single dashboard with either `dashboardId` (numeric) or `dashboardPublicId` (string), and/or to the notifications a single recipient receives — a user (`recipientUserId` or `recipientEmail`) or an embed user (`recipientEmbedTenantName` + `recipientExternalId`). At most one identifier per dashboard/recipient dimension. Each item describes the schedule only; recipients are managed through the `/recipients` sub-resource.
# Remove notification recipients
Source: https://docs.cube.dev/api-reference/notifications/remove-notification-recipients
/api-reference/api.yaml delete /api/v1/deployments/{deploymentId}/notifications/{id}/recipients
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Unsubscribes one or more recipients (1–1000 per request) from a notification. Each entry is identified by `type` plus the `id` returned by the API — `userId` for `USER`, `embedUserId` for `EMBED_USER`, or `channelId` for `SLACK`. `EMBED_USER` entries must also include `embedTenantName`, which locates the recipient’s storage partition. Removals are idempotent (deleting a recipient that isn’t subscribed is a no-op). Returns `204 No Content`, or `404` if the notification is not part of this deployment.
# Update a scheduled notification
Source: https://docs.cube.dev/api-reference/notifications/update-a-scheduled-notification
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/notifications/{id}
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Updates a notification’s schedule definition. Any subset of the schedule fields may be supplied (`scheduleType` + time fields or `customCron`, `timezone`), along with `isEnabled`, `notificationEnabled`, and `notificationFormat`. Omitted fields are left unchanged. This endpoint never modifies recipients — manage those through the `/recipients` sub-resource. Returns `404` if the notification is not part of this deployment.
# Create OAuth integration
Source: https://docs.cube.dev/api-reference/oauth-integrations/create-oauth-integration
/api-reference/api.yaml post /api/v1/oauth-integrations
# Delete OAuth integration
Source: https://docs.cube.dev/api-reference/oauth-integrations/delete-oauth-integration
/api-reference/api.yaml delete /api/v1/oauth-integrations/{id}
# Get OAuth integration
Source: https://docs.cube.dev/api-reference/oauth-integrations/get-oauth-integration
/api-reference/api.yaml get /api/v1/oauth-integrations/{id}
# List OAuth integrations
Source: https://docs.cube.dev/api-reference/oauth-integrations/list-oauth-integrations
/api-reference/api.yaml get /api/v1/oauth-integrations
# Update OAuth integration
Source: https://docs.cube.dev/api-reference/oauth-integrations/update-oauth-integration
/api-reference/api.yaml put /api/v1/oauth-integrations/{id}
# Create OIDC token config
Source: https://docs.cube.dev/api-reference/oidc-token-configs/create-oidc-token-config
/api-reference/api.yaml post /api/v1/oidc-token-configs
# Delete OIDC token config
Source: https://docs.cube.dev/api-reference/oidc-token-configs/delete-oidc-token-config
/api-reference/api.yaml delete /api/v1/oidc-token-configs/{id}
# Get OIDC token config
Source: https://docs.cube.dev/api-reference/oidc-token-configs/get-oidc-token-config
/api-reference/api.yaml get /api/v1/oidc-token-configs/{id}
# List OIDC token configs
Source: https://docs.cube.dev/api-reference/oidc-token-configs/list-oidc-token-configs
/api-reference/api.yaml get /api/v1/oidc-token-configs
# Update OIDC token config
Source: https://docs.cube.dev/api-reference/oidc-token-configs/update-oidc-token-config
/api-reference/api.yaml put /api/v1/oidc-token-configs/{id}
# Get the OpenAPI specification
Source: https://docs.cube.dev/api-reference/openapi-spec/get-the-openapi-specification
/api-reference/api.yaml get /api/v1/spec
The full OpenAPI 3.1 document for this API, as served by the build handling the request — every path, parameter, request body and schema. Intended for runtime discovery by clients and agents; `cube spec` wraps it with filtering.
# List regions
Source: https://docs.cube.dev/api-reference/regions/list-regions
/api-reference/api.yaml get /api/v1/regions
Lists the regions available to the account, cursor-paginated via `first`/`after`. The legacy `data` field returns the full, unpaginated list and is deprecated — use `items` + `pageInfo` instead.
# Connect a report to the calling spreadsheet
Source: https://docs.cube.dev/api-reference/reports/connect-a-report-to-the-calling-spreadsheet
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/reports/{reportId}/connect-workbook
Link a report to the caller's own spreadsheet by recording its placement
(workbook id + result location) so the add-in can list and refresh it there.
This upserts only the placement for the given workbook and never changes the
report's query, name, or other definition — so, like refresh, it requires only
read access to the report. A user who can view a report can place its data into
their own sheet without being its creator or having edit rights.
# Create a report
Source: https://docs.cube.dev/api-reference/reports/create-a-report
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/reports
Create a report in a deployment.
You may supply your own `publicId` (a 12-character `[0-9A-Za-z]` id) to key the report for as-code management; omit it to have one generated. It must be an id you chose — the auto-generated placeholder ids shown for reports that have no stable id yet are a reserved, non-unique shape and are rejected with `400`.
# Create or update a report by publicId
Source: https://docs.cube.dev/api-reference/reports/create-or-update-a-report-by-publicid
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/reports/by-public-id/{publicId}
Idempotent create-or-update ("upsert") of a report keyed by its portable `publicId` — the stable identity for managing dashboards as code (CI/CD pipelines re-applying the same definition against any environment).
If a report with this `publicId` exists in the deployment, it is updated with the fields you send (same semantics as `PUT /reports/:reportId`); otherwise a new report is created with this `publicId`. Re-applying the same request is a no-op. The `publicId` must be a 12-character alphanumeric id (`[0-9A-Za-z]`); mint one client-side when authoring the definition.
This endpoint returns `409` for three different reasons, distinguished by the response body's `code` — handle them differently:
* **No `code`** — the `publicId` already belongs to a report in a **different** deployment. Because `publicId` is unique across the account, it is rejected rather than silently updating across deployments (reported regardless of whether you have access to that report). This is **permanent**: retrying will not help; use a different `publicId`.
* **`code: "upsert_branch_changed"`** — whether this request creates or updates is decided, and access-checked, before it is applied; another writer changed that in between (created or deleted the report with this `publicId`). The request was rejected rather than applied under the wrong permission check. This is **transient: retry the request**. It only happens when two writers target the same `publicId` concurrently, e.g. two pipeline runs applying the same bundle at once.
* **`code: "ambiguous_legacy_id"`** — the id given is one of the ids synthesized for reports that predate stable ids (see below), and it matches more than one of them, so it cannot identify a single report. Nothing was changed. This is **permanent**, and the fix is different from the case above: assign the report you mean a `publicId` of your own with `PUT /deployments/{deploymentId}/reports/{reportId}`, then key on that.
**Reports created before stable ids** do not store a `publicId`; the API synthesizes one for them from the report's internal id. Those synthesized ids are **not unique** — several reports can share one, and in practice most do. This endpoint resolves such an id only when it is unambiguous, adopting it as the report's real `publicId` at that point; otherwise it returns the `ambiguous_legacy_id` conflict above.
For anything you intend to manage as code, do not rely on a synthesized id: give the report a `publicId` you chose. `PUT /deployments/{deploymentId}/reports/{reportId}` accepts `publicId` for exactly this, on any report that does not have one yet (it is write-once — a report's existing `publicId` cannot be changed, since clients may already have stored it).
`source` is **create-only**: it is recorded when the report is first created and ignored on subsequent updates (it describes where the report originated, not its current definition).
Access: updating requires **edit** access to the existing report; creating requires the same access as `POST /reports`.
# Delete report
Source: https://docs.cube.dev/api-reference/reports/delete-report
/api-reference/api.yaml delete /api/v1/deployments/{deploymentId}/reports/{reportId}
# Disconnect a report from one spot in a spreadsheet
Source: https://docs.cube.dev/api-reference/reports/disconnect-a-report-from-one-spot-in-a-spreadsheet
/api-reference/api.yaml delete /api/v1/deployments/{deploymentId}/reports/{reportId}/connect-workbook
Remove a single placement of a report — one sheet + anchor in one
spreadsheet — leaving every other placement, in this workbook and in others, alone.
Address the placement by `placementId`, or (for placements recorded before ids
existed) by `externalWorkbookId` plus `sheetName` and `anchorCell` — or a
`resultLocation` naming both. Neither the sheet name nor the anchor is optional in
that form: matching is by sheet + anchor, and a target missing either addresses
nothing.
`sheetId` is an optional extra, not a substitute for `sheetName`: a placement
written before that field existed has only a name, and no id the server could
resolve to it, so an id-only address would match nothing while looking correct.
Send a `sheetId` only when you know it is the one the placement was stored with.
It NARROWS the target — a placement recorded under a different id is not matched,
even when the sheet name and anchor agree, because the stored name may belong to a
tab that has since been renamed and the id is the only stable identity. If your id
may have moved on (a tab deleted and recreated under the same name reports a new
one), address the placement by `placementId`, or by sheet name with no
`sheetId` at all.
A target that matches no placement succeeds and changes nothing, so retrying is
safe.
Like connecting, this records nothing about the report's definition and so requires
only read access: unplacing a report from your own sheet is not an edit to the report.
# Get report
Source: https://docs.cube.dev/api-reference/reports/get-report
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/reports/{reportId}
# List reports
Source: https://docs.cube.dev/api-reference/reports/list-reports
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/reports
List the reports in a deployment, scoped to the reports the calling user can access.
Results are returned in pages using cursor-based pagination: pass `first` to set the page size and `after` (the previous response's `pageInfo.endCursor`) to fetch the next page.
Cursor pagination is not supported when sorting by `lastViewedAt` (a per-viewer sort with no stable cursor column).
# Refresh report
Source: https://docs.cube.dev/api-reference/reports/refresh-report
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/reports/{reportId}/refresh
# Update a report
Source: https://docs.cube.dev/api-reference/reports/update-a-report
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/reports/{reportId}
Update a report. All fields are optional; only the fields you send are changed.
`publicId` is **write-once**: you may assign one to a report that does not have one yet — which is how a report created before stable ids is adopted into as-code management, so `PUT /reports/by-public-id/{publicId}` can key on it — but a report's existing `publicId` cannot be changed, since clients may already have stored it. Attempting to change it returns `400`; a `publicId` already used by another report returns `409`.
The id you assign must be one you chose (any distinct 12-character `[0-9A-Za-z]` id). The auto-generated placeholder id shown for a report that has no stable id yet cannot be assigned — it is not unique — and is rejected with `400`.
Requires **edit** access to the report.
# Create group
Source: https://docs.cube.dev/api-reference/scim-groups/create-group
/api-reference/scim.yaml post /scim/v2/Groups
Provisions a new group.
# Delete group
Source: https://docs.cube.dev/api-reference/scim-groups/delete-group
/api-reference/scim.yaml delete /scim/v2/Groups/{id}
Deletes a group by its SCIM `id`. Member users are not deleted.
# Get group
Source: https://docs.cube.dev/api-reference/scim-groups/get-group
/api-reference/scim.yaml get /scim/v2/Groups/{id}
Returns a single group by its SCIM `id`.
# List groups
Source: https://docs.cube.dev/api-reference/scim-groups/list-groups
/api-reference/scim.yaml get /scim/v2/Groups
Returns a paginated ListResponse of groups.
# Patch group
Source: https://docs.cube.dev/api-reference/scim-groups/patch-group
/api-reference/scim.yaml patch /scim/v2/Groups/{id}
Applies a partial update to a group using a SCIM PatchOp. Commonly used
to add or remove members without resending the entire membership list.
# Replace group
Source: https://docs.cube.dev/api-reference/scim-groups/replace-group
/api-reference/scim.yaml put /scim/v2/Groups/{id}
Replaces all attributes of an existing group, including its full membership list.
# Create user
Source: https://docs.cube.dev/api-reference/scim-users/create-user
/api-reference/scim.yaml post /scim/v2/Users
Provisions a new user.
# Delete user
Source: https://docs.cube.dev/api-reference/scim-users/delete-user
/api-reference/scim.yaml delete /scim/v2/Users/{id}
De-provisions (deletes) a user by their SCIM `id`.
# Get user
Source: https://docs.cube.dev/api-reference/scim-users/get-user
/api-reference/scim.yaml get /scim/v2/Users/{id}
Returns a single user by their SCIM `id`.
# List users
Source: https://docs.cube.dev/api-reference/scim-users/list-users
/api-reference/scim.yaml get /scim/v2/Users
Returns a paginated [ListResponse](https://datatracker.ietf.org/doc/html/rfc7644#section-3.4.2)
of users. Use the `filter` parameter to look a user up by `userName`,
for example `userName eq "jane@example.com"`.
# Patch user
Source: https://docs.cube.dev/api-reference/scim-users/patch-user
/api-reference/scim.yaml patch /scim/v2/Users/{id}
Applies a partial update to a user using a SCIM
[PatchOp](https://datatracker.ietf.org/doc/html/rfc7644#section-3.5.2).
Commonly used to deactivate a user by setting `active` to `false`.
# Replace user
Source: https://docs.cube.dev/api-reference/scim-users/replace-user
/api-reference/scim.yaml put /scim/v2/Users/{id}
Replaces all attributes of an existing user with the supplied representation.
# Get tenant settings
Source: https://docs.cube.dev/api-reference/tenant-settings/get-tenant-settings
/api-reference/api.yaml get /api/v1/tenant/settings
# Update tenant settings
Source: https://docs.cube.dev/api-reference/tenant-settings/update-tenant-settings
/api-reference/api.yaml put /api/v1/tenant/settings
# Issue a Cube API token for Usage Analytics
Source: https://docs.cube.dev/api-reference/usage-analytics/issue-a-cube-api-token-for-usage-analytics
/api-reference/api.yaml post /api/v1/usage-analytics/token
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Issues a short-lived Cube API token for Cube's internal **Usage Analytics** deployment — the data shown on **Admin → Usage Analytics** — so you can query your own usage and billing data programmatically (REST API, BI tools, scripts) instead of only through the embedded page.
The token is scoped to your account: it carries the same server-resolved security context as the embedded Usage Analytics page, so queries only ever see your own data. Send it as `Authorization: Bearer ` on requests to the returned `apiUrl` (the deployment's [Cube REST API](https://cube.dev/docs/product/apis-integrations/rest-api)).
Tokens expire after 1 hour (see `expiresAt`). Issue a fresh one per session or scheduled refresh — issuing is idempotent and cheap.
Requires Usage Analytics to be enabled for the account, otherwise `404` is returned.
# Get user attribute values
Source: https://docs.cube.dev/api-reference/user-attribute-values/get-user-attribute-values
/api-reference/api.yaml get /api/v1/user-attribute-values/{userId}
# Upsert user attribute value
Source: https://docs.cube.dev/api-reference/user-attribute-values/upsert-user-attribute-value
/api-reference/api.yaml post /api/v1/user-attribute-values
# Create user attribute
Source: https://docs.cube.dev/api-reference/user-attributes/create-user-attribute
/api-reference/api.yaml post /api/v1/user-attributes
# Get user attributes
Source: https://docs.cube.dev/api-reference/user-attributes/get-user-attributes
/api-reference/api.yaml get /api/v1/user-attributes
List user attributes, newest first, cursor-paginated via `first`/`after`. The legacy `offset`/`limit` query params and the `data`/`count` response fields are deprecated — use `items` + `pageInfo` instead.
# Update user attribute
Source: https://docs.cube.dev/api-reference/user-attributes/update-user-attribute
/api-reference/api.yaml put /api/v1/user-attributes/{id}
# Get user OAuth token
Source: https://docs.cube.dev/api-reference/user-oauth-tokens/get-user-oauth-token
/api-reference/api.yaml get /api/v1/user-oauth-tokens/{integrationId}
# Initiate OAuth flow
Source: https://docs.cube.dev/api-reference/user-oauth-tokens/initiate-oauth-flow
/api-reference/api.yaml post /api/v1/user-oauth-tokens/{integrationId}/initiate
# List user OAuth tokens
Source: https://docs.cube.dev/api-reference/user-oauth-tokens/list-user-oauth-tokens
/api-reference/api.yaml get /api/v1/user-oauth-tokens
# Revoke user OAuth token
Source: https://docs.cube.dev/api-reference/user-oauth-tokens/revoke-user-oauth-token
/api-reference/api.yaml delete /api/v1/user-oauth-tokens/{integrationId}
# Apply an action to many users
Source: https://docs.cube.dev/api-reference/users-admin/apply-an-action-to-many-users
/api-reference/api.yaml post /api/v1/users/bulk
**🔒 Admin only.** Requires administrator privileges — the authenticated principal (API key, embed JWT, or any bearer token) must belong to a user with the admin role.
Apply one action to up to 100 users in a single request.
Request body:
* `action` — `DEACTIVATE` (revoke access and end the users' sessions), `ACTIVATE` (restore access to deactivated users), or `DELETE` (remove the users outright).
* `userIds` — the users to act on. Repeated ids are collapsed, so the response holds exactly one entry per distinct user requested.
**This endpoint is partially successful.** It returns `200` whenever the request itself is well formed, and reports each user's outcome separately:
* `succeeded` — ids of the users the action was applied to.
* `failed` — the users it was not applied to, each with an `error` carrying the `status` and `message` the single-user endpoint would have returned (`404` for an unknown id, `403` when a guard refuses the change).
An empty `failed` array means the whole batch applied. Users are processed one at a time in the order given, each in its own transaction, so earlier changes are **not** rolled back when a later user fails — the response is the record of what was applied.
Every guard of the single-user endpoints still applies, per user:
* The caller's own id fails rather than the request: `DEACTIVATE` on self is `403`, `DELETE` on self is `400`.
* A tenant must always keep at least one active, direct admin, so the action that would remove the last one fails with `403` while the rest of the batch still applies.
Reapplying an action a user is already in (deactivating a deactivated user) succeeds as a no-op.
# Create user
Source: https://docs.cube.dev/api-reference/users-admin/create-user
/api-reference/api.yaml post /api/v1/users
# Delete user
Source: https://docs.cube.dev/api-reference/users-admin/delete-user
/api-reference/api.yaml delete /api/v1/users/{id}
# Update user
Source: https://docs.cube.dev/api-reference/users-admin/update-user
/api-reference/api.yaml put /api/v1/users/{id}
# Update my settings
Source: https://docs.cube.dev/api-reference/users/update-my-settings
/api-reference/api.yaml patch /api/v1/users/me/settings
Merge a partial patch into the authenticated user's own settings. Omitted fields are left as they are; `sheets` merges field-by-field, and an explicit `null` clears it.
# Clone a workbook
Source: https://docs.cube.dev/api-reference/workbooks/clone-a-workbook
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/workbooks/{workbookId}/duplicate
Create a full copy of a workbook, including its reports and its published dashboard, and return the new workbook.
The clone is named "Copy of \{original name}". All report references inside the dashboard (and the dashboard draft) are re-pointed at the copied reports, so the duplicate is fully self-contained and independent of the original. Per-user view history (last-viewed timestamps) is **not** carried over.
The clone does **not** inherit the source's sharing: no user or group policies are copied, and signed embedding (`allowEmbed`) starts disabled on the cloned dashboard even when it is enabled on the source.
Placement follows the caller's access to the source's folder. With **edit** access to that folder the clone (and its reports) is created alongside the original, and therefore inherits the access that folder grants. With only **read** access — or when the source is at the root — the clone is created at the workspace root instead, so a viewer never writes into a folder they don't control.
Requires **read** access to the source workbook and the **AI BI User** role (or higher). The new workbook is owned by the calling user.
Creator-mode embed sessions can also clone a workbook from the **shared workspace** (the items listed by `GET /shared-workspace`) by passing `{ "shared": true }` in the request body — `workbookId` is then resolved against the shared workspace instead of the caller's own. Such a clone is built from the workbook's **published dashboard**: a fresh report is created per report snapshot, the dashboard is re-published onto the new workbook, and the same config seeds the workbook's draft — the source's live reports and unpublished draft are not copied. The clone is placed at the workspace root. Requires the workbook to have a published dashboard (`400` otherwise); `403` for non-creator-mode callers.
If the deployment's plan enforces a workbook limit and it has been reached, the request is rejected with `400`. Returns `404` if the workbook does not exist or belongs to a different deployment.
# Create or update a workbook by slug
Source: https://docs.cube.dev/api-reference/workbooks/create-or-update-a-workbook-by-slug
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/workbooks/by-slug/{slug}
Idempotent create-or-update ("upsert") of a workbook keyed by its deployment-scoped slug — the portable identity for managing dashboards as code (CI/CD pipelines re-applying the same definition against any environment).
If a workbook with this slug exists in the deployment, it is updated: only the fields you send are changed, and `meta` is merged into the existing metadata (a `meta.dashboardDraft` is validated like everywhere else). Otherwise a new workbook is created with this slug. The slug is normalized (trimmed, lowercased) before matching. Re-applying the same request is a no-op.
Access: updating requires **edit** access to the workbook holding the slug (requests that include `folderId` require **manage** access, plus **edit** access to the destination folder); creating requires the same access as `POST /workbooks` (AI BI User role or higher).
Whether this request creates or updates is decided — and access-checked — before it is applied. If another writer changes that in between (creates or deletes the workbook holding this slug), the request is rejected with `409` and `code: "upsert_branch_changed"` rather than applied under the wrong permission check. This is **transient: retry the request**. It only happens when two writers target the same slug concurrently — for example two pipeline runs applying the same bundle at once, where one creates the workbook and the other must then retry into the update path.
Returns `404` if the deployment does not exist.
# Create workbook
Source: https://docs.cube.dev/api-reference/workbooks/create-workbook
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/workbooks
# Delete workbook
Source: https://docs.cube.dev/api-reference/workbooks/delete-workbook
/api-reference/api.yaml delete /api/v1/deployments/{deploymentId}/workbooks/{workbookId}
# Get workbook
Source: https://docs.cube.dev/api-reference/workbooks/get-workbook
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/workbooks/{workbookId}
# Get workbooks
Source: https://docs.cube.dev/api-reference/workbooks/get-workbooks
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/workbooks
# Publish dashboard
Source: https://docs.cube.dev/api-reference/workbooks/publish-dashboard
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/workbooks/{workbookId}/publish
# Update a workbook
Source: https://docs.cube.dev/api-reference/workbooks/update-a-workbook
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/workbooks/{workbookId}
Update a workbook's properties. All fields are optional; only the fields you send are changed.
* `name` — rename the workbook.
* `meta` — replace the workbook's metadata object.
* `folderId` — move the workbook between folders. Set it to a folder id to move the workbook into that folder, or to `null` to move it back to the workspace root.
Access depends on what you change. Renaming or editing metadata requires **edit** access to the workbook. Moving the workbook (any request that includes `folderId`) requires **manage** access to the workbook, plus **edit** access to the destination folder when moving into one.
To move a workbook you can use either this endpoint or the unified `POST /workspace/move`. Returns `404` if the workbook does not exist or belongs to a different deployment.
# Update published dashboard AI widget thread
Source: https://docs.cube.dev/api-reference/workbooks/update-published-dashboard-ai-widget-thread
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/workbooks/{workbookId}/dashboard/ai-widget-thread
# Update workbook dashboard
Source: https://docs.cube.dev/api-reference/workbooks/update-workbook-dashboard
/api-reference/api.yaml put /api/v1/deployments/{deploymentId}/workbooks/{workbookId}/dashboard
# Delete workspace items in bulk
Source: https://docs.cube.dev/api-reference/workspace/delete-workspace-items-in-bulk
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/workspace/bulk-delete
Delete up to 100 workspace items — workbooks, reports, and folders, mixed freely — in a single request.
Request body:
* `items` — the items to delete, each `{ type, id }` with `type` one of `WORKBOOK`, `REPORT`, `FOLDER`. Repeated references to the same item are collapsed, so the response holds exactly one entry per distinct item requested.
**This endpoint is partially successful.** It returns `200` whenever the request itself is well formed, and reports each item's outcome separately:
* `deleted` — the `(type, id)` of each item that was deleted.
* `failed` — the items that were not deleted, each with the `(type, id)` requested and an `error` carrying the `status` and `message` the per-type delete endpoint would have returned for it (`403` when the caller lacks access, `404` when the item does not exist or belongs to another deployment, `400` when a folder still has sub-folders).
An empty `failed` array means the whole batch applied. Items are deleted one at a time in the order given, each in its own transaction, so **deletions are not rolled back** when a later item fails — the response is the record of what was applied. Deleting a folder does not delete the workbooks, dashboards, and reports directly inside it: they are detached and returned to the workspace root, matching `DELETE /folders/:folderId`.
Access is checked per item with the same rules as the per-type delete endpoints: **manage** access for a workbook or folder, and **edit** access plus the report-manage permission for a report.
# List shared workspace items
Source: https://docs.cube.dev/api-reference/workspace/list-shared-workspace-items
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/shared-workspace
List the workspace items (folders, workbooks, reports) that have been shared with embed users via the built-in "All embed users" group or the caller's per-embed-tenant `system:tenant:{embedTenantName}` group.
Unlike `GET /workspace`, this feed is not per-user: every valid embed caller from the same embed tenant sees the same set, scoped to what those groups have been granted. Owner identities are blanked out so they are not exposed to embed users.
Accepts the same `folderId`, `types`, `search`, sorting, and cursor-pagination parameters as `GET /workspace`. Folders reachable only because a nested item inside them was shared are surfaced as navigation, but their unshared siblings stay hidden.
# List workspace items
Source: https://docs.cube.dev/api-reference/workspace/list-workspace-items
/api-reference/api.yaml get /api/v1/deployments/{deploymentId}/workspace
List the items in a deployment's workspace — folders, workbooks, and reports — in a single unified, paginated feed.
By default this returns the items at the workspace root. Pass `folderId` to list the contents of a specific folder instead. Folders are always returned ahead of other item types so they render at the top of a listing.
Supported query parameters:
* `folderId` — list the contents of this folder (omit for the root).
* `types` — restrict the results to one or more item types (`FOLDER`, `WORKBOOK`, `REPORT`). Repeat the parameter to pass several.
* `search` — case-insensitive substring match on item names; searches across all folders, not just the current level.
* `orderByField` / `orderByDirection` — sort the non-folder items (e.g. by `name`, `updated_at`, `created_at`, or `viewer_last_viewed_at`).
* `first` / `after` — cursor-based pagination.
Results are scoped to the calling user: only items the user can access are returned.
# Move a workspace item
Source: https://docs.cube.dev/api-reference/workspace/move-a-workspace-item
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/workspace/move
Move any workspace item — a workbook, report, or folder — into a folder, in one unified endpoint. Returns the moved item.
Request body:
* `type` — the item type: `WORKBOOK`, `REPORT`, or `FOLDER`.
* `id` — the id of the item to move.
* `folderId` — the destination folder id, or `null` to move the item to the workspace root. For a folder, this becomes its new parent.
Access requirements depend on the item type and mirror the per-type rules:
* **Workbook** — manage access to the workbook, plus edit access to the destination folder when moving into one.
* **Report** — edit access to the report (with the report-manage permission).
* **Folder** — manage access to the folder being moved, plus manage or edit access to the destination parent folder.
Moving a folder into one of its own descendants, or exceeding the maximum folder depth, is rejected with `400`. Returns `404` if the item does not exist or belongs to a different deployment.
# Move workspace items in bulk
Source: https://docs.cube.dev/api-reference/workspace/move-workspace-items-in-bulk
/api-reference/api.yaml post /api/v1/deployments/{deploymentId}/workspace/bulk-move
Move up to 100 workspace items — workbooks, reports, and folders, mixed freely — into one destination folder in a single request.
Request body:
* `items` — the items to move, each `{ type, id }` with `type` one of `WORKBOOK`, `REPORT`, `FOLDER`. Repeated references to the same item are collapsed, so the response holds exactly one entry per distinct item requested.
* `folderId` — the destination folder id for the whole batch, or `null` to move the items to the workspace root. For a folder item this becomes its new parent.
**This endpoint is partially successful.** It returns `200` whenever the request itself is well formed, and reports each item's outcome separately:
* `moved` — the items that were moved, serialized exactly as `POST /workspace/move` returns them.
* `failed` — the items that were not moved, each with the `(type, id)` requested and an `error` carrying the `status` and `message` the single-item move would have returned for it (`403` when the caller lacks access, `404` when the item does not exist or belongs to another deployment, `400` for an invalid move such as a folder into its own descendant).
An empty `failed` array means the whole batch applied. Items are applied one at a time in the order given, each in its own transaction, so earlier moves are **not** rolled back when a later item fails — the response is the record of what was applied.
Access is checked per item with the same rules as the single-item move, so a batch mixing items the caller may and may not move applies the permitted ones and reports the rest.
# Architecture
Source: https://docs.cube.dev/cube-core/architecture
Learn about Cube Core's production deployment architecture, including API instances, refresh workers, and Cube Store.
A typical production deployment of Cube Core includes several components
working together. This page describes the architecture and how each
component fits in.
## Components
As shown in the diagram below, a typical production deployment of Cube includes
the following components:
* One or multiple API instances
* A Refresh Worker
* A Cube Store cluster
**API Instances** process incoming API requests and query either Cube Store for
pre-aggregated data or connected database(s) for raw data. The **Refresh
Worker** builds and refreshes pre-aggregations in the background. **Cube Store**
ingests pre-aggregations built by Refresh Worker and responds to queries from
API instances.
API instances and Refresh Workers can be configured via [environment
variables][ref-config-env] or the [`cube.js` configuration file][ref-config-js].
They also need access to the data model files. Cube Store clusters can be
configured via environment variables.
You can find an example Docker Compose configuration for a Cube deployment in
the platform-specific guide for \[Docker]\[ref-deploy-docker].
## API instances
API instances process incoming API requests and query either Cube Store for
pre-aggregated data or connected data sources for raw data. It is possible to
horizontally scale API instances and use a load balancer to balance incoming
requests between multiple API instances.
The [Cube Docker image][dh-cubejs] is used for API Instance.
API instances can be configured via environment variables or the `cube.js`
configuration file, and **must** have access to the data model files (as
specified by [`schema_path`][ref-conf-ref-schemapath].
## Refresh Worker
A Refresh Worker updates pre-aggregations and invalidates the in-memory cache in
the background. They also keep the refresh keys up-to-date for all data models
and pre-aggregations. Please note that the in-memory cache is just invalidated
but not populated by Refresh Worker. In-memory cache is populated lazily during
querying. On the other hand, pre-aggregations are eagerly populated and kept
up-to-date by Refresh Worker.
The [Cube Docker image][dh-cubejs] can be used for creating Refresh Workers; to
make the service act as a Refresh Worker, `CUBEJS_REFRESH_WORKER=true` should be
set in the environment variables.
## Cube Store
Cube Store is the purpose-built pre-aggregations storage for Cube.
Cube Store uses a distributed query engine architecture. In every Cube Store
cluster:
* a one or many [router nodes](#cube-store-router) handle incoming
connections, manages database metadata, builds query plans, and orchestrates
their execution
* multiple [worker nodes](#cube-store-worker) ingest warmed up data
and execute queries in parallel
* a local or cloud-based blob storage keeps pre-aggregated data in columnar
format
By default, Cube Store listens on the port `3030` for queries coming from Cube.
The port could be changed by setting `CUBESTORE_HTTP_PORT` environment variable.
In a case of using custom port, please make sure to change
[`CUBEJS_CUBESTORE_PORT`](/reference/configuration/environment-variables#cubejs_cubestore_port) environment variable for Cube API Instances and Refresh
Worker.
Both the router and worker use the [Cube Store Docker image][dh-cubestore]. The
following environment variables should be used to manage the roles:
| Environment Variable | Specify on Router? | Specify on Worker? |
| ----------------------- | ------------------ | ------------------ |
| `CUBESTORE_SERVER_NAME` | ✅ Yes | ✅ Yes |
| `CUBESTORE_META_PORT` | ✅ Yes | — |
| `CUBESTORE_WORKERS` | ✅ Yes | ✅ Yes |
| `CUBESTORE_WORKER_PORT` | — | ✅ Yes |
| `CUBESTORE_META_ADDR` | — | ✅ Yes |
Looking for a deeper dive on Cube Store architecture? Check out
[this presentation](https://docs.google.com/presentation/d/1oQ-koloag0UcL-bUHOpBXK4txpqiGl41rxhgDVrw7gw/)
by our CTO, [Pavel][gh-pavel].
### Cube Store Router
The Router in a Cube Store cluster is responsible for receiving queries from
Cube, managing metadata for the Cube Store cluster, and query planning and
distribution for the Workers. It also [provides a MySQL-compatible
interface][ref-caching-inspect-sql] that can be used to query pre-aggregations
from Cube Store directly. Cube **only** communicates with the Router, and does
not interact with Workers directly.
[ref-caching-inspect-sql]: /docs/pre-aggregations/using-pre-aggregations#inspecting-pre-aggregations
### Cube Store Worker
Workers in a Cube Store cluster receive and execute subqueries from the Router,
and directly interact with the underlying distributed storage for insertions,
selections and pre-aggregation warmup. Workers **do not** interact with each
other directly, and instead rely on the Router to distribute queries and manage
any associated metadata.
### Scaling
Although Cube Store *can* be run in single-instance mode, this is often
unsuitable for production deployments. For high concurrency and data throughput,
we **strongly** recommend running Cube Store as a cluster of multiple instances
instead. Because the storage layer is decoupled from the query processing
engine, you can horizontally scale your Cube Store cluster for as much
concurrency as you require.
A sample Docker Compose stack setting Cube Store cluster up might look like:
```yaml theme={"dark"}
services:
cubestore_router:
image: cubejs/cubestore:latest
environment:
- CUBESTORE_WORKERS=cubestore_worker_1:10001,cubestore_worker_2:10002
- CUBESTORE_REMOTE_DIR=/cube/data
- CUBESTORE_META_PORT=9999
- CUBESTORE_SERVER_NAME=cubestore_router:9999
volumes:
- .cubestore:/cube/data
depends_on:
- cubestore_worker_1
- cubestore_worker_2
cubestore_worker_1:
image: cubejs/cubestore:latest
environment:
- CUBESTORE_WORKERS=cubestore_worker_1:10001,cubestore_worker_2:10002
- CUBESTORE_SERVER_NAME=cubestore_worker_1:10001
- CUBESTORE_WORKER_PORT=10001
- CUBESTORE_REMOTE_DIR=/cube/data
- CUBESTORE_META_ADDR=cubestore_router:9999
volumes:
- .cubestore:/cube/data
cubestore_worker_2:
image: cubejs/cubestore:latest
environment:
- CUBESTORE_WORKERS=cubestore_worker_1:10001,cubestore_worker_2:10002
- CUBESTORE_SERVER_NAME=cubestore_worker_2:10002
- CUBESTORE_WORKER_PORT=10002
- CUBESTORE_REMOTE_DIR=/cube/data
- CUBESTORE_META_ADDR=cubestore_router:9999
volumes:
- .cubestore:/cube/data
```
### Storage
Cube Store makes use of a separate storage layer for storing metadata as well as
for persisting pre-aggregations as Parquet files. Cube Store can use both AWS S3
and Google Cloud, or if desired, a local path on the server if all nodes of a
cluster run on a single machine.
A simplified example using AWS S3 might look like:
```yaml theme={"dark"}
services:
cubestore_router:
image: cubejs/cubestore:latest
environment:
- CUBESTORE_SERVER_NAME=cubestore_router:9999
- CUBESTORE_META_PORT=9999
- CUBESTORE_WORKERS=cubestore_worker_1:9001
- CUBESTORE_S3_BUCKET=
- CUBESTORE_S3_REGION=
- CUBESTORE_AWS_ACCESS_KEY_ID=
- CUBESTORE_AWS_SECRET_ACCESS_KEY=
cubestore_worker_1:
image: cubejs/cubestore:latest
environment:
- CUBESTORE_SERVER_NAME=cubestore_worker_1:9001
- CUBESTORE_WORKER_PORT=9001
- CUBESTORE_META_ADDR=cubestore_router:9999
- CUBESTORE_WORKERS=cubestore_worker_1:9001
- CUBESTORE_S3_BUCKET=
- CUBESTORE_S3_REGION=
- CUBESTORE_AWS_ACCESS_KEY_ID=
- CUBESTORE_AWS_SECRET_ACCESS_KEY=
depends_on:
- cubestore_router
```
Instead of static access keys, Cube Store can authenticate to S3 using AWS Web
Identity (for example, IRSA on Amazon EKS). Leave `CUBESTORE_AWS_ACCESS_KEY_ID`
unset and provide
[`CUBESTORE_AWS_WEB_IDENTITY_TOKEN_FILE`][ref-config-env-cubestore-web-identity]
and [`CUBESTORE_AWS_ROLE_ARN`][ref-config-env]; Cube Store then assumes the role
via STS and refreshes credentials automatically when the token file changes.
[dh-cubejs]: https://hub.docker.com/r/cubejs/cube
[dh-cubestore]: https://hub.docker.com/r/cubejs/cubestore
[ref-config-env]: /reference/configuration/environment-variables
[ref-config-env-cubestore-web-identity]: /reference/configuration/environment-variables#cubestore_aws_web_identity_token_file
[ref-config-js]: /reference/configuration/config
[ref-conf-ref-schemapath]: /reference/configuration/config#schema_path
[gh-pavel]: https://github.com/paveltiunov
# Deploying Cube Core with Docker
Source: https://docs.cube.dev/cube-core/deployment
This guide walks you through deploying Cube with Docker.
This is an example of a production-ready deployment, but real-world deployments
can vary significantly depending on desired performance and scale.
If you'd like to deploy Cube to [Kubernetes](https://kubernetes.io), please
refer to the following resources with Helm charts:
[`gadsme/charts`](https://github.com/gadsme/charts) or
[`OpstimizeIcarus/cubejs-helm-charts-kubernetes`](https://github.com/OpstimizeIcarus/cubejs-helm-charts-kubernetes/tree/release-v0.1).
These resources are community-maintained, and they are not maintained by the
Cube team. Please direct questions related to these resources to their authors.
## Prerequisites
* [Docker Desktop][link-docker-app]
## Configuration
Create a Docker Compose stack by creating a `docker-compose.yml`. A
production-ready stack would at minimum consist of:
* One or more Cube API instance
* A Cube Refresh Worker
* A Cube Store Router node
* One or more Cube Store Worker nodes
An example stack using BigQuery as a data source is provided below:
**Using macOS or Windows?** Use `CUBEJS_DB_HOST=host.docker.internal` instead of
`localhost` if your database is on the same machine.
**Using macOS on Apple Silicon (arm64)?** Use the `arm64v8` tag for Cube Store
[Docker images](https://hub.docker.com/r/cubejs/cubestore/tags?page=\&page_size=\&ordering=\&name=arm64v8),
e.g., `cubejs/cubestore:arm64v8`.
Note that it's a best practice to use specific locked versions, e.g.,
`cubejs/cube:v0.36.0`, instead of `cubejs/cube:latest` in production.
```yaml theme={"dark"}
services:
cube_api:
restart: always
image: cubejs/cube:latest
ports:
- 4000:4000
environment:
- CUBEJS_DB_TYPE=bigquery
- CUBEJS_DB_BQ_PROJECT_ID=cube-bq-cluster
- CUBEJS_DB_BQ_CREDENTIALS=
- CUBEJS_DB_EXPORT_BUCKET=cubestore
- CUBEJS_CUBESTORE_HOST=cubestore_router
- CUBEJS_API_SECRET=secret
volumes:
- .:/cube/conf
depends_on:
- cube_refresh_worker
- cubestore_router
- cubestore_worker_1
- cubestore_worker_2
cube_refresh_worker:
restart: always
image: cubejs/cube:latest
environment:
- CUBEJS_DB_TYPE=bigquery
- CUBEJS_DB_BQ_PROJECT_ID=cube-bq-cluster
- CUBEJS_DB_BQ_CREDENTIALS=
- CUBEJS_DB_EXPORT_BUCKET=cubestore
- CUBEJS_CUBESTORE_HOST=cubestore_router
- CUBEJS_API_SECRET=secret
- CUBEJS_REFRESH_WORKER=true
volumes:
- .:/cube/conf
cubestore_router:
restart: always
image: cubejs/cubestore:latest
environment:
- CUBESTORE_WORKERS=cubestore_worker_1:10001,cubestore_worker_2:10002
- CUBESTORE_REMOTE_DIR=/cube/data
- CUBESTORE_META_PORT=9999
- CUBESTORE_SERVER_NAME=cubestore_router:9999
volumes:
- .cubestore:/cube/data
cubestore_worker_1:
restart: always
image: cubejs/cubestore:latest
environment:
- CUBESTORE_WORKERS=cubestore_worker_1:10001,cubestore_worker_2:10002
- CUBESTORE_SERVER_NAME=cubestore_worker_1:10001
- CUBESTORE_WORKER_PORT=10001
- CUBESTORE_REMOTE_DIR=/cube/data
- CUBESTORE_META_ADDR=cubestore_router:9999
volumes:
- .cubestore:/cube/data
depends_on:
- cubestore_router
cubestore_worker_2:
restart: always
image: cubejs/cubestore:latest
environment:
- CUBESTORE_WORKERS=cubestore_worker_1:10001,cubestore_worker_2:10002
- CUBESTORE_SERVER_NAME=cubestore_worker_2:10002
- CUBESTORE_WORKER_PORT=10002
- CUBESTORE_REMOTE_DIR=/cube/data
- CUBESTORE_META_ADDR=cubestore_router:9999
volumes:
- .cubestore:/cube/data
depends_on:
- cubestore_router
```
## Set up reverse proxy
In production, the Cube API should be served over an HTTPS connection to ensure
security of the data in-transit. We recommend using a reverse proxy; as an
example, let's use [NGINX][link-nginx].
You can also use a reverse proxy to enable HTTP 2.0 and GZIP compression
First we'll create a new server configuration file called `nginx/cube.conf`:
```nginx theme={"dark"}
server {
listen 443 ssl;
server_name cube.my-domain.com;
ssl_protocols TLSv1 TLSv1.1 TLSv1.2;
ssl_ecdh_curve secp384r1;
# Replace the ciphers with the appropriate values
ssl_ciphers "ECDHE-RSA-AES256-GCM-SHA512:DHE-RSA-AES256-GCM-SHA512:ECDHE-RSA-AES256-GCM-SHA384:DHE-RSA-AES256-GCM-SHA384:ECDHE-RSA-AES256-SHA384 OLD_TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305_SHA256 OLD_TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305_SHA256";
ssl_prefer_server_ciphers on;
ssl_certificate /etc/ssl/private/cert.pem;
ssl_certificate_key /etc/ssl/private/key.pem;
ssl_session_timeout 10m;
ssl_session_cache shared:SSL:10m;
ssl_session_tickets off;
ssl_stapling on;
ssl_stapling_verify on;
location / {
proxy_pass http://cube:4000/;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
}
}
```
Then we'll add a new service to our Docker Compose stack:
```yaml theme={"dark"}
services:
...
nginx:
image: nginx
ports:
- 443:443
volumes:
- ./nginx:/etc/nginx/conf.d
- ./ssl:/etc/ssl/private
```
Don't forget to create a `ssl` directory with the `cert.pem` and `key.pem` files
inside so the Nginx service can find them.
For automatically provisioning SSL certificates with LetsEncrypt, [this blog
post][medium-letsencrypt-nginx] may be useful.
## Security
### Use JSON Web Tokens
Cube can be configured to use industry-standard JSON Web Key Sets for securing
its API and limiting access to data. To do this, we'll define the relevant
options on our Cube API instance:
If you're using [`queryRewrite`][ref-config-queryrewrite] for access control,
then you must also configure
[`scheduledRefreshContexts`][ref-config-sched-ref-ctx] so the refresh workers
can correctly create pre-aggregations.
```yaml theme={"dark"}
services:
cube_api:
image: cubejs/cube:latest
ports:
- 4000:4000
environment:
- CUBEJS_DB_TYPE=bigquery
- CUBEJS_DB_BQ_PROJECT_ID=cube-bq-cluster
- CUBEJS_DB_BQ_CREDENTIALS=
- CUBEJS_DB_EXPORT_BUCKET=cubestore
- CUBEJS_CUBESTORE_HOST=cubestore_router
- CUBEJS_API_SECRET=secret
- CUBEJS_JWK_URL=https://cognito-idp..amazonaws.com//.well-known/jwks.json
- CUBEJS_JWT_AUDIENCE=
- CUBEJS_JWT_ISSUER=https://cognito-idp..amazonaws.com/
- CUBEJS_JWT_ALGS=RS256
- CUBEJS_JWT_CLAIMS_NAMESPACE=
volumes:
- .:/cube/conf
depends_on:
- cubestore_worker_1
- cubestore_worker_2
- cube_refresh_worker
```
### Securing Cube Store
All Cube Store nodes (both router and workers) should only be accessible to Cube
API instances and refresh workers. To do this with Docker Compose, we simply
need to make sure that none of the Cube Store services have any exposed
## Monitoring
All Cube logs can be found by through the Docker Compose CLI:
```bash theme={"dark"}
docker-compose ps
Name Command State Ports
---------------------------------------------------------------------------------------------------------------------------------
cluster_cube_1 docker-entrypoint.sh cubej ... Up 0.0.0.0:4000->4000/tcp,:::4000->4000/tcp
cluster_cubestore_router_1 ./cubestored Up 3030/tcp, 3306/tcp
cluster_cubestore_worker_1_1 ./cubestored Up 3306/tcp, 9001/tcp
cluster_cubestore_worker_2_1 ./cubestored Up 3306/tcp, 9001/tcp
docker-compose logs
cubestore_router_1 | 2021-06-02 15:03:20,915 INFO [cubestore::metastore] Creating metastore from scratch in /cube/.cubestore/data/metastore
cubestore_router_1 | 2021-06-02 15:03:20,950 INFO [cubestore::cluster] Meta store port open on 0.0.0.0:9999
cubestore_router_1 | 2021-06-02 15:03:20,951 INFO [cubestore::mysql] MySQL port open on 0.0.0.0:3306
cubestore_router_1 | 2021-06-02 15:03:20,952 INFO [cubestore::http] Http Server is listening on 0.0.0.0:3030
cube_1 | 🚀 Cube API server (vX.XX.XX) is listening on 4000
cubestore_worker_2_1 | 2021-06-02 15:03:24,945 INFO [cubestore::cluster] Worker port open on 0.0.0.0:9001
cubestore_worker_1_1 | 2021-06-02 15:03:24,830 INFO [cubestore::cluster] Worker port open on 0.0.0.0:9001
```
## Update to the latest version
Find the latest stable release version [from
Docker Hub][link-cubejs-docker]. Then update your `docker-compose.yml` to use
a specific tag instead of `latest`:
```yaml theme={"dark"}
services:
cube_api:
image: cubejs/cube:v0.36.0
ports:
- 4000:4000
environment:
- CUBEJS_DB_TYPE=bigquery
- CUBEJS_DB_BQ_PROJECT_ID=cube-bq-cluster
- CUBEJS_DB_BQ_CREDENTIALS=
- CUBEJS_DB_EXPORT_BUCKET=cubestore
- CUBEJS_CUBESTORE_HOST=cubestore_router
- CUBEJS_API_SECRET=secret
volumes:
- .:/cube/conf
depends_on:
- cubestore_router
- cube_refresh_worker
```
## Extend the Docker image
If you need to use dependencies (i.e., Python or npm packages) with native
extensions inside [configuration files][ref-config-files] or [dynamic data
models][ref-dynamic-data-models], build a custom Docker image.
You can do this by creating a `Dockerfile` and a corresponding
`.dockerignore` file:
```bash theme={"dark"}
touch Dockerfile
touch .dockerignore
```
Add this to the `Dockerfile`:
```dockerfile theme={"dark"}
FROM cubejs/cube:latest
COPY . .
RUN apt update && apt install -y pip
RUN pip install -r requirements.txt
RUN npm install
```
And this to the `.dockerignore`:
```gitignore theme={"dark"}
model
cube.py
cube.js
.env
node_modules
npm-debug.log
```
Then start the build process by running the following command:
```bash theme={"dark"}
docker build -t /cube-custom-image .
```
Finally, update your `docker-compose.yml` to use your newly-built image:
```yaml theme={"dark"}
services:
cube_api:
image: /cube-custom-image
ports:
- 4000:4000
environment:
- CUBEJS_API_SECRET=secret
# Other environment variables
volumes:
- .:/cube/conf
depends_on:
- cubestore_router
- cube_refresh_worker
# Other container dependencies
```
Note that you shoudn't mount the whole current folder (`.:/cube/conf`)
if you have dependencies in `package.json`. Doing so would effectively
hide the `node_modules` folder inside the container, where dependency files
installed with `npm install` reside, and result in errors like this:
`Error: Cannot find module 'my_dependency'`. In that case, mount individual files:
```yaml theme={"dark"}
# ...
volumes:
- ./model:/cube/conf/model
- ./cube.js:/cube/conf/cube.js
# Other necessary files
```
## Production checklist
Thinking of migrating to the cloud instead? [Click
here][blog-migrate-to-cube-cloud] to learn more about migrating a self-hosted
installation to [Cube Cloud][link-cube-cloud].
This is a checklist for configuring and securing Cube for a production
deployment.
### Disable Development Mode
When running Cube in production environments, make sure development mode is
disabled both on API Instances and Refresh Worker. You can read more about what
development mode changes [here][link-cubejs-dev-vs-prod].
**Development mode is an authentication bypass.** Cube is in development mode when
`CUBEJS_DEV_MODE=true`, and also whenever `NODE_ENV` is not `production`. When Cube
is started through the `cubejs` CLI — which is what the official Docker images run —
setting `CUBEJS_DEV_MODE=true` additionally forces `NODE_ENV=development`, which
switches off JWT verification on the
[REST (JSON)](/reference/core-data-apis/rest-api) and
[GraphQL](/reference/core-data-apis/graphql-api) APIs: they then accept requests with
no token at all.
Development mode also mounts [Playground](/docs/explore-analyze/playground) and its
supporting endpoints with no authentication whatsoever. Anyone who can reach the
instance is handed a ready-to-use API token, and can mint further ones carrying any
security context signed with your API secret — and so query every data API as any user,
bypassing [member-level access
control](/docs/data-modeling/access-control/member-level-security) and [row-level
security](/docs/data-modeling/access-control/row-level-security). The same endpoints
read your data model files and the table schema of every connected data source, and
overwrite your data model and your `.env`. With `CUBEJS_DEV_MODE=true` and no
[`CUBEJS_SQL_PASSWORD`](/reference/configuration/environment-variables#cubejs_sql_password)
set, the SQL API accepts any credentials as well, allowing arbitrary SQL against
connected data sources.
This is intentional. Development mode is designed to run on a developer's local
machine for ease of use and debugging. Never run it where anyone else can reach it,
never expose it to the internet, and never use it in production. Using development
mode in the Cube cloud platform is highly discouraged — it bypasses the platform's
security model.
To keep it off: `cubejs server` and the official Docker images already set
`NODE_ENV=production`, so leaving `CUBEJS_DEV_MODE` unset — its default — is enough
there. If you embed `@cubejs-backend/server-core` directly rather than starting Cube
through the `cubejs` CLI, set `NODE_ENV=production` yourself, since an unset `NODE_ENV`
puts the instance in development mode whatever the flag says.
`CUBEJS_DEV_MODE` defaults to `false`, and the official Docker images already set
`NODE_ENV=production` — as does `cubejs server` itself — so on a Docker deployment
leaving the flag unset is enough. Setting `NODE_ENV` below as well is belt-and-braces;
it matters only if you embed `@cubejs-backend/server-core` directly, where an unset
`NODE_ENV` puts the instance in development mode whatever the flag says.
```dotenv theme={"dark"}
# Keep development mode off. Under the `cubejs` CLI and the official images,
# setting this to `true` also forces NODE_ENV=development, switching off JWT
# verification on the data APIs
CUBEJS_DEV_MODE=false
# Already set by the official images and by `cubejs server`; set it explicitly if
# you start Cube some other way
NODE_ENV=production
```
### Set up Refresh Worker
To refresh in-memory cache and [pre-aggregations][ref-schema-ref-preaggs] in the
background, we recommend running a separate Cube Refresh Worker instance. This
allows your Cube API Instance to continue to serve requests with high
availability.
```dotenv theme={"dark"}
# Set to true so a Cube instance acts as a refresh worker
CUBEJS_REFRESH_WORKER=true
```
### Set up Cube Store
While Cube can operate with in-memory cache and queue storage, there're multiple
parts of Cube which require Cube Store in production mode. Replicating Cube
instances without Cube Store can lead to source database degraded performance,
various race conditions and cached data inconsistencies.
Cube Store manages in-memory cache, queue and pre-aggregations for Cube. Follow
the [instructions here][ref-caching-cubestore] to set it up.
Depending on your database, Cube may need to "stage" pre-aggregations inside
your database first before ingesting them into Cube Store. In this case, Cube
will require write access to a dedicated schema inside your database.
The schema name is `prod_pre_aggregations` by default. It can be set using the
[`pre_aggregations_schema` configration option][ref-conf-preaggs-schema].
You may consider enabling an export bucket which allows Cube to build large
pre-aggregations in a much faster manner. It is currently supported for
BigQuery, Redshift, Snowflake, and some other data sources. Check [the relevant
documentation for your configured database][ref-config-connect-db] to set it up.
### Secure the deployment
If you're using JWTs, you can configure Cube to correctly decode them and inject
their contents into the [Security Context][ref-sec-ctx]. Add your authentication
provider's configuration under [the `jwt` property of your `cube.js`
configuration file][ref-config-jwt], or if using environment variables, see
`CUBEJS_JWK_*`, `CUBEJS_JWT_*` in the [Environment Variables
reference][ref-env-vars].
### Set up health checks
Cube provides [Kubernetes-API compatible][link-k8s-healthcheck-api] health check
(or probe) endpoints that indicate the status of the deployment. Configure your
monitoring service of choice to use the [`/readyz`][ref-api-readyz] and
[`/livez`][ref-api-livez] API endpoints so you can check on the Cube
deployment's health and be alerted to any issues.
### Appropriate cluster sizing
There's no one-size-fits-all when it comes to sizing a Cube cluster and its
resources. Resources required by Cube significantly depend on the amount of
traffic Cube needs to serve and the amount of data it needs to process. The
following sizing estimates are based on default settings and are very generic,
which may not fit your Cube use case, so you should always tweak resources based
on consumption patterns you see.
#### Memory and CPU
Each Cube cluster should contain at least 2 Cube API instances. Every Cube API
instance should have at least 3GB of RAM and 2 CPU cores allocated for it.
Refresh workers tend to be much more CPU and memory intensive, so at least 6GB
of RAM is recommended. Please note that to take advantage of all available RAM,
the Node.js heap size should be adjusted accordingly by using the
[`--max-old-space-size` option][node-heap-size]:
```sh theme={"dark"}
NODE_OPTIONS="--max-old-space-size=6144"
```
[node-heap-size]: https://nodejs.org/api/cli.html#--max-old-space-sizesize-in-megabytes
The Cube Store router node should have at least 6GB of RAM and 4 CPU cores
allocated for it. Every Cube Store worker node should have at least 8GB of RAM
and 4 CPU cores allocated for it. The Cube Store cluster should have at least
two worker nodes.
#### RPS and data volume
Depending on data model size, every Core Cube API instance can serve 1 to 10
requests per second. Every Core Cube Store router node can serve 50-100 queries
per second. As a rule of thumb, you should provision 1 Cube Store worker node
per one Cube Store partition or 1M of rows scanned in a query. For example if
your queries scan 16M of rows per query, you should have at least 16 Cube Store
worker nodes provisioned. Please note that the number of raw data rows doesn't
usually equal the number of rows in pre-aggregation. At the same time, queries
don't usually scan all the data in pre-aggregations, as Cube Store uses
partition pruning to optimize queries. `EXPLAIN ANALYZE` can be used to see
scanned partitions involved in a Cube Store query. Cube Cloud ballpark
performance numbers can differ as it has different Cube runtime.
### Optimize usage
See [this recipe][ref-data-store-cost-saving-guide] to learn how to optimize
data source usage.
[blog-migrate-to-cube-cloud]: https://cube.dev/blog/migrating-from-self-hosted-to-cube-cloud/
[link-cube-cloud]: https://cubecloud.dev
[link-cubejs-dev-vs-prod]: /reference/configuration/environment-variables#cubejs_dev_mode
[link-k8s-healthcheck-api]: https://kubernetes.io/docs/reference/using-api/health-checks/
[ref-config-connect-db]: /admin/connect-to-data/data-sources
[ref-caching-cubestore]: /docs/pre-aggregations/running-in-production
[ref-conf-preaggs-schema]: /reference/configuration/config#pre_aggregations_schema
[ref-env-vars]: /reference/configuration/environment-variables
[ref-schema-ref-preaggs]: /reference/data-modeling/pre-aggregations
[ref-sec-ctx]: /docs/data-modeling/access-control/context
[ref-config-jwt]: /reference/configuration/config#jwt
[ref-api-readyz]: /reference/core-data-apis/rest-api/reference#readyz
[ref-api-livez]: /reference/core-data-apis/rest-api/reference#livez
[ref-data-store-cost-saving-guide]: /recipes/configuration/data-store-cost-saving-guide
[medium-letsencrypt-nginx]: https://pentacent.medium.com/nginx-and-lets-encrypt-with-docker-in-less-than-5-minutes-b4b8a60d3a71
[link-cubejs-docker]: https://hub.docker.com/r/cubejs/cube
[link-docker-app]: https://www.docker.com/products/docker-app
[link-nginx]: https://www.nginx.com/
[ref-config-files]: /admin/connect-to-data#cubepy-and-cubejs-files
[ref-dynamic-data-models]: /docs/data-modeling/dynamic
[ref-config-queryrewrite]: /reference/configuration/config#queryrewrite
[ref-config-sched-ref-ctx]: /reference/configuration/config#scheduledrefreshcontexts
# Add a pre-aggregation
Source: https://docs.cube.dev/cube-core/getting-started/add-a-pre-aggregation
Hands-on walkthrough of accepting a Playground pre-aggregation suggestion, updating the data model, and rerunning an accelerated query.
In this step, we'll add a pre-aggregation to optimize the performance of a
specific query. Pre-aggregations are a caching technique that massively reduces
query time from seconds to milliseconds. They are extremely useful for speeding
up queries that are run frequently.
From the **Build** tab, execute a query:
Just above the results, click on **Query was not accelerated with
pre-aggregation** to bring up the pre-aggregation suggestion:
A pre-aggregation will be automatically suggested for the query;
click **Add to the Data Model** and then retry the query in the
Playground. This time, the query should be accelerated with a pre-aggregation:
And with that, we conclude our Getting Started with Cube guide. If you'd like to
learn more about Cube, [check out this page][next].
[next]: /cube-core/getting-started/learn-more
# Create a project
Source: https://docs.cube.dev/cube-core/getting-started/create-a-project
In this step, we will create a Cube Core project on your computer, connect a data source, and generate data models.
## Scaffold a project
Start by opening your terminal to create a new folder for the project, then
create a `docker-compose.yml` file within it:
```bash theme={"dark"}
mkdir my-first-cube-project
cd my-first-cube-project
touch docker-compose.yml
```
Open the `docker-compose.yml` file and add the following content:
```yaml theme={"dark"}
services:
cube:
image: cubejs/cube:latest
ports:
- 4000:4000
- 15432:15432
environment:
- CUBEJS_DEV_MODE=true
volumes:
- .:/cube/conf
```
Note that we're setting the [`CUBEJS_DEV_MODE`](/reference/configuration/environment-variables#cubejs_dev_mode) environment variable to `true` to
enable [development mode](/reference/configuration/environment-variables#cubejs_dev_mode). This is
handy for local development but not suitable for
[production](/cube-core/running-in-production).
**Development mode is an authentication bypass.** Cube is in development mode when
`CUBEJS_DEV_MODE=true`, and also whenever `NODE_ENV` is not `production`. When Cube
is started through the `cubejs` CLI — which is what the official Docker images run —
setting `CUBEJS_DEV_MODE=true` additionally forces `NODE_ENV=development`, which
switches off JWT verification on the
[REST (JSON)](/reference/core-data-apis/rest-api) and
[GraphQL](/reference/core-data-apis/graphql-api) APIs: they then accept requests with
no token at all.
Development mode also mounts [Playground](/docs/explore-analyze/playground) and its
supporting endpoints with no authentication whatsoever. Anyone who can reach the
instance is handed a ready-to-use API token, and can mint further ones carrying any
security context signed with your API secret — and so query every data API as any user,
bypassing [member-level access
control](/docs/data-modeling/access-control/member-level-security) and [row-level
security](/docs/data-modeling/access-control/row-level-security). The same endpoints
read your data model files and the table schema of every connected data source, and
overwrite your data model and your `.env`. With `CUBEJS_DEV_MODE=true` and no
[`CUBEJS_SQL_PASSWORD`](/reference/configuration/environment-variables#cubejs_sql_password)
set, the SQL API accepts any credentials as well, allowing arbitrary SQL against
connected data sources.
This is intentional. Development mode is designed to run on a developer's local
machine for ease of use and debugging. Never run it where anyone else can reach it,
never expose it to the internet, and never use it in production. Using development
mode in the Cube cloud platform is highly discouraged — it bypasses the platform's
security model.
To keep it off: `cubejs server` and the official Docker images already set
`NODE_ENV=production`, so leaving `CUBEJS_DEV_MODE` unset — its default — is enough
there. If you embed `@cubejs-backend/server-core` directly rather than starting Cube
through the `cubejs` CLI, set `NODE_ENV=production` yourself, since an unset `NODE_ENV`
puts the instance in development mode whatever the flag says.
If you're using Linux as the Docker host OS, you'll also need to add
`network_mode: 'host'` to your `docker-compose.yml`.
## Start the development server
From the newly-created project directory, run the following command to start
Cube:
```bash theme={"dark"}
docker compose up -d
```
Using Windows? Remember to use [PowerShell][powershell-docs] or
[WSL2][wsl2-docs] to run the command below.
## Connect a data source
Head to [http://localhost:4000](http://localhost:4000) to open the [Developer
Playground][ref-devtools-playground].
The Playground has a database connection wizard that loads when Cube is first
started up and no `.env` file is found. After database credentials have been set
up, an `.env` file will automatically be created and populated with credentials.
Want to use a sample database instead? Select **PostgreSQL** and use the
credentials below:
| Field | Value |
| -------- | ------------------ |
| Host | `demo-db.cube.dev` |
| Port | `5432` |
| Database | `ecom` |
| Username | `cube` |
| Password | `12345` |
After selecting the data source, enter valid credentials for it and
click **Apply**. Check the [Connecting to Databases][ref-conf-db] page
for more details on specific data sources.
You should see tables available to you from the configured database; select the
`orders` table. After selecting the table, click **Generate Data Model**
and pick either **YAML** (recommended) or **JavaScript** format:
Finally, click **Build** in the dialog, which should take you to
the **Build** page.
You're now ready for the next step, [querying the
data][ref-getting-started-core-query-cube].
[powershell-docs]: https://learn.microsoft.com/en-us/powershell/
[ref-conf-db]: /admin/connect-to-data/data-sources
[ref-getting-started-core-query-cube]: /cube-core/getting-started/query-data
[ref-devtools-playground]: /docs/explore-analyze/playground
[wsl2-docs]: https://learn.microsoft.com/en-us/windows/wsl/install
# Getting Started with Cube Core
Source: https://docs.cube.dev/cube-core/getting-started/index
Create a project, connect a database, query data, and add a pre-aggregation in self-hosted Cube Core.
Getting started with Cube Core
First, we'll create a new project, connect it to a database and generate a data
model from it. Then, we'll run queries using the Developer Playground and APIs.
Finally, we'll add a pre-aggregation to optimize query latency down to
milliseconds.
This guide will walk you through the following tasks:
* [Create a new project](/cube-core/getting-started/create-a-project)
* [Run queries using the Developer Playground and APIs](/cube-core/getting-started/query-data)
* [Add a pre-aggregation to optimize query performance](/cube-core/getting-started/add-a-pre-aggregation)
If you'd prefer to try Cube Cloud, then you can refer to [Getting Started using
Cube Cloud][ref-getting-started-cloud-overview] instead.
[ref-getting-started-cloud-overview]: /docs/getting-started/cloud
# Learn more
Source: https://docs.cube.dev/cube-core/getting-started/learn-more
Now that you've set up your first project, learn more about what else Cube can do for you.
## Data Modeling
Learn more about [data modeling](/docs/data-modeling/overview)
and how to effectively define metrics in your data models.
## Querying
Cube can be queried in a variety of ways. Explore how to use
[REST (JSON) API](/reference/core-data-apis/rest-api), [GraphQL API](/reference/core-data-apis/graphql-api), and
[SQL API](/reference/core-data-apis/sql-api), or how to
[connect a BI or data visualization tool](/admin/connect-to-data/visualization-tools).
## Caching
Learn more about the [two-level cache](/docs/pre-aggregations) and how
[pre-aggregations help speed up queries](/docs/pre-aggregations/getting-started-pre-aggregations).
## Access Control
Cube uses [JSON Web Tokens](https://jwt.io/) to
[authenticate requests for the HTTP APIs](/docs/data-modeling/access-control), and
[`check_sql_auth`](/reference/configuration/config#check_sql_auth) to
[authenticate requests for the SQL API](/reference/core-data-apis/sql-api/security).
# Query data
Source: https://docs.cube.dev/cube-core/getting-started/query-data
Surveys ways to explore and consume modeled data—Playground, APIs, client SDKs, and BI or app integrations—after cubes are defined.
In this step, you will learn how to query your data using the data models you
created in the previous step. Cube provides several ways to query your data, and
we'll go over them here.
## Playground
[Playground](/docs/explore-analyze/playground) is a web-based tool which allows for
model generation and data exploration. On the **Build** tab, you can
select the measures and dimensions, and then run the query. Let's do this for
the `orders` cube you generated in the previous step.
Click **+ Measure** to display the available measures and add
`orders.count`, then click **+ Dimension** for available dimensions and
add `orders.status`:
Then, click **Run** to execute the query and see the results:
Please feel free to experiment: select other measures or dimensions, pick a
granularity for the time dimension instead of **w/o grouping**, or choose
another chart type instead of **Table**.
## APIs and integrations
Cube provides a [rich set of options][ref-downstream] to deliver data to other
tools: a suite of APIs, [JavaScript SDKs][ref-frontend-int], and integrations.
Connectivity to BI tools and data notebooks is enabled by the [SQL
API][ref-sql-api] which is Postgres-compatible: if something connects to
Postgres, it will work with Cube. Check the **Connect to BI** tab for
connection instructions for specific BI tools:
Connectivity to data applications is enabled by the [REST (JSON) API][ref-rest-api] and
the [GraphQL API][ref-graphql-api] as well as [JavaScript
SDKs][ref-frontend-int]. Check the **Frontend Integrations** tab for
usage instructions for these APIs:
Now that we've seen how to use Cube's APIs, let's take a look at [how to add
pre-aggregations][next] to speed up your queries.
[ref-graphql-api]: /reference/core-data-apis/graphql-api
[ref-rest-api]: /reference/core-data-apis/rest-api
[ref-downstream]: /admin/connect-to-data/visualization-tools
[ref-frontend-int]: /reference/javascript-sdk
[ref-sql-api]: /reference/core-data-apis/sql-api
[next]: /cube-core/getting-started/add-a-pre-aggregation
# Cube Core
Source: https://docs.cube.dev/cube-core/index
Cube Core is the open-source semantic layer you can self-host with Docker.
Cube Core is the open-source version of Cube that you can deploy and manage yourself using Docker.
Get Cube Core running locally in minutes
Deploy to production with Docker
Environment variables and config options
Set up JWT authentication
# Running in production
Source: https://docs.cube.dev/cube-core/running-in-production
Cube makes use of two different kinds of cache:
* In-memory storage of query results
* Pre-aggregations
Cube Store is enabled by default when running Cube in development mode. In
production, Cube Store **must** run as a separate process. The easiest way to do
this is to use the official Docker images for Cube and Cube Store.
Development mode is an authentication bypass. Cube is in development mode when
`CUBEJS_DEV_MODE=true` — which, under the `cubejs` CLI and the official Docker images,
also forces `NODE_ENV=development` and so switches off JWT verification on the
REST (JSON) and GraphQL APIs — and whenever `NODE_ENV` is not
`production`. Playground's endpoints are served with no authentication too: anyone who
can reach the instance is handed a ready-to-use API token, can mint others carrying any
security context, can read your data model files, and can overwrite your data model and
your `.env`. With `CUBEJS_DEV_MODE=true` and no
[`CUBEJS_SQL_PASSWORD`](/reference/configuration/environment-variables#cubejs_sql_password)
set, the SQL API accepts any credentials as well, allowing arbitrary SQL against
connected data sources. Use it only on a local development machine, never in
production. See
[`CUBEJS_DEV_MODE`](/reference/configuration/environment-variables#cubejs_dev_mode).
Using Windows? We **strongly** recommend using [WSL2 for Windows 10][link-wsl2]
to run the following commands.
You can run Cube Store with Docker with the following command:
```bash theme={"dark"}
docker run -p 3030:3030 cubejs/cubestore
```
Cube Store can further be configured via environment variables. To see a
complete reference, please consult the `CUBESTORE_*` environment variables in
the [Environment Variables reference][ref-config-env].
Next, run Cube and tell it to connect to Cube Store running on `localhost` (on
the default port `3030`):
```bash theme={"dark"}
docker run -p 4000:4000 \
-e CUBEJS_CUBESTORE_HOST=localhost \
-v ${PWD}:/cube/conf \
cubejs/cube
```
In the command above, we're specifying [`CUBEJS_CUBESTORE_HOST`](/reference/configuration/environment-variables#cubejs_cubestore_host) to let Cube know
where Cube Store is running.
You can also use Docker Compose to achieve the same:
```yaml theme={"dark"}
services:
cubestore:
image: cubejs/cubestore:latest
environment:
- CUBESTORE_REMOTE_DIR=/cube/data
volumes:
- .cubestore:/cube/data
cube:
image: cubejs/cube:latest
ports:
- 4000:4000
environment:
- CUBEJS_CUBESTORE_HOST=localhost
depends_on:
- cubestore
links:
- cubestore
volumes:
- ./model:/cube/conf/model
```
## Architecture
A Cube Store cluster consists of at least one Router and one or more Worker
instances. Cube sends queries to the Cube Store Router, which then distributes
the queries to the Cube Store Workers. The Workers then execute the queries and
return the results to the Router, which in turn returns the results to Cube.
## Scaling
Cube Store *can* be run in a single instance mode, but this is usually
unsuitable for production deployments. For high concurrency and data throughput,
we **strongly** recommend running Cube Store as a cluster of multiple instances
instead.
Scaling Cube Store for a higher concurrency is relatively simple when running in
cluster mode. Because [the storage layer](#storage) is decoupled from the query
processing engine, you can horizontally scale your Cube Store cluster for as
much concurrency as you require.
In cluster mode, Cube Store runs two kinds of nodes:
* one or more **router** nodes handle incoming client connections, manage
database metadata and serve simple queries.
* multiple **worker** nodes which execute SQL queries
Cube Store querying performance is optimal when the count of partitions in a
single query is less than or equal to the worker count. For example, you have a
200 million rows table that is partitioned by day, which is ten daily Cube
partitions or 100 Cube Store partitions in total. The query sent by the user
contains filters, and the resulting scan requires reading 16 Cube Store
partitions in total. Optimal query performance, in this case, can be achieved
with 16 or more workers. You can use `EXPLAIN` and `EXPLAIN ANALYZE` SQL
commands to see how many partitions would be used in a specific Cube Store
query.
Resources required for the main node and workers can vary depending on the
configuration. With default settings, you should expect to allocate at least 4
CPUs and up to 8GB per main or worker node.
The configuration required for each node can be found in the table below. More
information about these variables can be found [in the Environment Variables
reference][ref-config-env].
| Environment Variable | Specify on Router? | Specify on Worker? |
| ----------------------- | ------------------ | ------------------ |
| `CUBESTORE_SERVER_NAME` | ✅ Yes | ✅ Yes |
| `CUBESTORE_META_PORT` | ✅ Yes | — |
| `CUBESTORE_WORKERS` | ✅ Yes | ✅ Yes |
| `CUBESTORE_WORKER_PORT` | — | ✅ Yes |
| `CUBESTORE_META_ADDR` | — | ✅ Yes |
`CUBESTORE_WORKERS` and `CUBESTORE_META_ADDR` variables should be set with
stable addresses, which should not change. You can use stable DNS names and put
load balancers in front of your worker and router instances to fulfill stable
name requirements in environments where stable IP addresses can't be guaranteed.
To fully take advantage of the worker nodes in the cluster, we **strongly**
recommend using [partitioned pre-aggregations][ref-caching-partitioning].
A sample Docker Compose stack for the single machine setting this up might look
like:
```yaml theme={"dark"}
services:
cubestore_router:
restart: always
image: cubejs/cubestore:latest
environment:
- CUBESTORE_SERVER_NAME=cubestore_router:9999
- CUBESTORE_META_PORT=9999
- CUBESTORE_WORKERS=cubestore_worker_1:9001,cubestore_worker_2:9001
- CUBESTORE_REMOTE_DIR=/cube/data
volumes:
- .cubestore:/cube/data
cubestore_worker_1:
restart: always
image: cubejs/cubestore:latest
environment:
- CUBESTORE_SERVER_NAME=cubestore_worker_1:9001
- CUBESTORE_WORKER_PORT=9001
- CUBESTORE_META_ADDR=cubestore_router:9999
- CUBESTORE_WORKERS=cubestore_worker_1:9001,cubestore_worker_2:9001
- CUBESTORE_REMOTE_DIR=/cube/data
depends_on:
- cubestore_router
volumes:
- .cubestore:/cube/data
cubestore_worker_2:
restart: always
image: cubejs/cubestore:latest
environment:
- CUBESTORE_SERVER_NAME=cubestore_worker_2:9001
- CUBESTORE_WORKER_PORT=9001
- CUBESTORE_META_ADDR=cubestore_router:9999
- CUBESTORE_WORKERS=cubestore_worker_1:9001,cubestore_worker_2:9001
- CUBESTORE_REMOTE_DIR=/cube/data
depends_on:
- cubestore_router
volumes:
- .cubestore:/cube/data
cube:
image: cubejs/cube:latest
ports:
- 4000:4000
environment:
- CUBEJS_CUBESTORE_HOST=cubestore_router
depends_on:
- cubestore_router
volumes:
- .:/cube/conf
```
## Replication and High Availability
The open-source version of Cube Store doesn't support replicating any of its
nodes. The router node and every worker node should always have only one
instance copy if served behind the load balancer or service address. Replication
will lead to undefined behavior of the cluster, including connection errors and
data loss. If any cluster node is down, it'll lead to a complete cluster outage.
If Cube Store replication and high availability are required, please consider
using Cube Cloud.
## Storage
Cube Store cluster uses both persistent and scratch storage.
### Persistent storage
Cube Store makes use of a separate storage layer for storing metadata as well as
for persisting pre-aggregations as Parquet files.
Cube Store can be configured to use either AWS S3, Google Cloud Storage (GCS), or
Azure Blob Storage as persistent storage. If desired, a local path on
the server can also be used in case all Cube Store cluster nodes are
co-located on a single machine.
Cube Store can only use one type of remote storage at the same time.
Cube Store requires strong consistency guarantees from an underlying distributed
storage. AWS S3, Google Cloud Storage, and Azure Blob Storage are the only known
implementations that provide them. Using other implementations in production is
discouraged and can lead to consistency and data corruption errors.
Available on [Enterprise plan](https://cube.dev/pricing).
As an additional layer on top of standard AWS S3, Google Cloud Storage (GCS), or
Azure Blob Storage encryption, persistent storage can optionally use [Parquet
encryption](#data-at-rest-encryption) for data-at-rest protection.
A simplified example using AWS S3 might look like:
```yaml theme={"dark"}
services:
cubestore_router:
image: cubejs/cubestore:latest
environment:
- CUBESTORE_SERVER_NAME=cubestore_router:9999
- CUBESTORE_META_PORT=9999
- CUBESTORE_WORKERS=cubestore_worker_1:9001
- CUBESTORE_S3_BUCKET=
- CUBESTORE_S3_REGION=
- CUBESTORE_AWS_ACCESS_KEY_ID=
- CUBESTORE_AWS_SECRET_ACCESS_KEY=
cubestore_worker_1:
image: cubejs/cubestore:latest
environment:
- CUBESTORE_SERVER_NAME=cubestore_worker_1:9001
- CUBESTORE_WORKER_PORT=9001
- CUBESTORE_META_ADDR=cubestore_router:9999
- CUBESTORE_WORKERS=cubestore_worker_1:9001
- CUBESTORE_S3_BUCKET=
- CUBESTORE_S3_REGION=
- CUBESTORE_AWS_ACCESS_KEY_ID=
- CUBESTORE_AWS_SECRET_ACCESS_KEY=
depends_on:
- cubestore_router
```
Note that you can’t use the same bucket as an export bucket and persistent
storage for Cube Store. It's recommended to use two separate buckets.
### Scratch storage
Separately from persistent storage, Cube Store requires local scratch space
to warm up partitions by downloading Parquet files before querying them.
By default, this folder should be mounted to `.cubestore/data` inside the
container and can be configured by `CUBESTORE_DATA_DIR` environment variable.
It is advised to use local SSDs for this scratch space to maximize querying
performance.
### AWS
Cube Store can retrieve security credentials from instance metadata
automatically. This means you can skip defining the
`CUBESTORE_AWS_ACCESS_KEY_ID` and `CUBESTORE_AWS_SECRET_ACCESS_KEY` environment
variables.
Cube Store currently does not take the key expiration time returned from
instance metadata into account; instead the refresh duration for the key is
defined by `CUBESTORE_AWS_CREDS_REFRESH_EVERY_MINS`, which is set to `180` by
default.
### Garbage collection
Cleanup isn’t done in export buckets; however, it's done in the persistent
storage of Cube Store. The default time-to-live (TTL) for orphaned
pre-aggregation tables is one day.
Refresh worker should be able to finish pre-aggregation refresh before
garbage collection starts. It means that all pre-aggregation partitions
should be built before any tables are removed.
#### Supported file systems
The garbage collection mechanism relies on the ability of the underlying file
system to report the creation time of a file.
If the file system does not support getting the creation time, you will see the
following error message in Cube Store logs:
```text theme={"dark"}
ERROR [cubestore::remotefs::cleanup]
error while getting created time for file ".chunk.parquet":
creation time is not available for the filesystem
```
XFS is known to not support getting the creation time of a file.
Please see [this issue](https://github.com/cube-js/cube/issues/7905#issuecomment-2504212623)
for possible workarounds.
## Security
### Authentication
Cube Store does not have any in-built authentication mechanisms. For this reason,
we recommend running your Cube Store cluster with a network configuration that
only allows access from the Cube deployment.
### Data-at-rest encryption
[Persistent storage](#persistent-storage) is secured using the standard AWS S3,
Google Cloud Storage (GCS), or Azure Blob Storage encryption.
Cube Store also provides optional data-at-rest protection by utilizing the
[modular encryption mechanism][link-parquet-encryption] of Parquet files in its
persistent storage. Pre-aggregation data is secured using the [AES cipher][link-aes]
with 256-bit keys. Data encyption and decryption are completely seamless to Cube
Store operations.
Available on the [Enterprise plan](https://cube.dev/pricing).
Also requires the M [Cube Store Worker tier](/admin/account-billing/pricing#cube-store-worker-tiers).
You can provide, rotate, or drop your own [customer-managed keys][ref-cmk] (CMK)
for Cube Store via the **Encryption Keys** page in Cube Cloud.
## Troubleshooting
### Heavy pre-aggregations
When building some pre-aggregations, you might encounter the following error:
```text theme={"dark"}
Error: Error during create table: CREATE TABLE
Error: Query execution timeout after 10 min of waiting
```
It means that your pre-aggregation is too heavy and takes too long to build.
As a temporary solution, you can increase the timeout for all queies that Cube
runs against your data source via the [`CUBEJS_DB_QUERY_TIMEOUT`](/reference/configuration/environment-variables#cubejs_db_query_timeout) environment variable.
However, it is recommended that you optimize your pre-aggregations instead:
* Use an export bucket if your [data source][ref-data-sources] supports it. Cube will then load the pre-aggregation data in a much more efficient way.
* Use [partitions][ref-pre-agg-partitions]. Cube will then run a separate query to build each partition.
* Build pre-aggregations [incrementally][ref-pre-agg-incremental]. Cube will then build only the necessary partitions with each pre-aggregation refresh.
* Set an appropriate [build range][ref-pre-agg-build-range] if you don't need to query the whole date range. Cube will then include only the necessary data in the pre-aggregation.
* Check that your pre-aggregation includes only necessary dimensions. Each additional dimension usually increases the volume of the pre-aggregation data.
* If you include a high cardinality dimension, Cube needs to store a lot of data in the pre-aggregation. For example, if you include the primary key into the pre-aggregation, Cube will effectively need to store a copy of the original table in the pre-aggregation, which is rarely useful.
* If a single pre-aggregation is used by queries with different sets of dimensions, consider creating separate pre-aggregations for each set of dimensions. This way, Cube will only include necessary data in each pre-aggregation.
* Check if you have a heavy calculation in the [`sql` expression][ref-cube-sql] of your cubes (rather than a simple `sql_table` reference). If it's the case, you can build an additional [`original_sql` pre-aggregation][ref-pre-agg-original-sql] and [instruct][ref-pre-agg-use-original-sql] Cube to use it when building other pre-aggregations for this cube.
### MinIO
When using MinIO for persistent storage, you might encounter the following error:
```text theme={"dark"}
Error: Error during upload of
File can't be listed after upload.
Either there's Cube Store cluster misconfiguration,
or storage can't provide the required consistency.
```
Most likely, it happens because MinIO is not providing strong consistency guarantees,
as required for Cube Store's [persistent storage](#persistent-storage).
You can either try configuring MinIO to provide strong consistency or switch to
using S3 or GCS.
The support for MinIO as persistent storage in Cube Store was contributed by the community.
It's not supported by Cube or the vendor.
[link-wsl2]: https://docs.microsoft.com/en-us/windows/wsl/install-win10
[ref-caching-partitioning]: /docs/pre-aggregations/using-pre-aggregations#partitioning
[ref-config-env]: /reference/configuration/environment-variables
[link-parquet-encryption]: https://parquet.apache.org/docs/file-format/data-pages/encryption/
[link-aes]: https://en.wikipedia.org/wiki/Advanced_Encryption_Standard
[ref-cmk]: /admin/deployment/encryption-keys
[ref-data-sources]: /admin/connect-to-data/data-sources
[ref-pre-agg-partitions]: /docs/pre-aggregations/using-pre-aggregations#partitioning
[ref-pre-agg-incremental]: /reference/data-modeling/pre-aggregations#incremental
[ref-pre-agg-build-range]: /reference/data-modeling/pre-aggregations#build_range_start-and-build_range_end
[ref-cube-sql]: /reference/data-modeling/cube#sql
[ref-pre-agg-original-sql]: /reference/data-modeling/pre-aggregations#original_sql
[ref-pre-agg-use-original-sql]: /reference/data-modeling/pre-aggregations#use_original_sql_pre_aggregations
# Security context
Source: https://docs.cube.dev/docs/data-modeling/access-control/context
How Cube derives verified user claims from inbound JWTs and where those claims plug into query rewrite and compile-time logic.
Your authentication server issues JWTs to your client application, which, when
sent as part of the request, are verified and decoded by Cube to get security
context claims to evaluate access control rules. Inbound JWTs are decoded and
verified using industry-standard [JSON Web Key Sets (JWKS)][link-auth0-jwks].
For access control or authorization, Cube allows you to define granular access
control rules for every cube in your data model. Cube uses both the request and
security context claims in the JWT token to generate a SQL query, which includes
row-level constraints from the access control rules.
JWTs sent to Cube should be passed in the `Authorization: ` header to
authenticate requests.
JWTs can also be used to pass additional information about the user, known as a
**security context**. A security context is a verified set of claims about the
current user that the Cube server can use to ensure that users only have access
to the data that they are authorized to access.
It will be accessible as the [`securityContext`][ref-config-sec-ctx] property
inside:
* The [`query_rewrite`][ref-config-queryrewrite] configuration option in your
Cube configuration file.
* the [`COMPILE_CONTEXT`][ref-cubes-compile-ctx] global, which is used to
support [multi-tenant deployments][link-multitenancy].
## Contents
By convention, the contents of the security context should be an object (dictionary)
with nested structure:
```json theme={"dark"}
{
"sub": "1234567890",
"iat": 1516239022,
"user_name": "John Doe",
"user_id": 42,
"location": {
"city": "San Francisco",
"state": "CA"
}
}
```
### Reserved elements
Some features of Cube Cloud (e.g., [authentication integration][ref-auth-integration])
use the `cubeCloud` element in the security context.
This element is reserved and should not be used for other purposes.
## Using query\_rewrite
You can use [`query_rewrite`][ref-config-queryrewrite] to amend incoming queries
with filters. For example, let's take the following query:
```json theme={"dark"}
{
"measures": [
"orders_view.count"
],
"dimensions": [
"orders_view.status"
]
}
```
We'll also use the following as a JWT payload; `user_id`, `sub` and `iat` will
be injected into the security context:
```json theme={"dark"}
{
"sub": "1234567890",
"iat": 1516239022,
"user_id": 42
}
```
Cube expects the context to be an object. If you don't provide an object as the
JWT payload, you will receive the following error:
```bash theme={"dark"}
Cannot create proxy with a non-object as target or handler
```
To ensure that users making this query only receive their own orders, define
`query_rewrite` in the configuration file:
```python title="Python" theme={"dark"}
from cube import config
@config('query_rewrite')
def query_rewrite(query: dict, ctx: dict) -> dict:
if 'user_id' in ctx['securityContext']:
query['filters'].append({
'member': 'orders_view.users_id',
'operator': 'equals',
'values': [ctx['securityContext']['user_id']]
})
return query
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
queryRewrite: (query, { securityContext }) => {
if (securityContext.user_id) {
query.filters.push({
member: "orders_view.users_id",
operator: "equals",
values: [securityContext.user_id]
})
}
return query
}
}
```
To test this, we can generate an API token as follows:
```python title="Python" theme={"dark"}
# Install the PyJWT with pip install PyJWT
import jwt
import datetime
# Secret key to sign the token
CUBE_API_SECRET = 'secret'
# Create the token
token_payload = {
'user_id': 42
}
# Generate the JWT token
token = jwt.encode(token_payload, CUBE_API_SECRET, algorithm='HS256')
```
```javascript title="JavaScript" theme={"dark"}
const jwt = require("jsonwebtoken")
const CUBE_API_SECRET = "secret"
const cubeToken = jwt.sign({ user_id: 42 }, CUBE_API_SECRET, {
expiresIn: "30d"
})
```
Using this token, we authorize our request to the Cube API by passing it in the
Authorization HTTP header.
```bash theme={"dark"}
curl \
-H "Authorization: eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1Ijp7ImlkIjo0Mn0sImlhdCI6MTU1NjAyNTM1MiwiZXhwIjoxNTU4NjE3MzUyfQ._8QBL6nip6SkIrFzZzGq2nSF8URhl5BSSSGZYp7IJZ4" \
-G \
--data-urlencode 'query={"measures":["orders.count"]}' \
http://localhost:4000/cubejs-api/v1/load
```
And Cube will generate the following SQL:
```sql theme={"dark"}
SELECT
"orders".STATUS "orders_view__status",
count("orders".ID) "orders_view__count"
FROM
ECOM.ORDERS AS "orders"
LEFT JOIN ECOM.USERS AS "users" ON "orders".USER_ID = "users".ID
WHERE
("users".ID = 42)
GROUP BY
1
ORDER BY
2 DESC
LIMIT
5000
```
## Using COMPILE\_CONTEXT
`COMPILE_CONTEXT` can be used to create fully dynamic data models. It enables you to create multiple versions of data model based on the incoming security context.
The first thing you need to do is to define the mapping rule from a security context to the id of the compiled data model.
It is done with `context_to_app_id` configuration option.
```python theme={"dark"}
from cube import config
@config('context_to_app_id')
def context_to_app_id(ctx: dict) -> str:
return ctx['securityContext']['team']
```
It is common to use some field from the incoming security context as an id for your data model.
In our example, as illustrated below, we are using `team` property of the security context as a data model id.
Once you have this mapping, you can use `COMPILE_CONTEXT` inside your data model.
In the example below we are passing it as a variable into `masked` helper function.
```yaml theme={"dark"}
cubes:
- name: users
sql_table: ECOM.USERS
public: false
dimensions:
- name: last_name
sql: {{ masked('LAST_NAME', COMPILE_CONTEXT.securityContext) }}
type: string
```
This `masked` helper function is defined in `model/globals.py` as follows: it checks if the current `team` is inside the list of trusted teams.
If that's the case, it will render the SQL to get the value of the dimension; if not, it will return just the masked string.
```python theme={"dark"}
from cube import TemplateContext
template = TemplateContext()
@template.function('masked')
def masked(sql, security_context):
trusted_teams = ['cx', 'exec' ]
is_trusted_team = security_context.setdefault('team') in trusted_teams
if is_trusted_team:
return sql
else:
return "'--- masked ---'"
```
### Usage with pre-aggregations
To generate pre-aggregations that rely on `COMPILE_CONTEXT`, [configure
`scheduledRefreshContexts` in your `cube.js` configuration
file][ref-config-sched-refresh].
### Usage for member-level security
You can also use `COMPILE_CONTEXT` to control whether a data model entity should be
public or private [dynamically][ref-dynamic-data-modeling].
In the example below, the `customers` view would only be visible to a subset
of [tenants][ref-multitenancy] that have the `team` property set to `marketing`
in the security context:
```ymltitle="model/views/customers.yml" theme={"dark"}
views:
- name: customers
public: "{{ is_accessible_by_team('marketing', COMPILE_CONTEXT) }}"
```
```pythontitle="model/globals.py" theme={"dark"}
from cube import TemplateContext
template = TemplateContext()
@template.function('is_accessible_by_team')
def is_accessible_by_team(team: str, ctx: dict) -> bool:
return team == ctx['securityContext'].setdefault('team', 'default')
```
If you'd like to keep a data model entity public but prevent access to it
anyway, you can use the \[`query_rewrite` configuration option]\[ref-query-rewrite] for that.
## Testing during development
During development, it is often useful to be able to edit the security context
to test access control rules. The [Developer
Playground][ref-devtools-playground] allows you to set your own JWTs, or you can
build one from a JSON object.
## Enriching the security context
Sometimes it is convenient to enrich the security context with additional attributes
before it is used to evaluate access control rules.
### Extending the security context
You can use the [`extend_context`][ref-extend-context] configuration option to
enrich the security context with additional attributes.
### Authentication integration
When using Cube Cloud, you can enrich the security context with information about
an authenticated user, obtained during their authentication.
You can enable the authentication integration by navigating to the **Settings → Configuration**
of your Cube Cloud deployment and using the **Enable Cloud Auth Integration** toggle.
## Common patterns
### Enforcing mandatory filters
You can use `query_rewrite` to enforce mandatory filters that apply to all queries. This is useful when you need to ensure certain conditions are always met, such as filtering data by date range or restricting access to specific data subsets.
For example, if you want to only show orders created after a specific date across all queries, you can add a mandatory filter:
```python title="Python" theme={"dark"}
from cube import config
@config('query_rewrite')
def query_rewrite(query: dict, ctx: dict) -> dict:
query['filters'].append({
'member': 'orders.created_at',
'operator': 'afterDate',
'values': ['2019-12-30']
})
return query
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
queryRewrite: (query) => {
query.filters.push({
member: `orders.created_at`,
operator: "afterDate",
values: ["2019-12-30"]
})
return query
}
}
```
This filter will be automatically applied to all queries, ensuring that only orders created after December 30th, 2019 are returned, regardless of any other filters specified in the query.
### Enforcing role-based access
You can use `query_rewrite` to enforce role-based access control by filtering data based on the user's role from the security context.
For example, to restrict access so that users with the `operator` role can only view processing orders, while users with the `manager` role can only view shipped and completed orders:
```python title="Python" theme={"dark"}
from cube import config
@config('query_rewrite')
def query_rewrite(query: dict, ctx: dict) -> dict:
if not ctx['securityContext'].get('role'):
raise ValueError("No role found in Security Context!")
role = ctx['securityContext']['role']
if role == "manager":
query['filters'].append({
'member': 'orders.status',
'operator': 'equals',
'values': ['shipped', 'completed']
})
elif role == "operator":
query['filters'].append({
'member': 'orders.status',
'operator': 'equals',
'values': ['processing']
})
return query
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
queryRewrite: (query, { securityContext }) => {
if (!securityContext.role) {
throw new Error("No role found in Security Context!")
}
if (securityContext.role == "manager") {
query.filters.push({
member: "orders.status",
operator: "equals",
values: ["shipped", "completed"]
})
}
if (securityContext.role == "operator") {
query.filters.push({
member: "orders.status",
operator: "equals",
values: ["processing"]
})
}
return query
}
}
```
### Enforcing column-based access
You can use `query_rewrite` to enforce column-based access control by filtering data based on relationships and user attributes from the security context.
For example, to restrict suppliers to only see their own products based on their email:
```python title="Python" theme={"dark"}
from cube import config
@config('query_rewrite')
def query_rewrite(query: dict, ctx: dict) -> dict:
cube_names = [
*(query.get('dimensions', [])),
*(query.get('measures', []))
]
cube_names = [e.split(".")[0] for e in cube_names]
if "products" in cube_names:
email = ctx['securityContext'].get('email')
if not email:
raise ValueError("No email found in Security Context!")
query['filters'].append({
'member': 'suppliers.email',
'operator': 'equals',
'values': [email]
})
return query
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
queryRewrite: (query, { securityContext }) => {
const cubeNames = [
...(query.dimensions || []),
...(query.measures || [])
].map((e) => e.split(".")[0])
if (cubeNames.includes("products")) {
if (!securityContext.email) {
throw new Error("No email found in Security Context!")
}
query.filters.push({
member: `suppliers.email`,
operator: "equals",
values: [securityContext.email]
})
}
return query
}
}
```
### Controlling access to cubes and views
You can use `extend_context` and `COMPILE_CONTEXT` to control access to cubes and views based on user properties from the security context.
For example, to make a view accessible only to users with a `department` claim set to `finance`:
```python title="Python" theme={"dark"}
from cube import config
@config('extend_context')
def extend_context(ctx: dict) -> dict:
return {
'securityContext': {
**ctx['securityContext'],
'isFinance': ctx['securityContext'].get('department') == 'finance'
}
}
```
```javascript title="JavaScript" theme={"dark"}
module.exports = {
extendContext: ({ securityContext }) => {
return {
securityContext: {
...securityContext,
isFinance: securityContext.department === "finance"
}
}
}
}
```
Then in your data model, use `COMPILE_CONTEXT` to control visibility:
```yaml title="YAML" theme={"dark"}
views:
- name: total_revenue_per_customer
public: {{ COMPILE_CONTEXT['securityContext']['isFinance'] }}
# ...
```
```javascript title="JavaScript" theme={"dark"}
view(`total_revenue_per_customer`, {
public: COMPILE_CONTEXT.securityContext.isFinance,
// ...
})
```
[link-auth0-jwks]: https://auth0.com/docs/tokens/json-web-tokens/json-web-key-sets
[link-multitenancy]: /embedding/multitenancy
[ref-config-queryrewrite]: /reference/configuration/config#query_rewrite
[ref-config-sched-refresh]: /reference/configuration/config#scheduledrefreshcontexts
[ref-config-sec-ctx]: /reference/configuration/config#securitycontext
[ref-cubes-compile-ctx]: /reference/data-modeling/context-variables#compile_context
[ref-devtools-playground]: /docs/explore-analyze/playground#editing-the-security-context
[ref-auth-integration]: /docs/data-modeling/access-control#authentication-integration
[ref-extend-context]: /reference/configuration/config#extend_context
[ref-dynamic-data-modeling]: /docs/data-modeling/dynamic
[ref-multitenancy]: /embedding/multitenancy
# Access Control
Source: https://docs.cube.dev/docs/data-modeling/access-control/index
Learn how Cube separates authentication and authorization to control who can access your data and platform features.
Access control
Access control in Cube involves *authentication* and *authorization*.
## Authentication
Authentication determines if a user can access Cube.
* **Cube** cloud platform provides built-in authentication mechanisms. Users are assigned
[roles and permissions][ref-roles-perms] that determine available features of the Cube
platform.
* **Cube Core** provides several [authentication methods][ref-auth-methods] for its API
endpoints.
## Authorization
Authorization determines what data a user can access though Cube.
Authorization is managed declaratively via [access policies][ref-dap], a built-in
capability of Cube's data modeling layer. There are also programmatic controls for
advanced use cases, such as the [`query_rewrite`][ref-query-rewrite] configuration
parameter.
* **Cube** cloud platform applies access policies to users based on their
[groups][ref-user-groups] and [attributes][ref-user-attributes].
* **Cube Core** applies access policies to users based on their groups derived from the
[security context][ref-sec-ctx]. See the [`context_to_groups`][ref-ctx-to-groups]
configuration parameter for details.
[ref-roles-perms]: /admin/users-and-permissions/roles-and-permissions
[ref-auth-methods]: /embedding/authentication/jwt
[ref-user-groups]: /admin/users-and-permissions/user-groups
[ref-user-attributes]: /admin/users-and-permissions/user-attributes
[ref-sec-ctx]: /docs/data-modeling/access-control/context
[ref-dap]: /docs/data-modeling/data-access-policies
[ref-query-rewrite]: /reference/configuration/config#query_rewrite
[ref-ctx-to-groups]: /reference/configuration/config#context_to_groups
# Member-level security
Source: https://docs.cube.dev/docs/data-modeling/access-control/member-level-security
Covers restricting which cubes, views, and members end users can query, including defaults and policy configuration with includes and excludes.
The data model serves as a facade of your data. With member-level security,
you can define whether [data model entities][ref-data-modeling-concepts] (cubes, views,
and their members) are exposed to end users and can be queried via [APIs &
integrations][ref-apis].
Member-level security in Cube is similar to column-level security in SQL databases.
Defining whether users have access to [cubes][ref-cubes] and [views][ref-views] is
similar to defining access to database tables; defining whether they have access
to dimensions and measures — to columns.
**By default, all cubes, views, and their members are *public*,** meaning that they
can be accessed by any users and they are also visible during data model introspection.
## Managing member-level access
You can use [access policies][ref-dap] to configure member-level access
for different groups. With the `access_policy` parameter in
[cubes][ref-ref-cubes] and [views][ref-ref-views], you can define which members
are accessible to users with specific groups.
Use the `member_level` parameter to specify either:
* `includes`: a list of allowed members, or
* `excludes`: a list of disallowed members
You can use `"*"` as a shorthand to include or exclude all members.
When you define access policies for specific groups, access is automatically denied to all other groups. You don't need to create a default policy that denies access.
In the following example, member-level access is configured for different groups:
```yaml title="YAML" theme={"dark"}
views:
- name: orders_view
cubes:
- join_path: orders
includes:
- status
- created_at
- count
- count_7d
- count_30d
access_policy:
# Managers can access all members except for `count`
- group: manager
member_level:
excludes:
- count
# Observers can access all members except for `count` and `count_7d`
- group: observer
member_level:
excludes:
- count
- count_7d
# Guests can only access the `count_30d` measure
- group: guest
member_level:
includes:
- count_30d
```
```javascript title="JavaScript" theme={"dark"}
view(`orders_view`, {
cubes: [
{
join_path: orders,
includes: [
`status`,
`created_at`,
`count`,
`count_7d`,
`count_30d`
]
}
],
access_policy: [
{
// Managers can access all members except for `count`
group: `manager`,
member_level: {
excludes: [
`count`
]
}
},
{
// Observers can access all members except for `count` and `count_7d`
group: `observer`,
member_level: {
excludes: [
`count`,
`count_7d`
]
}
},
{
// Guests can only access the `count_30d` measure
group: `guest`,
member_level: {
includes: [
`count_30d`
]
}
}
]
})
```
This configuration results in the following access:
| Group | Access |
| --------------- | --------------------------------------------- |
| `manager` | All members except for `count` |
| `observer` | All members except for `count` and `count_7d` |
| `guest` | Only the `count_30d` measure |
| All other users | No access to this view at all |
Access policies also respect member-level security restrictions configured via
`public` parameters. For more details, see the [access policies
reference][ref-dap-ref].
If you want to return masked values for restricted members instead of hiding
them entirely, see [data masking][ref-data-masking] in access policies.
[ref-data-modeling-concepts]: /docs/data-modeling/overview
[ref-apis]: /reference
[ref-cubes]: /docs/data-modeling/cubes
[ref-views]: /docs/data-modeling/views
[ref-dap]: /docs/data-modeling/data-access-policies
[ref-ref-cubes]: /reference/data-modeling/cube
[ref-ref-views]: /reference/data-modeling/view
[ref-dev-mode]: /docs/data-modeling/dev-mode
[ref-playground]: /docs/explore-analyze/playground
[ref-dap-ref]: /reference/data-modeling/data-access-policies
[ref-cubes-public]: /reference/data-modeling/cube#public
[ref-views-public]: /reference/data-modeling/view#public
[ref-measures-public]: /reference/data-modeling/measures#public
[ref-dimensions-public]: /reference/data-modeling/dimensions#public
[ref-hierarchies-public]: /reference/data-modeling/hierarchies#public
[ref-segments-public]: /reference/data-modeling/segments#public
[ref-dynamic-data-modeling]: /docs/data-modeling/dynamic
[ref-security-context]: /docs/data-modeling/access-control/context
[ref-data-masking]: /docs/data-modeling/data-access-policies#data-masking
# Row-level security
Source: https://docs.cube.dev/docs/data-modeling/access-control/row-level-security
Covers applying group- and attribute-based filters so query results only include rows users are allowed to see.
The data model serves as a facade of your data. With row-level security,
you can define whether some [data model][ref-data-modeling-concepts] facts are exposed
to end users and can be queried via [APIs & integrations][ref-apis].
Row-level security in Cube is similar to row-level security in SQL databases.
Defining whether users have access to specific facts from [cubes][ref-cubes] and
[views][ref-views] is similar to defining access to rows in database tables.
**By default, all rows are *public*,** meaning that no filtering is applied to
data model facts when they are accessed by any users.
## Managing row-level access
You can use [access policies][ref-dap] to manage both [member-level][ref-mls]
and row-level security based on groups and [user attributes][ref-security-context].
Here's an example of how to filter rows by a user attribute using access policies:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
access_policy:
- group: manager
row_level:
filters:
- member: country
operator: equals
values: [ "{ userAttributes.country }" ]
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
access_policy: [
{
group: `manager`,
row_level: {
filters: [
{
member: `country`,
operator: `equals`,
values: [ userAttributes.country ]
}
]
}
}
]
})
```
[ref-data-modeling-concepts]: /docs/data-modeling/overview
[ref-apis]: /reference
[ref-cubes]: /docs/data-modeling/cubes
[ref-views]: /docs/data-modeling/views
[ref-cubes-sql]: /reference/data-modeling/cube#sql
[ref-dynamic-data-modeling]: /docs/data-modeling/dynamic
[ref-dap]: /docs/data-modeling/data-access-policies
[ref-mls]: /docs/data-modeling/access-control/member-level-security
[ref-security-context]: /docs/data-modeling/access-control/context
# Access Policies viewer
Source: https://docs.cube.dev/docs/data-modeling/access-policies-viewer
Audit row-level, member-level, and member-masking access policies that govern your data model from the Cube Cloud UI, grouped by user group.
The Access Policies viewer surfaces, in one place, every [access policy][ref-access-policies]
defined in your [data model][ref-data-modeling] — row-level filters, member-level
restrictions, and member masking — broken down by the user [groups][ref-user-groups]
they apply to.
Use it to audit who can see which cubes and views, and how each policy is composed,
without grepping through `cube` files or running test queries.
The viewer is read-only. Access policies themselves are authored in the
[data model][ref-access-policies] using `access_policy` blocks; this page
visualizes the resolved rules so you can review and debug them.
## Opening the viewer
In Cube Cloud, navigate to the **Model** module and click **Access Policies** in
the sub-sidebar. The viewer reflects whichever branch and build you are currently
viewing, so policies you are editing in [development mode][ref-dev-mode] appear
alongside what is live in production.
You need the `PlaygroundRead` permission to open the viewer.
## List view
The list view shows one row per group declared anywhere in the data model:
| Column | What it shows |
| ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Group** | Name of the group. The wildcard entry `*` is rendered as **All Groups** — this is the catch-all default policy applied when no other policy matches. |
| **Policies** | Number of cubes and views with an explicit policy for this group. Hover the cell to see the full list of cube and view names. |
| **Default Policy** | Number of cubes and views this group can access without an explicit policy — the union of cubes covered by the wildcard `*` policy and any cubes that have no policy at all. |
Cubes and views with no `access_policy` block defined are considered fully open;
they appear under **Default Policy** for every group.
Click a row to drill into the per-cube breakdown for that group.
## Per-policy detail view
The detail view shows one row per cube or view that the selected group can
access, with the resolved policy expanded across four columns:
| Column | What it shows |
| ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Cube / View** | Name of the cube or view, with an icon distinguishing the two. |
| **Condition** | The number of [`condition`][ref-policy-condition] expressions on the policy, or `—` if the policy applies unconditionally. Conditions are arbitrary expressions defined in the model. |
| **Member-level Access** | One of three states: **Allow All** (no member-level restrictions), **Deny All** (member access is fully denied), or **Allow:** followed by the resolved set of allowed dimensions, segments, and measures. |
| **Member Masking** | `—` if no [member masking][ref-mls-masking] applies, otherwise the list of masked dimensions. |
| **Row-level Access** | Either **Allow All**, or **Filters on:** followed by the dimensions referenced by the row-level filter. |
Member names are shortened to the last path segment for readability — for
example, `orders.user.email` is shown as `email`.
## What the viewer does not do
The viewer is intentionally scoped to inspecting policies that are already
defined in the model. It does not:
* Create, edit, or delete access policies. Edit `access_policy` blocks in your
data model and commit through your normal Git workflow.
* Show which individual users belong to a given group. See
[User groups][ref-user-groups] for membership management.
* Run preview queries against a policy. To verify behavior end-to-end, switch
the security context and issue queries against your development API.
[ref-data-modeling]: /docs/data-modeling/overview
[ref-access-policies]: /docs/data-modeling/data-access-policies
[ref-policy-condition]: /reference/data-modeling/data-access-policies#conditions
[ref-mls-masking]: /docs/data-modeling/data-access-policies#data-masking
[ref-dev-mode]: /docs/data-modeling/dev-mode
[ref-user-groups]: /admin/users-and-permissions/user-groups
# AI context
Source: https://docs.cube.dev/docs/data-modeling/ai-context
Improve AI accuracy and trust by enriching your semantic layer with descriptions and AI-specific context that helps agents generate better insights.
When using [Analytics Chat][ref-analytics-chat] or other AI-powered features,
the AI agent relies on your data model to understand your data. You can
optimize your data model to help the AI generate more accurate queries and
provide better insights.
There are two ways to provide additional context to the AI:
* **Descriptions** — visible to both end users and the AI agent.
* **AI context via `meta`** — only visible to the AI agent, not exposed in the
user interface.
## Using descriptions
The [`description`][ref-cube-description] parameter on cubes, views, measures,
dimensions, and segments provides human-readable context that is displayed in
the UI and also consumed by the AI agent.
Use descriptions to clarify the meaning of a member for both your team and
end users:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
description: All orders including pending, shipped, and completed
measures:
- name: total_revenue
sql: amount
type: sum
description: Total revenue from completed orders only
filters:
- sql: "{CUBE}.status = 'completed'"
dimensions:
- name: status
sql: status
type: string
description: "Current order status: pending, shipped, or completed"
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
description: `All orders including pending, shipped, and completed`,
measures: {
total_revenue: {
sql: `amount`,
type: `sum`,
description: `Total revenue from completed orders only`,
filters: [{ sql: `${CUBE}.status = 'completed'` }]
}
},
dimensions: {
status: {
sql: `status`,
type: `string`,
description: `Current order status: pending, shipped, or completed`
}
}
})
```
Descriptions are a good starting point because they serve double duty — they
help end users understand the data and also give the AI agent context for
query generation.
## Using AI context
If you want to provide context to the AI agent **without exposing it in the
user interface**, use the `ai_context` key inside the
[`meta`][ref-cube-meta] parameter. The `meta` parameter accepts custom
metadata on views, measures, and dimensions.
`ai_context` must be defined on **views** or on **individual members**
(measures, dimensions). `ai_context` defined at the cube level is **not
consumed by the AI agent**.
Each `ai_context` value is limited to 2,000 characters; anything longer is
silently truncated before it reaches the agent.
Use `ai_context` on [views][ref-view-meta] to provide high-level guidance,
and on individual members for member-specific instructions:
```yaml title="YAML" theme={"dark"}
views:
- name: revenue_overview
description: Revenue metrics and breakdowns
meta:
ai_context: >
This is the primary view for revenue analysis. It combines
order, product, and user data. Use this view when users ask
about sales, revenue, or product performance.
cubes:
- join_path: order_items
includes:
- total_sale_price
- count
- status
- created_at
- join_path: order_items.products
includes:
- brand
- category
```
```javascript title="JavaScript" theme={"dark"}
view(`revenue_overview`, {
description: `Revenue metrics and breakdowns`,
meta: {
ai_context: `This is the primary view for revenue analysis. It combines
order, product, and user data. Use this view when users ask about
sales, revenue, or product performance.`
},
cubes: [
{
join_path: order_items,
includes: [
`total_sale_price`,
`count`,
`status`,
`created_at`
]
},
{
join_path: order_items.products,
includes: [
`brand`,
`category`
]
}
]
})
```
For member-level context, define `ai_context` directly on the measure or
dimension:
```yaml title="YAML" theme={"dark"}
cubes:
- name: order_items
sql_table: ECOMMERCE.ORDER_ITEMS
measures:
- name: total_sale_price
sql: sale_price
type: sum
format: currency
meta:
ai_context: >
Use this measure for any revenue-related questions.
It includes all line items regardless of order status.
dimensions:
- name: created_at
sql: created_at
type: time
meta:
ai_context: >
This is the order creation timestamp in UTC.
For delivery analysis, use delivered_at instead.
```
```javascript title="JavaScript" theme={"dark"}
cube(`order_items`, {
sql_table: `ECOMMERCE.ORDER_ITEMS`,
measures: {
total_sale_price: {
sql: `sale_price`,
type: `sum`,
format: `currency`,
meta: {
ai_context: `Use this measure for any revenue-related questions.
It includes all line items regardless of order status.`
}
}
},
dimensions: {
created_at: {
sql: `created_at`,
type: `time`,
meta: {
ai_context: `This is the order creation timestamp in UTC.
For delivery analysis, use delivered_at instead.`
}
}
}
})
```
You can also override member-level `ai_context` when including members in
a view — for example, to define synonyms or acronyms that only apply in the
context of that view:
```yaml title="YAML" theme={"dark"}
views:
- name: sales_overview
description: Sales metrics and breakdowns
meta:
ai_context: >
This view is for sales performance analysis across brands.
cubes:
- join_path: order_items
includes:
- total_sale_price
- join_path: order_items.products
includes:
- name: brand
meta:
ai_context: >
Common acronyms: LC = Lucky Charms,
HNC = Honey Nut Cheerios.
```
```javascript title="JavaScript" theme={"dark"}
view(`sales_overview`, {
description: `Sales metrics and breakdowns`,
meta: {
ai_context: `This view is for sales performance analysis across brands.`
},
cubes: [
{
join_path: order_items,
includes: [`total_sale_price`]
},
{
join_path: order_items.products,
includes: [
{
name: `brand`,
meta: {
ai_context: `Common acronyms: LC = Lucky Charms,
HNC = Honey Nut Cheerios.`
}
}
]
}
]
})
```
## Descriptions vs. AI context
| | `description` | `meta.ai_context` |
| -------------------- | -------------------------------------------- | --------------------------- |
| Visible in the UI | Yes | No |
| Used by the AI agent | Yes | Yes |
| Supported on | Cubes, views, measures, dimensions, segments | Views, measures, dimensions |
| Length limit | None | 2,000 characters |
Use `description` when the context is useful to both end users and the AI
agent. Use `ai_context` when you want to provide additional instructions or
context that is only relevant to the AI agent — for example, guidance on
which measures to prefer, nuances about data quality, or business logic that
would be confusing in a user-facing description.
You can use both together. The AI agent reads both the `description` and
`ai_context` when generating queries:
```yaml title="YAML" theme={"dark"}
cubes:
- name: order_items
sql_table: ECOMMERCE.ORDER_ITEMS
description: Line items for all orders
measures:
- name: total_sale_price
sql: sale_price
type: sum
format: currency
description: Total revenue across all line items
meta:
ai_context: >
This is the primary revenue metric. Always use this
instead of summing the sale_price column directly.
When users ask about "sales", they mean this measure.
```
```javascript title="JavaScript" theme={"dark"}
cube(`order_items`, {
sql_table: `ECOMMERCE.ORDER_ITEMS`,
description: `Line items for all orders`,
measures: {
total_sale_price: {
sql: `sale_price`,
type: `sum`,
format: `currency`,
description: `Total revenue across all line items`,
meta: {
ai_context: `This is the primary revenue metric. Always use this
instead of summing the sale_price column directly.
When users ask about "sales", they mean this measure.`
}
}
}
})
```
## Best practices
* **Add descriptions to all public members.** Descriptions help both end users
and the AI agent understand your data model.
* **Use AI context for agent-specific guidance.** If you need to tell the AI
agent which measure to prefer or how to interpret ambiguous terms, use
`ai_context`.
* **Define context on views or individual members.** `ai_context` defined
at the cube level is not consumed by the AI agent. Place it on the view
itself or on individual measures and dimensions.
* **Be specific.** Vague context like "important metric" is less helpful than
"use this measure when users ask about monthly recurring revenue."
* **Document relationships.** Use AI context to explain how cubes relate to each
other and which views to prefer for common questions.
* **Keep it up to date.** As your data model evolves, update descriptions and AI
context to reflect the current state.
[ref-analytics-chat]: /docs/explore-analyze/analytics-chat
[ref-cube-description]: /reference/data-modeling/cube#description
[ref-cube-meta]: /reference/data-modeling/cube#meta
[ref-view-meta]: /reference/data-modeling/view#meta
# Syntax
Source: https://docs.cube.dev/docs/data-modeling/concepts/syntax
Specifies the model directory layout, file naming, and when to use YAML versus JavaScript for Cube data model files.
Entities within the data model (e.g., cubes, views, etc.) should be placed under
the `model` [folder][self-folder-structure], follow [naming
conventions][self-naming], and be defined using a supported
[syntax][self-syntax].
## Folder structure
Data model files should be placed inside the `model` folder. You can use the
[`schema_path` configuration option][ref-config-model-path] to override the
folder name or the [`repository_factory` configuration
option][ref-config-repository-factory] to dynamically define the folder name and
data model file contents.
It's recommended to place each cube or view in a separate file, in `model/cubes`
and `model/views` folders, respectively. [View groups][ref-view-groups] can be
defined alongside views or in their own files. Example:
```tree theme={"dark"}
model
├── cubes
│ ├── orders.yml
│ ├── products.yml
│ └── users.yml
└── views
├── revenue.yml
└── view_groups.yml
```
## Model syntax
Cube supports two ways to define data model files: with [YAML][wiki-yaml] or
JavaScript syntax. YAML data model files should have the `.yml` extension,
whereas JavaScript data model files should end with `.js`. You can mix YAML and
JavaScript files within a single data model.
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql: |
SELECT *
FROM orders, line_items
WHERE orders.id = line_items.order_id
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql: `
SELECT *
FROM orders, line_items
WHERE orders.id = line_items.order_id
`
})
```
You can define the data model statically or build [dynamic data
models][ref-dynamic-data-models] programmatically. YAML data models use
[Jinja and Python][ref-dynamic-data-models-jinja] whereas JavaScript data
models use [JavaScript][ref-dynamic-data-models-js].
It is recommended to default to YAML syntax because of its
simplicity and readability. However, JavaScript might provide more flexibility
for dynamic data modeling.
See [Cube style guide][ref-style-guide] for more recommendations on syntax and structure.
## Naming
Common rules apply to names of entities within the data model. All names must:
* Start with a letter.
* Consist of letters, numbers, and underscore (`_`) symbols only.
* Not be a [reserved keyword in Python][link-python-reserved-words], e.g., `from`, `return`, or `yield`.
* When using the DAX API, not clash with the names of columns in date hierarchies.
* Be unique within their scope. Cube and view names must be unique across the data
model; members (measures, dimensions, segments, pre-aggregations, hierarchies) must
be unique within their cube; and folder names must be unique within their view. Cube
will report an error for duplicates.
It is also recommended that names use [snake case][wiki-snake-case].
Good examples of names:
* `orders`, `stripe_invoices`, or `base_payments` (cubes)
* `opportunities`, `cloud_accounts`, or `arr` (views)
* `count`, `avg_price`, or `total_amount_shipped` (measures)
* `name`, `is_shipped`, or `created_at` (dimensions)
* `main`, `orders_by_status`, or `lambda_invoices` (pre-aggregations)
## SQL expressions
When defining cubes, you would often provide SQL snippets in `sql` and
`sql_table` parameters.
Provided SQL expressions should match your database SQL dialect, e.g.,
to aggregate a list of strings, you would probably pick the [`LISTAGG`
function][link-snowflake-listagg] in Snowflake and the [`STRING_AGG`
function][link-bigquery-stringagg] in BigQuery.
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
measures:
- name: statuses
sql: "STRING_AGG(status)"
type: string
dimensions:
- name: status
sql: "UPPER(status)"
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
measures: {
statuses: {
sql: `STRING_AGG(status)`,
type: `string`
}
},
dimensions: {
status: {
sql: `UPPER(status)`,
type: `string`
}
}
})
```
Currently, Cube does not wrap your SQL snippets in parentheses during SQL
generation. In case of non-trivial snippets, this may lead to unexpected results.
Please [track this issue](https://github.com/cube-js/cube/issues/6373).
### User-defined functions
If you have created a [user-defined function][link-sql-udf] (UDF) in your data
source, you can use it in the `sql` parameter as well.
### Case sensitivity
If your database uses case-sensitive identifiers, make sure to properly
quote table and column names. For example, here's how you can reference
a Postgres table that contains uppercase letters in its name:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: 'public."Orders"'
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `public."Orders"`
})
```
## References
To write versatile data models, it is important to be able to reference
members of cubes and views, such as measures or dimensions, as well as
table columns. Cube supports the following syntax for references.
### `column`
Most commonly, you would use bare column names in the `sql` parameter of
measures or dimensions. In the following example, `name` references
the respective column of the `users` table.
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
sql_table: users
dimensions:
- name: name
sql: name
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
sql_table: `users`,
dimensions: {
name: {
sql: `name`,
type: `string`
}
}
})
```
This syntax works great for simple use cases. However, if your cubes have
joins and joined cubes have columns with the same name, the generated SQL
query might become ambiguous. See below how to work around that.
### `{member}`
When defining measures and dimensions, you can also reference other members
of the same cube by wrapping their names in curly braces. In the following
example, the `full_name` dimension references `name` and `surname` dimensions
of the same cube.
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
sql_table: users
dimensions:
- name: name
sql: name
type: string
- name: surname
sql: "UPPER(surname)"
type: string
- name: full_name
sql: "CONCAT({name}, ' ', {surname})"
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
sql_table: `users`,
dimensions: {
name: {
sql: `name`,
type: `string`
},
surname: {
sql: `UPPER(surname)`,
type: `string`
},
full_name: {
sql: `CONCAT(${name}, ' ', ${surname})`,
type: `string`
}
}
})
```
You can safely reference a member whose `sql` is a compound expression, such as
arithmetic (`{price} + {tax}`) or boolean logic (`{is_paid} OR {is_pending}`). When you
reference such a member inside an arithmetic or logical expression, Cube wraps it in
parentheses automatically, so operator precedence is preserved. For example, if
`price_with_tax` is defined as `{price} + {tax}`, then referencing it as
`{price_with_tax} * {quantity}` generates `(price + tax) * quantity`. Cube only adds these
parentheses where they are needed — a reference that already sits in a safe position, such as
a function argument (`ABS({price_with_tax})`) or a `CAST(...)`, is left unwrapped.
This syntax works great for simple use cases. However, there are cases
(like [subquery][ref-subquery]) when you'd like to reference members of
other cubes. See below how to do that.
### `{time_dimension.granularity}`
When referencing a [time dimension][ref-time-dimension], you can specificy a
granularity to refer to a time value with that specific granularity. It can be
one of the [default granularities][ref-default-granularities] (e.g., `year` or
`week`) or a [custom granularity][ref-custom-granularities]:
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
sql_table: users
dimensions:
- name: created_at
sql: created_at
type: time
granularities:
- name: sunday_week
interval: 1 week
offset: -1 day
- name: created_at__year
sql: "{created_at.year}"
type: time
- name: created_at__sunday_week
sql: "{created_at.sunday_week}"
type: time
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
sql_table: `users`,
dimensions: {
created_at: {
sql: `created_at`,
type: `time`,
granularities: {
sunday_week: {
interval: `1 week`,
offset: `-1 day`
}
}
},
created_at__year: {
sql: `${created_at.year}`,
type: `time`
},
created_at__sunday_week: {
sql: `${created_at.sunday_week}`,
type: `time`
}
}
})
```
### `{cube}.column`, `{cube.member}`
You can qualify column and member names with the name of a cube to remove
the ambiguity when cubes are joined and reference members of other cubes.
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
sql_table: users
joins:
- name: contacts
sql: "{users}.contact_id = {contacts.id}"
relationship: one_to_one
dimensions:
- name: id
sql: "{users}.id"
type: number
primary_key: true
- name: name
sql: "COALESCE({users.name}, {contacts.name})"
type: string
- name: contacts
sql_table: contacts
dimensions:
- name: id
sql: "{contacts}.id"
type: number
primary_key: true
- name: name
sql: "{contacts}.name"
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
sql_table: `users`,
joins: {
contacts: {
sql: `${users}.contact_id = ${contacts.id}`,
relationship: `one_to_one`
}
},
dimensions: {
id: {
sql: `${users}.id`,
type: `number`,
primary_key: true
},
name: {
sql: `COALESCE(${users}.name, ${contacts.name})`,
type: `string`
}
}
})
cube(`contacts`, {
sql_table: `contacts`,
dimensions: {
id: {
sql: `${contacts}.id`,
type: `number`,
primary_key: true
},
name: {
sql: `${contacts}.name`,
type: `string`
}
}
})
```
In production, using fully-qualified names is generally encouraged since it
removes the ambiguity and keeps data model code maintainable as it grows.
However, always referring to the current cube by its name leads to code
repetition and violates the DRY principle. See below how to solve that.
### `{cube1.cube2.member}`
You can also qualify member names with more than one cube name, separated by a dot, to
provide a [join path][ref-join-paths] and remove the ambiguity of join resolution. Join
paths can be used in [calculated members][ref-calculated-members], [views][ref-views],
and [pre-aggregation][ref-preaggs] definitions.
In the following example, four cubes are joined together, forming a [*diamond
subgraph*][ref-diamond-subgraphs] in the join tree: `a` joins to `b` and `c`, and both
`b` and `c` join to `d`. As you can see, join paths in the `d_via_b` and `d_via_c`
dimensions of `a` are then used to specify which intermediate cube to use when resolving
the dimension from `d`.
```yaml title="YAML" theme={"dark"}
cubes:
- name: a
sql: |
SELECT 1 AS id UNION ALL
SELECT 2 AS id UNION ALL
SELECT 3 AS id
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: d_via_b
sql: "{b.d.id}"
type: number
- name: d_via_c
sql: "{c.d.id}"
type: number
joins:
- name: b
sql: "{a.id} = {b.id}"
relationship: one_to_one
- name: c
sql: "{a.id} = {c.id}"
relationship: one_to_one
- name: b
sql: |
SELECT 1 AS id UNION ALL
SELECT 2 AS id UNION ALL
SELECT 3 AS id
dimensions:
- name: id
sql: id
type: number
primary_key: true
joins:
- name: d
sql: "{b.id} = {d.id}"
relationship: one_to_one
- name: c
sql: |
SELECT 1 AS id UNION ALL
SELECT 2 AS id UNION ALL
SELECT 3 AS id
dimensions:
- name: id
sql: id
type: number
primary_key: true
joins:
- name: d
sql: "{c.id} = {d.id}"
relationship: one_to_one
- name: d
sql: |
SELECT 1 AS id UNION ALL
SELECT 2 AS id UNION ALL
SELECT 3 AS id
dimensions:
- name: id
sql: id
type: number
primary_key: true
```
```javascript title="JavaScript" theme={"dark"}
cube(`a`, {
sql: `
SELECT 1 AS id UNION ALL
SELECT 2 AS id UNION ALL
SELECT 3 AS id
`,
dimensions: {
id: {
sql: `id`,
type: `number`,
primary_key: true
},
d_via_b: {
sql: `${b.d.id}`,
type: `number`
},
d_via_c: {
sql: `${c.d.id}`,
type: `number`
}
},
joins: {
b: {
sql: `${a.id} = ${b.id}`,
relationship: `one_to_one`
},
c: {
sql: `${a.id} = ${c.id}`,
relationship: `one_to_one`
}
}
})
cube(`b`, {
sql: `
SELECT 1 AS id UNION ALL
SELECT 2 AS id UNION ALL
SELECT 3 AS id
`,
dimensions: {
id: {
sql: `id`,
type: `number`,
primary_key: true
}
},
joins: {
d: {
sql: `${b.id} = ${d.id}`,
relationship: `one_to_one`
}
}
})
cube(`c`, {
sql: `
SELECT 1 AS id UNION ALL
SELECT 2 AS id UNION ALL
SELECT 3 AS id
`,
dimensions: {
id: {
sql: `id`,
type: `number`,
primary_key: true
}
},
joins: {
d: {
sql: `${c.id} = ${d.id}`,
relationship: `one_to_one`
}
}
})
cube(`d`, {
sql: `
SELECT 1 AS id UNION ALL
SELECT 2 AS id UNION ALL
SELECT 3 AS id
`,
dimensions: {
id: {
sql: `id`,
type: `number`,
primary_key: true
}
}
})
```
### `{CUBE}` variable
You can use a handy `{CUBE}` [context variable][ref-context-variables]
(mind the uppercase) to reference the current cube so you don't have to
repeat the its name over and over. It works both for column and member
references.
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
sql_table: users
joins:
- name: contacts
sql: "{CUBE}.contact_id = {contacts.id}"
relationship: one_to_one
dimensions:
- name: id
sql: "{CUBE}.id"
type: number
primary_key: true
- name: name
sql: "COALESCE({CUBE.name}, {contacts.name})"
type: string
- name: contacts
sql_table: contacts
dimensions:
- name: id
sql: "{CUBE}.id"
type: number
primary_key: true
- name: name
sql: "{CUBE}.name"
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
sql_table: `users`,
joins: {
contacts: {
sql: `${CUBE}.contact_id = ${contacts.id}`,
relationship: `one_to_one`
}
},
dimensions: {
id: {
sql: `${CUBE}.id`,
type: `number`,
primary_key: true
},
name: {
sql: `COALESCE(${CUBE}.name, ${contacts.name})`,
type: `string`
}
}
})
cube(`contacts`, {
sql_table: `contacts`,
dimensions: {
id: {
sql: `${CUBE}.id`,
type: `number`,
primary_key: true
},
name: {
sql: `${CUBE}.name`,
type: `string`
}
}
})
```
Check the `{users.name}` dimension. Referencing another cube in the
dimension definition instructs Cube to make an implicit join to that cube.
For example, using the data model above, we can make the following query:
```json theme={"dark"}
{
"dimensions": ["users.name"]
}
```
The resulting generated SQL query would look like this:
```sql theme={"dark"}
SELECT COALESCE("users".name, "contacts".name) "users__name"
FROM users "users"
LEFT JOIN contacts "contacts"
ON "users".contact_id = "contacts".id
```
### `{cube.sql()}` function
When defining a cube, you can reference the `sql` parameter of another cube,
effectively reusing the SQL query it's defined on. This is particularly useful
when defining [polymorphic cubes][ref-polymorphism] or using [data blending][ref-data-blending].
Consider the following data model:
```yaml title="YAML" theme={"dark"}
cubes:
- name: organisms
sql_table: organisms
- name: animals
sql: |
SELECT *
FROM {organisms.sql()}
WHERE kingdom = 'animals'
- name: dogs
sql: |
SELECT *
FROM {animals.sql()}
WHERE species = 'dogs'
measures:
- name: count
type: count
```
```javascript title="JavaScript" theme={"dark"}
cube(`organisms`, {
sql_table: `organisms`
})
cube(`animals`, {
sql: `
SELECT *
FROM ${organisms.sql()}
WHERE kingdom = 'animals'
`
})
cube(`dogs`, {
sql: `
SELECT *
FROM ${animals.sql()}
WHERE species = 'dogs'
`,
measures: {
count: {
type: `count`
}
}
})
```
If you query for `dogs.count`, Cube will generate the following SQL:
```sql theme={"dark"}
SELECT count(*) "dogs__count"
FROM (
SELECT *
FROM (
SELECT *
FROM organisms
WHERE kingdom = 'animals'
)
WHERE species = 'dogs'
) AS "dogs"
```
### Curly braces and escaping
As you can see in the examples above, within [SQL expressions][self-sql-expressions],
curly braces are used to reference cubes and members.
In YAML data models, use `{reference}`:
```yaml theme={"dark"}
cubes:
- name: orders
sql: |
SELECT id, created_at
FROM {other_cube.sql()}
dimensions:
- name: status
sql: status
type: string
- name: status_x2
sql: "{status} || ' ' || {status}"
type: string
```
In JavaScript data models, use `${reference}` in [JavaScript template
literals][link-js-template-literals] (mind the dollar sign):
```javascript theme={"dark"}
cube(`orders`, {
sql: `
SELECT id, created_at
FROM ${other_cube.sql()}
`,
dimensions: {
status: {
sql: `status`,
type: `string`
},
status_x2: {
sql: `${status} || ' ' || ${status}`,
type: `string`
}
}
})
```
If you need to use literal, non-referential curly braces in YAML, e.g.,
to define a JSON object, you can escape them with a backslash:
```yaml theme={"dark"}
cubes:
- name: json_object_in_postgres
sql: SELECT CAST('\{"key":"value"\}'::JSON AS TEXT) AS json_column
- name: csv_from_s3_in_duckdb
sql: |
SELECT *
FROM read_csv(
's3://bbb/aaa.csv',
delim = ',',
header = true,
columns=\{'time':'DATE','count':'NUMERIC'\}
)
```
### Non-SQL references
Outside [SQL expressions][self-sql-expressions], `column` is not recognized
as a column name; it is rather recognized as a member name. It means that,
outside `sql` and `sql_table` parameters, you can skip the curly braces and
reference members by their names directly: `member`, `cube_name.member`, or
`CUBE.member`.
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
dimensions:
- name: status
sql: status
type: string
measures:
- name: count
type: count
pre_aggregations:
- name: orders_by_status
dimensions:
- CUBE.status
measures:
- CUBE.count
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
dimensions: {
status: {
sql: `status`,
type: `string`
}
},
measures: {
count: {
type: `count`
}
},
pre_aggregations: {
orders_by_status: {
dimensions: [CUBE.status],
measures: [CUBE.count]
}
}
})
```
## Context variables
In addition to the `CUBE` variable, you can also use a few more [context
variables][ref-context-variables] within your data model. They are
generally useful for two purposes: optimizing generated SQL queries and
defining dynamic data models.
## Troubleshooting
### `Can't parse timestamp`
Sometimes, you might come across the following error message: `Can't parse timestamp:
2023-11-07T14:33:23.16.000`.
It indicates that the data source was unable to recognize the value of a [time
dimension][ref-time-dimension] as a timestamp. Please check that the [SQL
expression](#sql-expressions) of this time dimension evaluates to a `TIMESTAMP` type.
Check [this recipe][ref-recipe-string-time-dimensions] to see how you can work around
string values in time dimensions.
[self-folder-structure]: #folder-structure
[self-naming]: #naming
[self-syntax]: #code-syntax
[self-sql-expressions]: #sql-expressions
[ref-dynamic-data-models]: /docs/data-modeling/dynamic
[ref-dynamic-data-models-jinja]: /docs/data-modeling/dynamic/jinja
[ref-dynamic-data-models-js]: /docs/data-modeling/dynamic/javascript
[ref-context-variables]: /reference/data-modeling/context-variables
[ref-config-model-path]: /reference/configuration/config#schemapath
[ref-config-repository-factory]: /reference/configuration/config#repositoryfactory
[ref-subquery]: /docs/data-modeling/dimensions#subquery-dimensions
[wiki-snake-case]: https://en.wikipedia.org/wiki/Snake_case
[wiki-yaml]: https://en.wikipedia.org/wiki/YAML
[link-snowflake-listagg]: https://docs.snowflake.com/en/sql-reference/functions/listagg
[link-bigquery-stringagg]: https://cloud.google.com/bigquery/docs/reference/standard-sql/functions-and-operators#string_agg
[link-sql-udf]: https://en.wikipedia.org/wiki/User-defined_function#Databases
[ref-time-dimension]: /reference/data-modeling/dimensions#type
[ref-default-granularities]: /docs/data-modeling/dimensions#time-dimensions
[ref-custom-granularities]: /reference/data-modeling/dimensions#granularities
[ref-style-guide]: /recipes/data-modeling/style-guide
[ref-polymorphism]: /recipes/data-modeling/polymorphic-cubes
[ref-data-blending]: /docs/data-modeling/concepts/data-blending
[link-js-template-literals]: https://developer.mozilla.org/en-US/docs/Learn_web_development/Core/Scripting/Strings#embedding_javascript
[link-python-reserved-words]: https://docs.python.org/3/reference/lexical_analysis.html#keywords
[ref-time-dimension]: /docs/data-modeling/dimensions#time-dimensions
[ref-recipe-string-time-dimensions]: /recipes/data-modeling/string-time-dimensions
[ref-view-groups]: /reference/data-modeling/view-group
[ref-views]: /docs/data-modeling/views
[ref-preaggs]: /reference/data-modeling/pre-aggregations
[ref-join-paths]: /docs/data-modeling/joins#join-paths
[ref-calculated-members]: /docs/data-modeling/measures#calculated-measures
[ref-diamond-subgraphs]: /docs/data-modeling/joins#diamond-subgraphs
# Configuration
Source: https://docs.cube.dev/docs/data-modeling/configuration
Configure model-level query defaults for a Cube deployment — query time zone, row limits, and query timeout.
The **Model Configuration** section in deployment settings is where you set
model-level defaults that every query in a [deployment][ref-deployment-types]
inherits unless it overrides them explicitly: the query time zone, default
and maximum row limits, and the per-query timeout.
To open it, go to **Settings** → **Configuration** in the deployment
sidebar.
Model Configuration is a curated, typed UI for the most common
model-level environment variables. Anything set here can also be set
through the [Environment Variables][ref-env-vars] page directly.
## Default time zone
The time zone used to interpret and return time dimensions when a query
doesn't specify one. Accepts any TZ database name, such as
`America/Los_Angeles` or `Europe/Berlin`. Defaults to `UTC`.
Backed by the
[`CUBEJS_DEFAULT_TIMEZONE`](/reference/configuration/environment-variables#cubejs_default_timezone)
environment variable. See [time zone][ref-time-zone] in the queries
reference for how Cube applies it to time dimensions and date ranges.
## Default row limit
The [row limit][ref-row-limit] applied to a query when it doesn't include
an explicit `LIMIT`. Defaults to `10,000`.
Lower this if most of your dashboards aggregate to a small number of rows
and you want to catch run-away "select everything" queries earlier. Raise
it if you're frequently truncating legitimate result sets.
Backed by
[`CUBEJS_DB_QUERY_DEFAULT_LIMIT`](/reference/configuration/environment-variables#cubejs_db_query_default_limit).
## Maximum row limit
The hard cap on row count for any query. Any explicit `LIMIT` is reduced
to this value, regardless of what the client requested. Defaults to
`50,000`.
This is the guardrail that protects the deployment from out-of-memory
crashes and accidental full-table extracts.
Raising the maximum row limit can cause out-of-memory crashes and makes
the deployment more vulnerable to denial-of-service attacks if the APIs
are exposed to untrusted clients. Increase it only when you have a
specific use case that requires it.
Backed by
[`CUBEJS_DB_QUERY_LIMIT`](/reference/configuration/environment-variables#cubejs_db_query_limit).
[SQL API][ref-sql-api] queries running in streaming mode can exceed this
cap by design.
## Query timeout
The per-query timeout applied to the upstream data source. Accepts a
duration string (`10m`, `30s`, `2h`) or a plain number of seconds.
Defaults to `10m`.
If a query exceeds this timeout, Cube cancels it and returns an error to
the client. Tune this based on the slowest legitimate query you expect to
run against the data source.
Backed by
[`CUBEJS_DB_QUERY_TIMEOUT`](/reference/configuration/environment-variables#cubejs_db_query_timeout).
## Auto-run
The deployment-wide default for whether [Explore][ref-explore] and
[Workbooks][ref-workbooks] tabs run their query automatically as you build it,
or wait for you to click **Run query**. Defaults to on.
A semantic view's [`meta.auto_run`][ref-view-auto-run] setting and a user's
own per-tab toggle both take precedence over this deployment default.
Backed by
[`CUBEJS_AUTO_RUN_MODE`](/reference/configuration/environment-variables#cubejs_auto_run_mode).
## Applying changes
Saving Model Configuration restarts the deployment's development-mode
worker so the new values take effect immediately for the
[development environment][ref-environments-dev]. Production environments
pick up the changes on the next build or deploy.
[ref-deployment-types]: /admin/deployment/deployment-types
[ref-env-vars]: /reference/configuration/environment-variables
[ref-time-zone]: /reference/core-data-apis/queries#time-zone
[ref-row-limit]: /reference/core-data-apis/queries#row-limit
[ref-sql-api]: /reference/core-data-apis/sql-api
[ref-environments-dev]: /admin/deployment/environments#development-environments
[ref-explore]: /docs/explore-analyze/explore
[ref-workbooks]: /docs/explore-analyze/workbooks/querying-data
[ref-view-auto-run]: /reference/data-modeling/view#auto_run
# Content Validator
Source: https://docs.cube.dev/docs/data-modeling/content-validator
Find workbook reports that reference data model fields or views that no longer exist, and fix or remove them before a model change breaks live dashboards.
The Content Validator finds workbook reports whose queries reference data model
members — individual fields or whole views — that no longer exist in your
current data model. Use it to catch and fix broken content while you are
changing the model, before the change ships and breaks live
[dashboards][ref-dashboards].
## Opening the Content Validator
In Cube Cloud, navigate to the **Data Model** module and select the
**Content Validator** tab.
Validation runs in real time, right when you open the tab — there is no
background job to schedule or wait for.
## How validation works
The Content Validator checks every [workbook][ref-workbooks] in the
deployment. For each workbook, it validates the workbook's draft reports as
well as the report snapshots of its published [dashboard][ref-dashboards].
Each report's query is checked with the same query parser the report editor
uses, so a reference flagged here is exactly the broken reference a viewer
would hit when opening that report.
Validation is branch-aware: reports are checked against the data model of
whatever context you are currently in — your
[development mode][ref-dev-mode], a shared branch, or production. The page
header names the context you are validating against, so you always know which
version of the model the findings apply to.
The Content Validator uses the full data model, ignoring visibility settings
and access policies. A report referencing a member that is hidden with
[`public: false`][ref-cube-public] or restricted by an
[access policy][ref-access-policies] is **not** flagged — the member still
exists in the model. Only references to members that are truly gone are
reported.
## Reviewing findings
Each broken reference is one row in the findings table:
| Column | What it shows |
| ----------- | ----------------------------------------------------------------------------------------- |
| **Content** | The report name and the workbook it belongs to. |
| **State** | Whether the finding is in a **Draft** report or a **Published** dashboard snapshot. |
| **Field** | The broken member reference, e.g., `orders.total`. |
| **Error** | The parser error, e.g., `No such field "orders.total"` or `No such view "sessions_view"`. |
| **Status** | The state of the reference — **Missing**. |
There are two kinds of findings:
* **Missing field** — the report references a member (a measure or a
dimension) that no longer exists in an otherwise valid view.
* **Missing view** — the report queries a view that no longer exists in the
data model.
Use the search box above the table to filter findings by report, workbook, or
member name, and the **All** / **Draft** / **Published** filter to focus on
one kind of content.
If every reference resolves, the page shows *No broken references found*.
## Fixing broken references
How you fix a finding depends on whether it is in a draft report or a
published dashboard snapshot:
* **Draft reports** are editable in place. Each draft finding offers
**Open**, **Replace**, and **Delete** actions.
* **Published dashboards** are immutable snapshots — they only change when
their dashboard is republished. Published findings offer only an **Open**
action; to fix one, update the draft report in its workbook and republish
the dashboard.
### Open
**Open** navigates to the content itself: the report in its workbook for a
draft finding, or the live published dashboard for a published finding.
### Replace
**Replace** swaps the broken reference for an existing one across the
report's saved query. Clicking it opens a dialog with a searchable picker of
valid replacements — members of the same view for a missing field, or other
views for a missing view.
* If several draft reports share the same broken reference, the dialog offers
a **Replace in all affected reports** checkbox (enabled by default) so you
can fix them all in one action.
* If a published dashboard references the same member, the dialog notes that
it will keep its current snapshot until the dashboard is republished.
When you apply a replacement, Cube rewrites each affected report's saved
query, persists it, and re-runs validation, so fixed rows clear from the
table immediately. The rewrite is precise: only references to the broken
member are updated, while same-named members of other views are left
untouched.
### Delete
**Delete** permanently removes the report from its workbook — use it for
reports that are no longer worth fixing. A confirmation dialog is shown
before anything is deleted, and validation re-runs afterwards.
## Scope and limitations
* The Content Validator validates reports whose queries are stored as
semantic SQL. A report whose query the parser cannot read is counted as
*not checked* — it is neither flagged as broken nor asserted to be clean.
When this happens, the summary shows a *N reports could not be checked*
note.
* Validation covers each workbook's draft reports and its published dashboard
snapshots.
* **Replace** and **Delete** act on draft reports only. Fixing published
content is a two-step process: fix the draft, then republish the dashboard.
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-dashboards]: /docs/explore-analyze/dashboards
[ref-dev-mode]: /docs/data-modeling/dev-mode
[ref-cube-public]: /reference/data-modeling/cube#public
[ref-access-policies]: /docs/data-modeling/data-access-policies
# Cubes
Source: https://docs.cube.dev/docs/data-modeling/cubes
Cubes represent the tables in your database. Each cube maps to a table or query in your data source and contains measures, dimensions, joins, and pre-aggregations.
Cubes represent the tables in your database. Each cube maps to a single
table in your [data source][ref-data-sources] and contains the business
logic — [measures][ref-measures], [dimensions][ref-dimensions],
[joins][ref-joins], and [pre-aggregations][ref-pre-aggs] — that defines
how that data can be queried.
See the [cube reference][ref-cube-reference] for the full list of
parameters and configuration options.
## Defining a cube
A cube points to a table in your data source using `sql_table`:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`
})
```
You can also use the `sql` property for more complex queries:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql: |
SELECT *
FROM orders, line_items
WHERE orders.id = line_items.order_id
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql: `
SELECT *
FROM orders, line_items
WHERE orders.id = line_items.order_id
`
})
```
If you're using dbt, see [this recipe][ref-cube-with-dbt] to streamline
defining cubes on top of dbt models.
## Cube members
Each cube contains definitions for its members: dimensions, measures,
and segments.
### Dimensions
[Dimensions][ref-dimensions] represent the properties of a single data
point — the attributes you group by and filter on, such as `status`,
`city`, or `created_at`:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: status
sql: status
type: string
- name: created_at
sql: created_at
type: time
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
dimensions: {
id: {
sql: `id`,
type: `number`,
primary_key: true
},
status: {
sql: `status`,
type: `string`
},
created_at: {
sql: `created_at`,
type: `time`
}
}
})
```
Time dimensions enable grouping by granularity (year, quarter, month,
week, day, hour, minute, second) and are essential for
[partitioned pre-aggregations][ref-partition-preaggs].
### Measures
[Measures][ref-measures] represent aggregated values over a set of data
points — counts, sums, averages, and custom calculations:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
measures:
- name: count
type: count
- name: total_amount
sql: amount
type: sum
- name: average_amount
sql: amount
type: avg
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
measures: {
count: {
type: `count`
},
total_amount: {
sql: `amount`,
type: `sum`
},
average_amount: {
sql: `amount`,
type: `avg`
}
}
})
```
Measures can reference other measures to create
[calculated measures][ref-calculated-measures], and you can apply
[filters][ref-measure-filters] to create filtered aggregations like
"count of completed orders."
### Segments
[Segments][ref-segments] are predefined filters on a cube. They allow
you to define commonly used filter logic once and reuse it across
queries:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
segments:
- name: completed
sql: "{CUBE}.status = 'completed'"
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
segments: {
completed: {
sql: `${CUBE}.status = 'completed'`
}
}
})
```
## Joins
[Joins][ref-joins] define relationships between cubes, forming the data
graph that Cube uses to generate multi-table SQL queries:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
joins:
- name: users
relationship: many_to_one
sql: "{CUBE}.user_id = {users.id}"
- name: line_items
relationship: one_to_many
sql: "{CUBE}.id = {line_items.order_id}"
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
joins: {
users: {
relationship: `many_to_one`,
sql: `${CUBE}.user_id = ${users.id}`
},
line_items: {
relationship: `one_to_many`,
sql: `${CUBE}.id = ${line_items.order_id}`
}
}
})
```
Cube supports `one_to_one`, `many_to_one`, and `one_to_many` relationship
types. See [working with joins][ref-working-with-joins] for advanced
patterns like cross-database joins and join direction control.
## Pre-aggregations
[Pre-aggregations][ref-pre-aggs] are materialized summaries of cube data
that dramatically speed up query execution. Cube automatically matches
incoming queries to the best available pre-aggregation:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
pre_aggregations:
- name: main
measures:
- count
- total_amount
dimensions:
- status
time_dimension: created_at
granularity: day
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
pre_aggregations: {
main: {
measures: [count, total_amount],
dimensions: [status],
time_dimension: created_at,
granularity: `day`
}
}
})
```
Pre-aggregations support [partitioning][ref-partition-preaggs] by time
and [incremental refreshes][ref-incremental-preaggs] to keep materialized
data up-to-date efficiently.
## Designing effective cubes
### One cube per entity
Map each cube to a single business entity — `orders`, `users`,
`products`, `line_items`. Use [joins](#joins) to connect them rather than
creating wide cubes with data from multiple tables.
### Keep naming consistent
Use clear, consistent naming for members. Dimensions should describe
attributes (`status`, `city`, `created_at`), and measures should describe
aggregations (`count`, `total_revenue`, `average_order_value`). Add
[`description`][ref-cube-description] and [`title`][ref-cube-title] for
user-friendly display.
### Control visibility
Use [`public`][ref-cube-public] to hide cubes that should not be
directly queried by end-users. In most data models, cubes are internal
building blocks and [views][ref-views] are the public interface:
```yaml title="YAML" theme={"dark"}
cubes:
- name: base_orders
public: false
sql_table: orders
# ...
```
```javascript title="JavaScript" theme={"dark"}
cube(`base_orders`, {
public: false,
sql_table: `orders`,
// ...
})
```
### Scale with extension and polymorphism
When cubes share common members, use [`extends`][ref-extending-cubes] to
avoid duplication. For data models with many similar entities,
[polymorphic cubes][ref-polymorphic-cubes] let you define a base cube
and specialize it per entity.
## Next steps
* See the [cube reference][ref-cube-reference] for the full list of
parameters
* Learn about [views][ref-views] to expose cubes to end-users
* Explore [calculated measures][ref-calculated-measures] for derived metrics
* Use the [Semantic Model IDE][ref-ide] to develop cubes interactively
[wiki-view-sql]: https://en.wikipedia.org/wiki/View_\(SQL\)
[ref-data-sources]: /admin/connect-to-data
[ref-cube-reference]: /reference/data-modeling/cube
[ref-cube-description]: /reference/data-modeling/cube#description
[ref-cube-title]: /reference/data-modeling/cube#title
[ref-cube-public]: /reference/data-modeling/cube#public
[ref-measures]: /reference/data-modeling/measures
[ref-dimensions]: /reference/data-modeling/dimensions
[ref-segments]: /reference/data-modeling/segments
[ref-joins]: /reference/data-modeling/joins
[ref-pre-aggs]: /reference/data-modeling/pre-aggregations
[ref-views]: /docs/data-modeling/views
[ref-extending-cubes]: /docs/data-modeling/extending-cubes
[ref-polymorphic-cubes]: /recipes/data-modeling/polymorphic-cubes
[ref-calculated-measures]: /docs/data-modeling/measures#calculated-measures
[ref-measure-filters]: /reference/data-modeling/measures#filters
[ref-working-with-joins]: /docs/data-modeling/joins
[ref-partition-preaggs]: /docs/pre-aggregations/matching-pre-aggregations#partitioning
[ref-incremental-preaggs]: /reference/data-modeling/pre-aggregations#incremental
[ref-cube-with-dbt]: /recipes/data-modeling/dbt
[ref-ide]: /docs/data-modeling/data-model-ide
# Access policies
Source: https://docs.cube.dev/docs/data-modeling/data-access-policies
Declares group-scoped policies that combine member access, row filters, and masking rules directly in the data model.
Access policies provide a holistic mechanism to manage [member-level](#member-level-access),
[row-level](#row-level-access) security, and [data masking](#data-masking) for
different user groups. You can define access control rules in data model files,
allowing for an organized and maintainable approach to security.
## Policies
You can define policies that target specific groups and contain member-level and (or)
row-level security rules:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
access_policy:
# For the `manager` group,
# allow access to all members
# but filter rows by the user's country
- group: manager
member_level:
includes: "*"
row_level:
filters:
- member: country
operator: equals
values: [ "{ userAttributes.country }" ]
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
access_policy: [
{
// For all groups, restrict access entirely
group: `*`,
member_level: {
includes: []
}
},
{
// For the `manager` group,
// allow access to all members
// but filter rows by the user's country
group: `manager`,
member_level: {
includes: `*`
},
row_level: {
filters: [
{
member: `country`,
operator: `equals`,
values: [ userAttributes.country ]
}
]
}
}
]
})
```
While you can define access policies on both cubes and views, it is more common to define them on views.
For more details on available parameters, check out the [access policies reference][ref-ref-dap].
## Policy evaluation
When processing a request, Cube will evaluate the access policies and combine them
with relevant custom security rules, e.g., [`public` parameters][ref-mls-public] for member-level security
and `query_rewrite` filters for row-level security.
### The permission space
It helps to think of access control as a **two-dimensional permission space** —
a grid of **members** (the columns a user may query: dimensions and measures)
and **rows** (the records a user may see):
* one axis is **members** — *what* a user can look at;
* the other axis is **rows** — *which* records they can look at.
Each access policy grants visibility over a rectangular region of this space:
its `member_level` chooses the members (the horizontal extent) and its
`row_level` chooses the rows (the vertical extent). Defaults *widen* the region —
a policy with no `row_level` (or with `row_level: { allow_all: true }`) spans
**every row**, and a policy with no `member_level` spans **every member**.
`member_masking` marks part of a region as **visible but masked** rather than
fully readable.
A user usually matches more than one policy (for example, through multiple
[groups](#custom-mapping)), so their effective access is the combination of
every region granted by every matching policy:
* **Members are unioned.** A member is accessible if **any** matching policy
grants it. A user who matches several policies sees every member those
policies expose, even when no single policy exposes all of them.
* **Rows are intersected across the queried members.** For each queried member,
the visible rows are the **union** of the row filters of the policies that
grant that member (a policy with no row filter adds no restriction). A row is
returned only when it is visible for **every** queried member.
* **A member is masked** when no granting policy gives it unconditional full
access through `member_level`, but some matching policy lists it under
`member_masking`.
* **Access is denied** (an empty result) only when a queried member is granted
by **no** matching policy at all.
#### Diagram and behavior
Consider an `orders_view` matched by two of a user's groups:
* the `support` group — `member_level: [status, count]`, `row_level` restricted
to `region = 'US'`;
* the `finance` group — `member_level: [count, revenue]`, `row_level` restricted
to `region = 'EU'`.
The two policies cover overlapping regions of the permission space. `count` sits
in the overlap (both policies grant it); `status` and `revenue` are each granted
by only one policy:
```text theme={"dark"}
members
▲
│ ┌───────────────────────────────────────┐
revenue │ │ finance policy │
│ ┌────────┼──────────────┐ │
count │ │ │ overlap │ │
│ │ └──────────────┼────────────────────────┘
status │ │ support policy │
│ └───────────────────────┘
└───────────────────────────────────────────────────▶ rows
US region EU region
```
For a user in **both** groups, the readable cells (✓) of the permission space are:
| | `US` rows | `EU` rows |
| --------- | :---------: | :---------: |
| `status` | ✓ (support) | — |
| `count` | ✓ (support) | ✓ (finance) |
| `revenue` | — | ✓ (finance) |
Because rows are intersected across the queried members, the visible rows depend
on *which* members the query selects:
| Queried members | How rows resolve | Visible rows |
| ---------------------------- | -------------------- | ------------------- |
| `status`, `count` | `US` ∩ (`US` ∪ `EU`) | `US` rows |
| `count`, `revenue` | (`US` ∪ `EU`) ∩ `EU` | `EU` rows |
| `count` | `US` ∪ `EU` | all rows |
| `status`, `count`, `revenue` | `US` ∩ `EU` | none (empty result) |
* Querying `status` and `count` returns only `US` rows: `status` is granted only
by the `support` policy, so records outside the US can never satisfy the query.
* Querying `count` alone returns all rows: both policies grant `count`, so its
visible rows are the union of the two regions.
* Querying `status`, `count`, and `revenue` returns nothing: `status` is visible
only on `US` rows and `revenue` only on `EU` rows, and no record is in both.
The result is empty rather than leaking US-only members onto EU rows.
A policy without a `row_level` filter defaults to **all rows** (allow-all). So
when every policy that grants the queried members is filter-less, there is no
row restriction at all — the members are simply unioned and all rows are
returned. Row filters only narrow the result when a granting policy defines them.
### Member-level access
Member-level security rules in access policies are *combined together*
with `public` parameters of cube and view members using the *AND* semantics.
Both will apply to the request.
*When querying a view,* member-level security rules defined in the view are ***not** combined together*
with member-level security rules defined in relevant cubes.
**Only the ones from the view will apply to the request.**
This is consistent with how column-level security works in SQL databases. If you have
a view that exposes a subset of columns from a table, it doesnt matter if the
columns in the table are public or not, the view will expose them anyway.
### Row-level access
Row-level filters in access policies are *combined together* with filters defined
using the `query_rewrite` configuration option.
Both will apply to the request.
*When querying a view,* row-level filters defined in the view are *combined together*
with row-level filters defined in relevant cubes. Both will apply to the request.
This is consistent with how row-level security works in SQL databases. If you have
a view that exposes a subset of rows from another view, the result set will be
filtered by the row-level security rules of both views.
### Data masking
With data masking, you can return masked values for restricted members instead
of denying access entirely. Users who don't have full access to a member will
see a transformed value (e.g., `***`, `-1`, `NULL`) rather than receiving an error.
To use data masking, define a [`mask` parameter][ref-ref-mask-dim] on dimensions
or measures, and add `member_masking` to your access policy alongside `member_level`.
Members in `member_level` get real values; members not in `member_level` but in
`member_masking` get masked values; members in neither are denied.
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
dimensions:
- name: status
sql: status
type: string
- name: secret_code
sql: secret_code
type: string
mask:
sql: "CONCAT('***', RIGHT({CUBE}.secret_code, 3))"
- name: revenue
sql: revenue
type: number
mask: -1
measures:
- name: count
type: count
mask: 0
access_policy:
- group: manager
member_level:
includes:
- status
- count
member_masking:
includes: "*"
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
dimensions: {
status: {
sql: `status`,
type: `string`
},
secret_code: {
sql: `secret_code`,
type: `string`,
mask: {
sql: `CONCAT('***', RIGHT(${CUBE}.secret_code, 3))`
}
},
revenue: {
sql: `revenue`,
type: `number`,
mask: -1
}
},
measures: {
count: {
type: `count`,
mask: 0
}
},
access_policy: [
{
group: `manager`,
member_level: {
includes: [`status`, `count`]
},
member_masking: {
includes: `*`
}
}
]
})
```
With this policy, users in the `manager` group will see:
| Member | Value |
| ------------- | ------------------------------------------- |
| `status` | Real value (full access via `member_level`) |
| `count` | Real value (full access via `member_level`) |
| `secret_code` | Masked via SQL: `***xyz` |
| `revenue` | Masked: `-1` |
If no `mask` is defined on a member, the default mask value is `NULL`. You can
customize defaults with the `CUBEJS_ACCESS_POLICY_MASK_STRING`,
`CUBEJS_ACCESS_POLICY_MASK_NUMBER`, `CUBEJS_ACCESS_POLICY_MASK_BOOLEAN`, and
`CUBEJS_ACCESS_POLICY_MASK_TIME` environment variables.
SQL masks (`mask: { sql: "..." }`) on measures are not applied in ungrouped
queries (e.g., `SELECT *` via the SQL API), because SQL mask expressions
typically reference columns that are not meaningful in a per-row context.
Static masks (`mask: -1`, `mask: 0`) are applied in all cases.
If you need to mask a measure in ungrouped queries with a dynamic expression,
define it as a dimension with an SQL mask instead, and reference that masked
dimension in your query.
#### Masking across multiple policies
Because member access is [unioned](#the-permission-space), **full access wins
over masking**. If any matching policy grants a member unconditional full access
through `member_level` (with no `row_level` filter), the user sees the **real**
value — even if another matching policy lists that member under `member_masking`.
Masking only takes effect when **no** matching policy grants unconditional full
access. There are two sub-cases:
* **Masked only.** The member is exposed solely through `member_masking` (or any
full-access policy is itself row-restricted). The member is masked for **all**
rows.
* **Conditionally unmasked.** Another policy grants full access *and* defines a
`row_level` filter — full access is conditional on that filter. Masking then
becomes **conditional on the row filter**: rows matching the filter show the
real value, while the rest show the masked value. The generated SQL is roughly
`CASE WHEN {rowFilter} THEN {value} ELSE {mask} END`.
This lets you combine a broad masking policy (e.g. the `*` group sees masked
values) with a narrower policy that reveals real values only for the rows a
group is entitled to (its `row_level` range).
If the query itself already constrains rows to a subset of the conditional
mask's row filter — an equally or more restrictive filter on the same member
(for example, the query filters `country = 'US'` and the mask filter is
`country = 'US'`) — then every returned row would show the real value anyway.
In that case the `CASE WHEN` is unnecessary and the member is **unmasked**. This
also lets a conditionally-masked *aggregate measure* render its real value
instead of being masked, even without grouping by the filter's member.
#### Conditional masking on measures
Conditional masking is evaluated **per row**, which works naturally for
dimensions (the `CASE WHEN` expression is part of the `GROUP BY`). For an
**aggregate measure** (e.g. `sum`, `count`), that per-row expression can only be
applied when the members referenced by the row filter are part of the query's
`GROUP BY`.
When a query selects a conditionally-masked measure but **does not group by the
members referenced in the row filter**, Cube cannot decide the condition per
aggregated group. Rather than emit invalid SQL (where the filter column is
neither grouped nor aggregated — which fails on strict engines like BigQuery),
it renders the **mask value for the entire measure** (`NULL` by default) instead
of the conditional expression.
| Query groups by the row filter's members? | Result for the measure |
| ----------------------------------------- | ----------------------------------------------------------- |
| Yes | Conditional: real value for matching rows, masked otherwise |
| No | Fully masked (the mask value, e.g. `NULL`) |
If you need a row-aware value for a measure regardless of grouping, add the row
filter's member (the dimension it filters on) to your query's dimensions so it
becomes part of the `GROUP BY`.
*When querying a view,* data masking follows the same pattern as row-level
security: masking rules from both the view and relevant cubes are applied.
For more details on available parameters, check out the
[`member_masking` reference][ref-ref-dap-masking].
## Common patterns
### Restrict access to specific groups
To restrict access to a view to only specific groups, define access policies for those groups. Access is automatically denied to all other groups:
```yaml title="YAML" theme={"dark"}
views:
- name: sensitive_data_view
# ...
access_policy:
# Allow access only to the `analysts` group
- group: analysts
member_level:
includes: "*"
```
```javascript title="JavaScript" theme={"dark"}
view(`sensitive_data_view`, {
// ...
access_policy: [
{
// Allow access only to the `analysts` group
group: `analysts`,
member_level: {
includes: `*`
}
}
]
})
```
You can also use the `groups` parameter (plural) to apply the same policy to multiple groups at once:
```yaml title="YAML" theme={"dark"}
views:
- name: sensitive_data_view
# ...
access_policy:
# Allow access to multiple groups using groups array
- groups: [analysts, managers]
member_level:
includes: "*"
```
```javascript title="JavaScript" theme={"dark"}
view(`sensitive_data_view`, {
// ...
access_policy: [
{
// Allow access to multiple groups using groups array
groups: [`analysts`, `managers`],
member_level: {
includes: `*`
}
}
]
})
```
### Filter by user attribute
You can filter data based on user attributes to ensure users only see data they're authorized to access. For example, sales people can see only their own deals, while sales managers can see all deals:
```yaml title="YAML" theme={"dark"}
views:
- name: deals_view
# ...
access_policy:
# Sales people can only see their own deals
- group: sales
member_level:
includes: "*"
row_level:
filters:
- member: sales_person_id
operator: equals
values: [ "{ userAttributes.userId }" ]
# Sales managers can see all deals
- group: sales_manager
member_level:
includes: "*"
# No row-level filters - full access to all rows
```
```javascript title="JavaScript" theme={"dark"}
view(`deals_view`, {
// ...
access_policy: [
{
// Sales people can only see their own deals
group: `sales`,
member_level: {
includes: `*`
},
row_level: {
filters: [
{
member: `sales_person_id`,
operator: `equals`,
values: [ userAttributes.userId ]
}
]
}
},
{
// Sales managers can see all deals
group: `sales_manager`,
member_level: {
includes: `*`
}
// No row-level filters - full access to all rows
}
]
})
```
### Filter by multiple user attributes
You can pass multiple values in the `values` array to match a dimension against
more than one user attribute. This is useful when users may have access based on
multiple properties, such as a country and a custom country property:
```yaml title="YAML" theme={"dark"}
views:
- name: deals_view
# ...
access_policy:
- group: sales
member_level:
includes: "*"
row_level:
filters:
- member: users_country
operator: equals
values: [ "{ userAttributes.country }", "{ userAttributes.customCountryProperty }" ]
```
```javascript title="JavaScript" theme={"dark"}
view(`deals_view`, {
// ...
access_policy: [
{
group: `sales`,
member_level: {
includes: `*`
},
row_level: {
filters: [
{
member: `users_country`,
operator: `equals`,
values: [
userAttributes.country,
userAttributes.customCountryProperty
]
}
]
}
}
]
})
```
### Mask sensitive members
You can mask sensitive members for most users while granting full access to
privileged groups:
```yaml title="YAML" theme={"dark"}
views:
- name: orders_view
# ...
access_policy:
# Default: all members masked
- group: "*"
member_level:
includes: []
member_masking:
includes: "*"
# Admins: full access
- group: admin
member_level:
includes: "*"
```
```javascript title="JavaScript" theme={"dark"}
view(`orders_view`, {
// ...
access_policy: [
{
// Default: all members masked
group: `*`,
member_level: {
includes: []
},
member_masking: {
includes: `*`
}
},
{
// Admins: full access
group: `admin`,
member_level: {
includes: `*`
}
}
]
})
```
### Mandatory filters
You can apply mandatory row-level filters to specific groups to ensure they only see data matching certain criteria:
```yaml title="YAML" theme={"dark"}
views:
- name: country_data_view
# ...
access_policy:
# Allow access only to the `sales` and `marketing` groups with country filtering
- groups: [sales, marketing]
member_level:
includes: "*"
row_level:
filters:
- member: users_country
operator: equals
values: ["Brasil"]
```
```javascript title="JavaScript" theme={"dark"}
view(`country_data_view`, {
// ...
access_policy: [
{
// Allow access only to the `sales` and `marketing` groups with country filtering
groups: [`sales`, `marketing`],
member_level: {
includes: `*`
},
row_level: {
filters: [
{
member: `users_country`,
operator: `equals`,
values: [`Brasil`]
}
]
}
}
]
})
```
## Custom mapping
Cube cloud platform automatically maps authenticated users to groups for access policies.
If you are using Cube Core or authenticating against [Core Data APIs][ref-core-data-apis] directly, you might need to map the security context to groups manually.
```python title="Python" theme={"dark"}
# cube.py
from cube import config
@config('context_to_groups')
def context_to_groups(ctx: dict) -> list[str]:
return ctx['securityContext'].get('groups', ['default'])
```
```javascript title="JavaScript" theme={"dark"}
// cube.js
module.exports = {
contextToGroups: ({ securityContext }) => {
return securityContext.groups || ['default']
}
}
```
A user can have more than one group.
## Using securityContext
The [`userAttributes`][ref-sec-ctx] object is only available in Cube Cloud platform. If you are using Cube Core or authenticating against [Core Data APIs][ref-core-data-apis] directly, you won't have access to `userAttributes`. Instead, you need to use `securityContext` directly when referencing user attributes in access policies (e.g., in `row_level` filters or `conditions`). For example, use `securityContext.userId` instead of `userAttributes.userId`.
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
access_policy:
- group: manager
row_level:
filters:
- member: country
operator: equals
values: [ "{ securityContext.country }" ]
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
access_policy: [
{
group: `manager`,
row_level: {
filters: [
{
member: `country`,
operator: `equals`,
values: [ securityContext.country ]
}
]
}
}
]
})
```
[ref-mls-public]: /docs/data-modeling/access-control/member-level-security#managing-member-level-access
[ref-sec-ctx]: /docs/data-modeling/access-control/context
[ref-ref-dap]: /reference/data-modeling/data-access-policies
[ref-ref-dap-masking]: /reference/data-modeling/data-access-policies#member-masking
[ref-ref-mask-dim]: /reference/data-modeling/dimensions#mask
[ref-core-data-apis]: /reference/core-data-apis
# Data Model IDE
Source: https://docs.cube.dev/docs/data-modeling/data-model-ide
Covers the browser code editor for full data modeling features, branch-based dev APIs, and testing model changes before they reach production traffic.
Data model editor provides the code-first experience for building and enhancing the
[data model][ref-data-modeling] of your semantic layer from within your web browser.
Unlike the [Visual Modeler][ref-visual-model] editor, it provides the freedom to use all
available data modeling features at the expense of a code-centric experience.
Cube Cloud can create branch-based development API instances to quickly test
changes in the data model in your frontend applications before pushing them into
production.
## Development Mode
In development mode, you can safely make changes to your project without
affecting production deployment. Development mode uses a separate Git branch and
allows testing your changes in Playground or via a separate API endpoint
specific to this branch. This development API hot-reloads your data model
changes, allowing you to quickly test API changes from your applications.
To enter development mode, navigate to the **Data Model** screen and
click **Dev Mode**.
If you are entering development mode from the main branch, you will be asked to
choose an existing branch or create a new one. You can configure whether it's possible
to commit directly to the main branch in the deployment settings.
When development mode is active, the status of the development API will be shown
at the top of the screen. After any changes to the project, the API
will hot-reload, and the API status will indicate when it's ready.
You can exit development mode by clicking **Dev Mode X** in the purple banner. If
you've been editing a data model and navigate away, Cube Cloud will warn you if
there are any unsaved changes:
## Git integration
To add more Git branches to your Cube Cloud deployment and/or switch between
them, click the branch name in the status bar:
Speaking of Git branches, you can now easily add and remove branches with the
same switcher; click **Add Branch** and enter a name for the new branch
in the popup:
These branches are shared, meaning everyone who has access to the deployment can
see and edit them. This makes them extremely useful for out-of-band experiments
where you can quickly test things in Cube Cloud without having to go through a
CI/CD process.
Unused branches can also be deleted. Ensure you are already on the branch you
want to delete, then open the switcher and click **Remove Branch**:
To share a link to the branch you're currently viewing, click the copy-link
button next to the branch switcher. It copies the current URL with
`?branch=` appended; opening that link switches the recipient to
the same branch.
## Generating data model files
You can generate data model files when [creating a new
deployment][ref-creating-deployment] or at later moment. To open the data model
wizard, click on the **...** button on the right and select **Generate
data model**:
It is safe to generate data model files for tables that already have matching cubes.
In that case, existing files will be renamed by appending `.backup` to their names.
You would also be able to review changes to files on the **Changes** tab:
[ref-data-modeling]: /docs/data-modeling/overview
[ref-creating-deployment]: /admin/deployment#creating-a-new-deployment
[ref-visual-model]: /docs/data-modeling/visual-modeler
# Development mode
Source: https://docs.cube.dev/docs/data-modeling/dev-mode
Outlines development mode in Cube Cloud—branch-scoped APIs, save and commit flows, and safe iteration before promoting changes to production.
Development mode allows to test and debug the data model in an isolated
[development environment][ref-environments-dev] before releasing any changes to production.
This page describes development mode in the Cube cloud platform. It is unrelated to
Cube Core's development mode — which is on when
[`CUBEJS_DEV_MODE=true`](/reference/configuration/environment-variables#cubejs_dev_mode)
or whenever `NODE_ENV` is not `production` — and which is an **authentication bypass**.
Under the `cubejs` CLI and the official Docker images, setting the flag also forces
`NODE_ENV=development`, switching off JWT verification on the data APIs. Playground's
endpoints are served with no authentication at all, so
anyone who can reach the instance can mint an API token carrying any security context
(signed with the API secret), read the data model, and overwrite it and `.env`. With
`CUBEJS_DEV_MODE=true` and no
[`CUBEJS_SQL_PASSWORD`](/reference/configuration/environment-variables#cubejs_sql_password)
set, the SQL API accepts any credentials as well. That is intentional — it is designed
to run on a developer's local machine for ease of use and debugging. Never use it in
production, and using it in the Cube cloud platform is highly discouraged because it
bypasses the platform's security model.
When you enter the development mode, you'll have access to your personal API endpoints
that will track the branch you're on and will be updated automatically when you make
changes to the data model.
## Development flow
### Save changes
When you make changes to the data model, you can save them by clicking the "Save all" button in the actions bar at the top of the screen.
### Commit & Sync
After saving changes, you can commit them to the branch you're working on by clicking the "Commit & Sync" button in the actions bar.
### Merge or Create a Pull Request
After committing changes, you can merge them into the parent branch or create a pull request by clicking the "Merge" or "Create a Pull Request" button in the actions bar.
### Updating the branch
When there are changes in the parent branch, you can update your active branch by pulling changes from the parent branch.
If your active branch got updated in the connected repository, you can update it by pulling changes from the remote.
### Resolving conflicts
In some cases, you may encounter conflicts when merging or pulling changes from the parent branch or remote.
If this happens, you'll see a warning about conflicts in the actions bar.
You'll then be able to either resolve them manually or discard the changes.
## Deep-link into development mode
You can link directly into development mode on the Semantic Model IDE by
adding query parameters to the page URL:
* `?branch=` switches to the given branch, entering development mode
if you're already in it (or staying read-only otherwise).
* `?devmode=true` enters development mode on the current branch. Combine it
with `?branch=` to enter development mode on a specific branch in one step:
```text theme={"dark"}
https://your-tenant.cubecloud.dev/d/DEPLOYMENT_ID/schema?branch=my-branch&devmode=true
```
Both parameters are stripped from the URL once applied. They're a no-op for
users without permission to use development mode.
## Finding endpoints
In Development Mode, you'll have dedicated endpoints that will track the branch you're on.
They will be shown on the Overview page:
Read more about the available endpoints on the [Environments][ref-environments-endpoints] page.
## Behavior with CLI deployments
Development mode is fully available on deployments that are configured to
[deploy with CLI][ref-deploy-with-cli]: you can enter dev mode, switch
branches, save and commit changes, and merge into the production branch from
the Cube Cloud UI.
However, on CLI deployments, **none of these actions trigger a production
build or redeploy** — production only changes when somebody explicitly runs
`cubejs-cli deploy` against the deployment, and the next deploy overwrites
any changes made through the UI. See [Deploy with CLI →
Development mode][ref-deploy-with-cli-dev-mode] for the full behavior.
## Limitations
The Development API has some limitations compared to Production, and it is
important to be aware of these as they may affect your development process.
### Autoscaling
The Development API **does not** support autoscaling. If you need to test
autoscaling, we recommend creating a separate deployment using a different Git
branch of the same repository.
### Scheduled refresh
Scheduled refresh is also disabled in the Development API. If you need to test
scheduled refreshes, we recommend creating a separate deployment using a
different Git branch of the same repository.
[ref-environments-dev]: /admin/deployment/environments#development-environments
[ref-environments-endpoints]: /admin/deployment/environments#api-endpoints
[ref-deploy-with-cli]: /admin/deployment/continuous-deployment#deploy-with-cli
[ref-deploy-with-cli-dev-mode]: /admin/deployment/continuous-deployment#development-mode-with-cli-deployments
# Dimensions
Source: https://docs.cube.dev/docs/data-modeling/dimensions
Dimensions are attributes that describe individual rows of data — the fields you group by and filter on, such as status, city, or created_at.
Dimensions represent attributes of individual rows in your data. They are
the fields you group by and filter on — things like `status`, `city`,
`product_name`, or `created_at`. Each dimension maps to a column or SQL
expression in your data source.
See the [dimensions reference][ref-dimensions-ref] for the full list of
parameters and configuration options.
## Defining dimensions
A dimension specifies the SQL expression and its type:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: status
sql: status
type: string
- name: created_at
sql: created_at
type: time
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true },
status: { sql: `status`, type: `string` },
created_at: { sql: `created_at`, type: `time` }
}
})
```
### Dimension types
| Data type in SQL | Dimension type in Cube |
| ------------------------------ | ---------------------- |
| `timestamp`, `date`, `time` | [`time`][ref-type] |
| `text`, `varchar` | [`string`][ref-type] |
| `integer`, `bigint`, `decimal` | [`number`][ref-type] |
| `boolean` | [`boolean`][ref-type] |
### Primary keys
Every cube that participates in [joins][ref-joins] should define a
[`primary_key`][ref-primary-key] dimension. Cube uses primary keys to avoid
fanouts — when rows get duplicated during joins and aggregates are
over-counted. Composite primary keys can be created by concatenating columns:
```yaml theme={"dark"}
dimensions:
- name: composite_key
sql: "CONCAT({CUBE}.order_id, '-', {CUBE}.product_id)"
type: string
primary_key: true
```
## Time dimensions
Time dimensions are dimensions of the [`time` type][ref-type]. They enable
grouping by time granularity (year, quarter, month, week, day, hour, minute,
second) and are essential for time-series analysis.
```yaml theme={"dark"}
dimensions:
- name: created_at
sql: created_at
type: time
```
When queried, you can group by any built-in granularity without defining
additional dimensions.
### Custom granularities
You can define [custom granularities][ref-granularities] for time dimensions
when the built-in ones don't fit — for example, weeks starting on Sunday
or fiscal years:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
dimensions:
- name: created_at
sql: created_at
type: time
granularities:
- name: sunday_week
interval: 1 week
offset: -1 day
- name: fiscal_year
interval: 1 year
offset: 1 month
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
dimensions: {
created_at: {
sql: `created_at`,
type: `time`,
granularities: {
sunday_week: { interval: `1 week`, offset: `-1 day` },
fiscal_year: { interval: `1 year`, offset: `1 month` }
}
}
}
})
```
Time dimensions are essential for performance features like
[partitioned pre-aggregations][ref-partition-preaggs] and
[incremental refreshes][ref-incremental-preaggs].
See the following recipes:
* For a [custom granularity][ref-custom-granularity-recipe] example.
* For a [custom calendar][ref-custom-calendar-recipe] example.
## Proxy dimensions
Proxy dimensions reference dimensions from the same cube or other cubes,
providing a way to reuse existing definitions and reduce code duplication.
### Within the same cube
Reference existing dimensions to build derived ones without duplicating SQL:
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
sql_table: users
dimensions:
- name: initials
sql: "SUBSTR(first_name, 1, 1)"
type: string
- name: last_name
sql: "UPPER(last_name)"
type: string
- name: full_name
sql: "{initials} || '. ' || {last_name}"
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
sql_table: `users`,
dimensions: {
initials: { sql: `SUBSTR(first_name, 1, 1)`, type: `string` },
last_name: { sql: `UPPER(last_name)`, type: `string` },
full_name: { sql: `${initials} || '. ' || ${last_name}`, type: `string` }
}
})
```
### From other cubes
If cubes are [joined][ref-joins], you can bring a dimension from one cube
into another. Cube generates the necessary joins automatically:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
joins:
- name: users
sql: "{CUBE}.user_id = {users.id}"
relationship: many_to_one
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: user_name
sql: "{users.name}"
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
joins: {
users: {
sql: `${CUBE}.user_id = ${users.id}`,
relationship: `many_to_one`
}
},
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true },
user_name: { sql: `${users.name}`, type: `string` }
}
})
```
### Time dimension granularity references
When referencing a time dimension, you can specify a granularity to create
a proxy dimension at that specific granularity — including
[custom granularities](#custom-granularities):
```yaml theme={"dark"}
dimensions:
- name: created_at
sql: created_at
type: time
granularities:
- name: sunday_week
interval: 1 week
offset: -1 day
- name: created_at_year
sql: "{created_at.year}"
type: time
- name: created_at_sunday_week
sql: "{created_at.sunday_week}"
type: time
```
## Subquery dimensions
Subquery dimensions reference [measures][ref-measures-page] from other cubes,
effectively turning an aggregate into a per-row value. This enables nested
aggregations — for example, calculating the average of per-customer order counts.
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
joins:
- name: users
sql: "{users}.id = {CUBE}.user_id"
relationship: many_to_one
dimensions:
- name: id
sql: id
type: number
primary_key: true
measures:
- name: count
type: count
- name: users
sql_table: users
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: name
sql: name
type: string
- name: order_count
sql: "{orders.count}"
type: number
sub_query: true
measures:
- name: avg_order_count
sql: "{order_count}"
type: avg
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
joins: {
users: {
sql: `${users}.id = ${CUBE}.user_id`,
relationship: `many_to_one`
}
},
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true }
},
measures: {
count: { type: `count` }
}
})
cube(`users`, {
sql_table: `users`,
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true },
name: { sql: `name`, type: `string` },
order_count: {
sql: `${orders.count}`,
type: `number`,
sub_query: true
}
},
measures: {
avg_order_count: {
sql: `${order_count}`,
type: `avg`
}
}
})
```
The `order_count` subquery dimension computes the order count per user.
The `avg_order_count` measure then averages those per-user values. Cube
implements this as a correlated subquery via joins for optimal performance.
See the following recipes:
* How to calculate [nested aggregates][ref-nested-aggregates-recipe].
* How to calculate [filtered aggregates][ref-filtered-aggregates-recipe].
## Links
Dimensions can declare **links** — clickable navigation targets that supporting
tools (such as [Cube Cloud Workbooks][ref-workbooks]) surface next to the
dimension's values.
`links` is a **list**, so a single dimension can declare **any number of
links** — they all appear together in the cell menu. Each link points either to
an **external URL** (`url`) or to **another Cube Cloud dashboard** (`dashboard`,
a drill-in), and its URL is built per row from the dimension's data. The example
below declares two links on one dimension (an external search and a drill-in).
Dimension `links` require Cube **v1.6.53** or newer.
### Parameters
`links` is a list of link objects. Each link accepts:
| Parameter | Required? | Description |
| ----------- | ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `name` | **Required** | Identifier, unique within the dimension. Also used in the synthetic dimension name (see [Behavior](#behavior)). |
| `label` | **Required** | The text shown for the link in the UI. |
| `url` | **Either `url` or `dashboard`** | SQL expression that builds an **external** URL per row. May [reference][ref-references] columns and other dimensions. |
| `dashboard` | **Either `url` or `dashboard`** | The target Cube Cloud dashboard's **slug** (a **drill-in**). Each link sets exactly one of `url` or `dashboard` — never both — but different links on the same dimension can mix the two. |
| `icon` | Optional | A [Tabler icon][link-tabler] name (see [Icons](#icons)). Defaults to a generic link icon. |
| `primary` | Optional | Marks the link that renders **inline** on the cell value (see [Behavior](#behavior)). At most one link per dimension may set it. |
| `target` | Optional | Where to open the link: `blank` (new tab) or `self` (same tab). The default depends on the link kind — `url:` links default to `blank`, `dashboard:` drill-ins default to `self` (in-app). An explicit `target` is honored for both (on `dashboard:` drill-ins this requires Cube **v1.7.5** or newer; see [Behavior](#behavior)). |
| `params` | Optional | Extra per-row parameters. For `url:` they are appended as query parameters. For `dashboard:` they become **equality filters** on the target dashboard — each `key` is a member of the target dashboard's view (a cube path such as `orders.status` is auto-resolved to the matching view member), and `value` is the per-row value. |
### Example
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
dimensions:
- name: status
sql: status
type: string
links:
# External link — opens a URL built from the row's value.
# `primary` makes it render inline on the cell value.
- name: search
label: Search the web
url: "CONCAT('https://www.google.com/search?q=order+', {CUBE}.status)"
icon: brand-google
target: blank
primary: true
# Drill-in link — opens another Cube Cloud dashboard, filtered by the row
- name: details
label: Open order details
dashboard: orders-detail # the target dashboard's slug
params:
- key: orders_view.status
value: "{CUBE}.status"
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
dimensions: {
status: {
sql: `status`,
type: `string`,
links: [
// External link — opens a URL built from the row's value.
// `primary` makes it render inline on the cell value.
{
name: `search`,
label: `Search the web`,
url: `CONCAT('https://www.google.com/search?q=order+', ${CUBE}.status)`,
icon: `brand-google`,
target: `blank`,
primary: true
},
{
name: `details`,
label: `Open order details`,
dashboard: `orders-detail`,
params: [{ key: `orders_view.status`, value: `${CUBE}.status` }]
}
]
}
}
})
```
### Icons
`icon` accepts any name from the [Tabler icon set][link-tabler] — the
kebab-case name **without** any prefix, for example `brand-google`,
`external-link`, `layout-dashboard`, or `send`. Browse and search the available
names at [tabler.io/icons][link-tabler]. If `icon` is omitted, a default link
icon is shown.
### Setting a dashboard slug
A `dashboard:` link targets another dashboard by its **slug** — a short, stable,
human-readable identifier (e.g. `orders-detail`). The slug is **portable across
environments**: it is resolved within the current deployment, so the same model
works in development and production without hardcoding dashboard IDs.
To set a slug in Cube Cloud, open the target dashboard, open its **options
sidebar**, and fill the **Slug** field (see
[Dashboards → Dashboard slug](/docs/explore-analyze/dashboards#dashboard-slug)).
Slugs are **unique per deployment**, and the `dashboard:` value in your link must
match the slug exactly.
**There is no required order.** You can write the `dashboard:` slug in the model
first (it won't break anything — a link whose slug doesn't yet resolve is simply
skipped in the cell menu) and set the dashboard's slug later, or set the
dashboard slug first and reference it from the model afterwards. The link starts
working as soon as both sides use the same slug.
### Behavior
* **Per-row resolution** — each link's `url` (and its `params`) is evaluated for
every row, so the destination reflects the clicked cell. Internally a link
compiles to a hidden **synthetic dimension** named
`___link__url` holding the resolved URL; it is added to the
query automatically and is never shown as a column.
* **Null values** — if the source value (and thus the resolved URL) is null for
a row, that link is omitted from the menu for that row.
* **Where links appear** — in Cube Cloud, links surface in the table
[**cell menu**](/docs/explore-analyze/charts/chart-types/table#cell-menu) on
every results table (dashboard and workbook table charts, embedded dashboards,
and Explore / SQL results). A left-click on a cell opens the menu listing the
dimension's links, alongside **Copy value** and, on measures, **Drill down**.
* **Inline rendering** — the link marked `primary` also renders **inline**, as a
clickable link on the cell value itself, with no chart configuration. The
remaining links stay in the cell menu, which still lists every link. In a
workbook table chart, the **Style** tab can override which declared link
renders inline for a column; see
[Inline links](/docs/explore-analyze/charts/chart-types/table#inline-links).
* **Drill-in vs external** — a `dashboard:` link navigates **in-app, in the same
tab** by default; an explicit `target` is honored (`blank` opens a new tab,
`self` stays in-app), and a Cmd/Ctrl-click always opens a new tab. A `url:`
link opens per its `target` (`blank` by default). Honoring an explicit `target`
on a drill-in requires Cube **v1.7.5** or newer — older runtimes always navigate
drill-ins in-app regardless of `target`.
* **Filters** — for `dashboard:` links, each `params` entry becomes an
**equality filter** on the target (the `key` is a view member, the `value` is
the per-row value). An unknown or inaccessible target slug is skipped
gracefully — the user is notified and no navigation happens.
See the [`links` reference][ref-dimensions-ref] for the canonical parameter
list.
## Hierarchies
Dimensions can be organized into [hierarchies][ref-hierarchies] to define
drill-down paths (e.g., Country → State → City):
```yaml theme={"dark"}
cubes:
- name: users
# ...
dimensions:
- name: country
sql: country
type: string
- name: state
sql: state
type: string
- name: city
sql: city
type: string
hierarchies:
- name: location
levels:
- country
- state
- city
```
## Next steps
* See the [dimensions reference][ref-dimensions-ref] for all parameters
* Learn about [measures][ref-measures-page] for aggregated calculations
* Explore [custom granularities][ref-granularities] for fiscal calendars
and non-standard time periods
[ref-dimensions-ref]: /reference/data-modeling/dimensions
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-references]: /docs/data-modeling/concepts/syntax#references
[link-tabler]: https://tabler.io/icons
[ref-measures-page]: /docs/data-modeling/measures
[ref-joins]: /docs/data-modeling/joins
[ref-type]: /reference/data-modeling/dimensions#type
[ref-primary-key]: /reference/data-modeling/dimensions#primary_key
[ref-granularities]: /reference/data-modeling/dimensions#granularities
[ref-hierarchies]: /reference/data-modeling/hierarchies
[ref-partition-preaggs]: /docs/pre-aggregations/matching-pre-aggregations#partitioning
[ref-incremental-preaggs]: /reference/data-modeling/pre-aggregations#incremental
[ref-custom-granularity-recipe]: /recipes/data-modeling/custom-granularity
[ref-custom-calendar-recipe]: /recipes/data-modeling/custom-calendar
[ref-nested-aggregates-recipe]: /recipes/data-modeling/nested-aggregates
[ref-filtered-aggregates-recipe]: /recipes/data-modeling/filtered-aggregates
# Export and import
Source: https://docs.cube.dev/docs/data-modeling/dynamic/code-reusability-export-and-import
This functionality only works with data models written in JavaScript, not YAML.
This functionality only works with data models written in JavaScript, not YAML.
In Cube, your data model is code, and code is much easier to manage when it is
in small, digestible chunks. It is best practice to keep files small and
containing only relevant and non-duplicated code. As your data model grows,
maintaining and debugging is much easier with a well-organized codebase.
Cube data models in JavaScript supports ES6-style [`export`][mdn-js-es6-export]
and [`import`][mdn-js-es6-import] statements, which allow writing code in one
file and sharing it, so it can be used by another file or files.
There are several typical use cases in Cube where it is considered best practice
to extract some variables or functions and then import it when needed.
## Managing constants
Quite often, you may want to have an array of test user IDs to exclude from your
analysis. You can define it once and `export` it like this:
```javascript theme={"dark"}
// in constants.js
export const TEST_USER_IDS = [1, 2, 3, 4, 5]
```
Later, you can `import` into the cube whenever needed:
```javascript theme={"dark"}
// in Users.js
import { TEST_USER_IDS } from "./constants"
cube(`users`, {
// ...
measures: {
// ...
},
dimensions: {
// ...
},
segments: {
exclude_test_users: {
sql: `${CUBE}.id NOT IN (${TEST_USER_IDS.join(", ")})`
}
}
})
```
## Helper functions
You can assign some commonly used SQL snippets to JavaScript functions. The
example below shows a parsing helper function, which can be used across any
number of cubes to correctly parse a date if it was stored as a string.
You can read more about working with [string time dimensions
here][ref-schema-string-time-dims].
```javascript theme={"dark"}
// in helpers.js
export const parseDateWithTimeZone = (column) =>
`PARSE_TIMESTAMP('%F %T %Ez', ${column})`
```
```javascript theme={"dark"}
// in events.js
import { parseDateWithTimeZone } from "./helpers"
cube(`events`, {
sql_table: `events`,
// ...
dimensions: {
date: {
sql: `${parseDateWithTimeZone("date")}`,
type: `time`
}
}
})
```
## Import from parent directories
You may need to import from parent directories as Cube flattens nested
directories. The example below shows a correct way to import a helper function,
which is located in a parent directory.
```tree theme={"dark"}
.
├── README.md
├── cube.js
├── package.json
└── model/
├── shared_utils/
│ └── utils.js
└── sales/
└── orders.js
```
```javascript theme={"dark"}
// in model/sales/orders.js
import { capitalize } from "./shared_utils/utils"
```
```javascript theme={"dark"}
// in model/shared_utils/utils.js
export const capitalize = (s) => s.charAt(0).toUpperCase() + s.slice(1)
```
[mdn-js-es6-export]: https://developer.mozilla.org/en-US/docs/web/javascript/reference/statements/export
[mdn-js-es6-import]: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Statements/import
[ref-schema-string-time-dims]: /recipes/data-modeling/string-time-dimensions
# Dynamic Data Models
Source: https://docs.cube.dev/docs/data-modeling/dynamic/index
Generate data models programmatically using Jinja with Python or JavaScript for dynamic schema creation.
Cube supports authoring data models dynamically — useful for de-duplicating
common patterns across cubes, generating models from a remote source, or
adapting the schema per tenant at runtime.
Pick the approach that matches the language your data models are written in:
Template YAML data models with Jinja, and use Python for loops, includes,
and runtime generation.
Generate cubes and views on-the-fly from JavaScript data models using
`asyncModule()`.
## How it fits together
The diagrams below show how YAML and JavaScript data models are parsed and
compiled before they're served by Cube.
# Data modeling with JavaScript
Source: https://docs.cube.dev/docs/data-modeling/dynamic/javascript
This functionality only works with data models written in JavaScript, not YAML.
This functionality only works with data models written in JavaScript, not YAML.
For similar functionality in YAML, see [Dynamic data models with Jinja and Python](/docs/data-modeling/dynamic/jinja).
Cube allows data models to be created on-the-fly using a special
[`asyncModule()`][ref-async-module] function only available in the
[execution environment][ref-schema-env]. `asyncModule()` allows registering an
async function to be executed at the end of the data model compile phase so
additional definitions can be added. This is often useful in situations where
data model properties can be dynamically updated through an API, for example.
Each `asyncModule` call will be invoked only once per data model compilation.
[ref-schema-env]: /docs/data-modeling/dynamic/schema-execution-environment
[ref-async-module]: /docs/data-modeling/dynamic/schema-execution-environment#asyncmodule
When creating data models via `asyncModule()`, it is important to be aware of
the following differences compared to statically defined ones with `cube()`:
* The `sql` and `drill_members` properties for both dimensions and measures must
be of type `() => string` and `() => string[]` accordingly
Cube supports importing JavaScript logic from other files in a data model, so it
is useful to declare utility functions for handling the above differences in a
separate file:
```javascript theme={"dark"}
// model/utils.js
export const convertStringPropToFunction = (propNames, dimensionDefinition) => {
let newResult = { ...dimensionDefinition };
propNames.forEach((propName) => {
const propValue = newResult[propName];
if (!propValue) {
return;
}
newResult[propName] = () => propValue;
});
return newResult;
};
export const transformDimensions = (dimensions) => {
return Object.keys(dimensions).reduce((result, dimensionName) => {
const dimensionDefinition = dimensions[dimensionName];
return {
...result,
[dimensionName]: convertStringPropToFunction(
["sql"],
dimensionDefinition
),
};
}, {});
};
export const transformMeasures = (measures) => {
return Object.keys(measures).reduce((result, dimensionName) => {
const dimensionDefinition = measures[dimensionName];
return {
...result,
[dimensionName]: convertStringPropToFunction(
["sql", "drill_members"],
dimensionDefinition
),
};
}, {});
};
```
## Generation
In the following example, we retrieve a JSON object representing all our cubes
using `fetch()`, transform some of the properties to be functions that return a
string, and then finally use the [`cube()` global function][ref-globals] to
generate data models from that data:
[ref-globals]: /docs/data-modeling/dynamic/schema-execution-environment#cube-js-globals-cube-and-others
```javascript theme={"dark"}
// model/cubes/DynamicDataModel.js
const fetch = require("node-fetch");
import {
convertStringPropToFunction,
transformDimensions,
transformMeasures,
} from "./utils";
asyncModule(async () => {
const dynamicCubes = await (
await fetch("http://your-api-endpoint/dynamicCubes")
).json();
console.log(dynamicCubes);
// [
// {
// name: 'dynamic_cube_model',
// sql_table: 'my_table',
//
// measures: {
// price: {
// sql: `price`,
// type: `number`,
// }
// },
//
// dimensions: {
// color: {
// sql: `color`,
// type: `string`,
// },
// },
// },
// ]
dynamicCubes.forEach((dynamicCube) => {
const dimensions = transformDimensions(dynamicCube.dimensions);
const measures = transformMeasures(dynamicCube.measures);
cube(dynamicCube.name, {
sql: dynamicCube.sql,
dimensions,
measures,
pre_aggregations: {
main: {
// ...
},
},
});
});
});
```
## Usage with `schema_version`
It is also useful to be able to recompile the data model when there are changes
in the underlying input data. For this purpose, the [`schema_version`
][link-config-schema-version] value in the `cube.js` configuration options can
be specified as an asynchronous function:
```javascript theme={"dark"}
// cube.js
module.exports = {
schemaVersion: async ({ securityContext }) => {
const schemaVersions = await (
await fetch("http://your-api-endpoint/schema_version")
).json();
return schemaVersions[securityContext.tenantId];
},
};
```
[link-config-schema-version]: /reference/configuration/config#schema_version
## Usage with COMPILE\_CONTEXT
The `COMPILE_CONTEXT` global object can also be used in conjunction with async
data model creation to allow for multi-tenant deployments of Cube.
In an example scenario where all tenants share the same cube, but see different
dimensions and measures, you could do the following:
```javascript theme={"dark"}
// model/cubes/DynamicDataModel.js
const fetch = require("node-fetch");
import {
convertStringPropToFunction,
transformDimensions,
transformMeasures,
} from "./utils";
asyncModule(async () => {
const {
securityContext: { tenantId },
} = COMPILE_CONTEXT;
const dynamicCubes = await (
await fetch(`http://your-api-endpoint/dynamicCubes`)
).json();
const allowedDimensions = await (
await fetch(`http://your-api-endpoint/dynamicDimensions/${tenantId}`)
).json();
const allowedMeasures = await (
await fetch(`http://your-api-endpoint/dynamicMeasures/${tenantId}`)
).json();
dynamicCubes.forEach((dynamicCube) => {
const dimensions = transformDimensions(allowedDimensions);
const measures = transformMeasures(allowedMeasures);
cube(dynamicCube.name, {
sql: dynamicCube.sql,
title: `${dynamicCube.title}-${tenantId}`,
dimensions,
measures,
pre_aggregations: {
main: {
// ...
},
},
});
});
});
```
## Usage with data\_source
When using multiple databases, you'll need to ensure you set the
[`data_source`][ref-schema-datasource] property for any asynchronously-created
data models, as well as ensuring the corresponding database drivers are set up with
[`driverFactory()`][ref-config-driverfactory] in your [`cube.js` configuration
file][ref-config].
[ref-schema-datasource]: /reference/data-modeling/cube#data_source
[ref-config-driverfactory]: /reference/configuration/config#driver_factory
[ref-config]: /reference/configuration/config
For an example scenario where data models may use either MySQL or Postgres
databases, you could do the following:
```javascript theme={"dark"}
// model/cubes/DynamicDataModel.js
const fetch = require("node-fetch");
import {
convertStringPropToFunction,
transformDimensions,
transformMeasures,
} from "./utils";
asyncModule(async () => {
const dynamicCubes = await (
await fetch("http://your-api-endpoint/dynamicCubes")
).json();
dynamicCubes.forEach((dynamicCube) => {
const dimensions = transformDimensions(dynamicCube.dimensions);
const measures = transformMeasures(dynamicCube.measures);
cube(dynamicCube.name, {
data_source: dynamicCube.data_source,
sql: dynamicCube.sql,
dimensions,
measures,
pre_aggregations: {
main: {
// ...
},
},
});
});
});
```
```javascript theme={"dark"}
// cube.js
const MySQLDriver = require("@cubejs-backend/mysql-driver");
const PostgresDriver = require("@cubejs-backend/postgres-driver");
module.exports = {
driverFactory: ({ dataSource }) => {
if (dataSource === "mysql") {
return new MySQLDriver({ database: dataSource });
}
return new PostgresDriver({ database: dataSource });
},
};
```
# Data modeling with YAML, Jinja, and Python
Source: https://docs.cube.dev/docs/data-modeling/dynamic/jinja
Jinja and Python techniques for templating YAML models—loops, includes, and runtime generation—to keep large or multi-tenant schemas maintainable.
Cube supports authoring dynamic data models using the [Jinja templating
language][jinja] and Python. This allows de-duplicating common patterns in your data models
as well as dynamically generating data models from a remote data source.
Jinja is supported in all YAML data model files.
## YAML
It is recommended to default to YAML syntax because of its simplicity and readability.
### Folded and literal strings
Sometimes you might want to use multi-line strings in YAML-based data models, e.g.,
in parameters such as `sql` or `description`. It is recommended to use [literal][ref-yaml-literal]
(`|`) string style in such cases as it preserves line breaks.
```yaml theme={"dark"}
cubes:
- name: orders
description: |
This cube represents customer orders.
It includes measures for total sales and order count.
sql: |
-- Fetch only relevant columns
SELECT id, created_at, total_amount
FROM staging.orders
```
## Jinja
Please check the [Jinja documentation][jinja-docs] for details on Jinja syntax.
### Previewing YAML
You can preview the data model code after applying Jinja templates in the **[Data
Model][ref-data-model-editor]** editor by clicking **... → Jinja Preview**
on files that contain Jinja templates in the sidebar.
Currently, there's no way to preview the data model code in YAML after applying
Jinja templates in Cube Core. Please [track this issue](https://github.com/cube-js/cube/issues/8134).
You can also view the resulting data model in [Playground][ref-payground] and [Visual
Model][ref-visual-model]. Also, you can introspect the data model using the
[`/v1/meta` REST (JSON) API endpoint][ref-meta-api].
### Loops
Jinja supports [looping][jinja-docs-for-loop] over lists and dictionaries. In
the following example, we loop over a list of nested properties and generate a
`LEFT JOIN UNNEST` clause for each one: for each one:
```yaml theme={"dark"}
{%- set nested_properties = [
"referrer",
"href",
"host",
"pathname",
"search"
] -%}
cubes:
- name: analytics
sql: |
SELECT
{%- for prop in nested_properties %}
{{ prop }}_prop.value AS {{ prop }}
{%- endfor %}
FROM public.events
{%- for prop in nested_properties %}
LEFT JOIN UNNEST(properties) AS {{ prop }}_prop ON {{ prop }}_prop.key = '{{ prop }}'
{%- endfor %}
```
Another useful pattern is to loop over a dictionary of values and generate a
measure for each one, as in the following example:
```yaml theme={"dark"}
{%- set metrics = {
"mau": 30,
"wau": 7,
"day": 1
} %}
cubes:
- name: orders
sql_table: public.orders
measures:
{%- for name, days in metrics | items %}
- name: {{ name | safe }}
type: count_distinct
sql: user_id
rolling_window:
trailing: {{ days }} day
offset: start
{% endfor %}
```
### Macros
Cube data models also support Jinja macros, which allow you to define reusable
snippets of code. You can read more about macros in the [Jinja
documentation][jinja-docs-macros].
In the following example, we define a macro called `dimension()` which generates
a dimension definition in Cube. This macro is then invoked multiple times to
generate multiple dimensions:
```yaml theme={"dark"}
{# Declare the macro before using it, otherwise Jinja will throw an error. #}
{%- macro dimension(column_name, type='string', primary_key=False) -%}
- name: {{ column_name }}
sql: {{ column_name }}
type: {{ type }}
{% if primary_key -%}
primary_key: true
{% endif -%}
{% endmacro -%}
cubes:
- name: orders
sql_table: public.orders
dimensions:
{{ dimension('id', 'number', primary_key=True) }}
{{ dimension('status') }}
{{ dimension('created_at', 'time') }}
{{ dimension('completed_at', 'time') }}
```
You could also use macros to generate SQL snippets for use in the `sql`
property:
```yaml theme={"dark"}
{%- macro cents_to_dollars(column_name, precision=2) -%}
({{ column_name }} / 100)::NUMERIC(16, {{ precision }})
{%- endmacro -%}
cubes:
- name: payments
sql: |
SELECT
id AS payment_id,
{{ cents_to_dollars('amount') }} AS amount_usd
FROM app_data.payments
```
### Reusing macros across files
You can define macros in dedicated `.jinja` files and import them into your
data model files using Jinja's [`import`][jinja-docs-import] statement. This
is useful for sharing common patterns across multiple cubes and views.
Consider the following project structure:
```tree theme={"dark"}
.
└── cube/
├── model/
│ ├── cubes/
│ │ └── orders.yml
│ ├── views/
│ └── macros/
│ └── common_dimensions.jinja
└── cube.py
```
First, define reusable macros in a `.jinja` file under the `macros/` directory:
```yamltitle="model/macros/common_dimensions.jinja" theme={"dark"}
{%- macro dimension(column_name, type='string', primary_key=False) -%}
- name: {{ column_name }}
sql: {{ column_name }}
type: {{ type }}
{% if primary_key -%}
primary_key: true
{% endif -%}
{% endmacro -%}
{%- macro cents_to_dollars(column_name, precision=2) -%}
({{ column_name }} / 100)::NUMERIC(16, {{ precision }})
{%- endmacro -%}
```
Then, import and use those macros in your data model files:
```yamltitle="model/cubes/orders.yml" theme={"dark"}
{%- import "macros/common_dimensions.jinja" as common -%}
cubes:
- name: orders
sql_table: public.orders
dimensions:
{{ common.dimension('id', 'number', primary_key=True) }}
{{ common.dimension('status') }}
{{ common.dimension('created_at', 'time') }}
measures:
- name: amount_usd
type: sum
sql: "{{ common.cents_to_dollars('amount') }}"
```
The import path is relative to the `model/` directory.
### Escaping unsafe strings
[Auto-escaping][jinja-docs-autoescaping] of unsafe string values in Jinja
templates is enabled by default. It means that any strings coming from Python
might get wrapped in quotes, potentially breaking YAML syntax.
You can work around that by using the [`safe` Jinja
filter][jinja-docs-filters-safe] with such string values:
```yaml theme={"dark"}
cubes:
- name: my_cube
description: {{ get_unsafe_string() | safe }}
```
Alternatively, you can wrap unsafe strings into instances of the following
class in your Python code, effectively marking them as safe. This is
particularly useful for library code, e.g., similar to the
[`cube_dbt`][ref-cube-dbt] package.
```python theme={"dark"}
class SafeString(str):
is_safe: bool
def __init__(self, v: str):
self.is_safe = True
```
## Python
### Template context
You can use Python to declare functions that can be invoked and variables that can be
referenced from within a Jinja template. These functions and variables must be defined
in `model/globals.py` file and registered in the `TemplateContext` instance.
See the [`TemplateContext` reference][ref-cube-template-context] for more details.
In the following example, we declare a function called `load_data` that supposedly loads
data from a remote API endpoint. We will then use the function to generate a data model:
```python theme={"dark"}
from cube import TemplateContext
template = TemplateContext()
@template.function('load_data')
def load_data():
client = MyApiClient("example.com")
return client.load_data()
class MyApiClient:
def __init__(self, api_url):
self.api_url = api_url
# mock API call
def load_data(self):
api_response = {
"cubes": [
{
"name": "cube_from_api",
"measures": [
{ "name": "count", "type": "count" },
{ "name": "total", "type": "sum", "sql": "amount" }
],
"dimensions": []
},
{
"name": "cube_from_api_with_dimensions",
"measures": [
{ "name": "active_users", "type": "count_distinct", "sql": "user_id" }
],
"dimensions": [
{ "name": "city", "sql": "city_column", "type": "string" }
]
}
]
}
return api_response
```
Now that we've decorated our function with the `@template.function` decorator, we can
call it from within a Jinja template. In the following example, we'll call the
`load_data()` function and use the result to generate a data model.
```yaml theme={"dark"}
cubes:
{# Here we use the decorated function from earlier #}
{%- for cube in load_data()["cubes"] %}
- name: {{ cube.name }}
{%- if cube.measures is not none and cube.measures|length > 0 %}
measures:
{%- for measure in cube.measures %}
- name: {{ measure.name }}
type: {{ measure.type }}
{%- if measure.sql %}
sql: {{ measure.sql }}
{%- endif %}
{%- endfor %}
{%- endif %}
{%- if cube.dimensions is not none and cube.dimensions|length > 0 %}
dimensions:
{%- for dimension in cube.dimensions %}
- name: {{ dimension.name }}
type: {{ dimension.type }}
sql: {{ dimension.sql }}
{%- endfor %}
{%- endif %}
{%- endfor %}
```
### Imports
In the `model/globals.py` file (or the `cube.py` configuration file), you can
import modules from the current directory. In the following example, we import a function
from the `utils` module and use it to populate a variable in the template context:
```pythontitle="model/utils.py" theme={"dark"}
def answer_to_main_question() -> str:
return "42"
```
```pythontitle="model/globals.py" theme={"dark"}
from cube import TemplateContext
from utils import answer_to_main_question
template = TemplateContext()
answer = answer_to_main_question()
template.add_variable('answer', answer)
```
### Dependencies
If you need to use dependencies in your dynamic data model (or your `cube.py`
configuration file), you can list them in the `requirements.txt` file in the root
directory of your Cube deployment. They will be automatically installed with `pip` on
the startup.
[`cube` package][ref-cube-package] is available out of the box, it doesn't need to be
listed in `requirements.txt`.
If you use dbt for data transformation, you might find the [`cube_dbt`
package][ref-cube-dbt-package] useful. It provides a set of utilities that simplify
defining the data model in YAML [based on dbt models][ref-cube-with-dbt].
If you need to use dependencies with native extensions, build a [custom Docker
image][ref-docker-image-extension].
[jinja]: https://jinja.palletsprojects.com/
[jinja-docs]: https://jinja.palletsprojects.com/en/3.1.x/templates/
[jinja-docs-for-loop]: https://jinja.palletsprojects.com/en/3.1.x/templates/#for
[jinja-docs-macros]: https://jinja.palletsprojects.com/en/3.1.x/templates/#macros
[jinja-docs-import]: https://jinja.palletsprojects.com/en/3.1.x/templates/#import
[jinja-docs-autoescaping]: https://jinja.palletsprojects.com/en/3.1.x/api/#autoescaping
[jinja-docs-filters-safe]: https://jinja.palletsprojects.com/en/3.1.x/templates/#jinja-filters.safe
[ref-cube-dbt]: /reference/data-modeling/cube_dbt
[ref-visual-model]: /docs/data-modeling/visual-modeler
[ref-docker-image-extension]: /admin/deployment/core#extend-the-docker-image
[ref-cube-package]: /reference/data-modeling/cube-package
[ref-cube-template-context]: /reference/data-modeling/cube-package#templatecontext-class
[ref-cube-dbt-package]: /reference/data-modeling/cube_dbt
[ref-cube-with-dbt]: /recipes/data-modeling/dbt
[ref-data-model-editor]: /docs/data-modeling/data-model-ide
[ref-payground]: /docs/explore-analyze/playground
[ref-meta-api]: /reference/core-data-apis/rest-api/reference#base_path/v1/meta
[ref-yaml-literal]: https://yaml.org/spec/1.2.2/#812-literal-style
[ref-yaml-folded]: https://yaml.org/spec/1.2.2/#813-folded-style
# Execution Environment (JavaScript)
Source: https://docs.cube.dev/docs/data-modeling/dynamic/schema-execution-environment
How JavaScript model code runs inside Cube’s Node VM, including custom require, forbidden globals, and constraints compared to normal Node scripts.
Cube Data Model Compiler uses [Node.js VM][nodejs-vm] to execute data model
compiler code. It gives required flexibility allowing transpiling data model
files before they get executed, storing data models in external databases and
executing untrusted code in a safe manner. Cube data model JavaScript is
standard JavaScript supported by Node.js starting in version 8 with the
following exceptions.
## Require
Being executed in VM, data model JavaScript code doesn't have access to [Node.js
require][nodejs-require] directly. Instead `require()` is implemented by Data
Model Compiler to provide access to other data model files and to regular
Node.js modules. Besides that, the data model `require()` can resolve Cube
packages such as `Funnels` unlike standard Node.js `require()`.
## Node.js globals (process.env, console.log and others)
Data model JavaScript code doesn't have access to any standard Node.js globals
like `process` or `console`. In order to access `process.env`, utility functions
can be added outside the `model/` directory:
**tablePrefix.js:**
```javascript theme={"dark"}
exports.tableSchema = () => process.env.TABLE_SCHEMA
```
**model/cubes/Users.js**:
```javascript theme={"dark"}
import { tableSchema } from "../tablePrefix"
cube(`users`, {
sql_table: `${tableSchema()}.users`,
// ...
})
```
## console.log
Data models cannot access `console.log` due to a separate [VM
instance][nodejs-vm] that runs it. Suppose you find yourself writing complex
logic for SQL generation that depends on a lot of external input. In that case,
you probably want to introduce a helper service outside of the [data model
directory][ref-schema-path] that you can debug as usual Node.js code.
## Cube globals (cube and others)
Cube defines `cube()`, `context()` and `asyncModule()` global variable functions
in order to provide API for data model configuration which aren't normally
accessible outside of a Cube data model.
## Import / Export
Data model JavaScript files are transpiled to convert ES6 `import` and `export`
expressions to corresponding Node.js calls. In fact `import` is routed to
[Require][self-require] method.
`export` can be used to define named exports as well as default ones:
**constants.js:**
```javascript theme={"dark"}
export const TEST_USER_IDS = [1, 2, 3, 4, 5]
```
**usersSql.js:**
```javascript theme={"dark"}
export default (usersTable) => `select * from ${usersTable}`
```
Later, you can `import` into the cube, wherever needed:
**Users.js**:
```javascript theme={"dark"}
// in users.js
import { TEST_USER_IDS } from "./constants"
import usersSql from "./usersSql"
cube(`users`, {
sql: usersSql(`users`),
measures: {
/* ... */
},
dimensions: {
/* ... */
},
segments: {
excludeTestUsers: {
sql: `${CUBE}.id NOT IN (${TEST_USER_IDS.join(", ")})`
}
}
})
```
## asyncModule
Data models can be externally stored and retrieved through an asynchronous
operation using the `asyncModule()`. For more information, consult the [dynamic
data model creation][ref-dynamic-schemas].
## Context symbols transpile
Cube uses a custom transpiler to optimize boilerplate code around referencing
cubes and cube members. There are reserved property names inside `cube`
definition that undergo reference resolve transpiling process:
* `sql`
* `measures`
* `dimensions`
* `segments`
* `time_dimension`
* `drill_members`
* `context_members`
Each of these properties inside `cube` and `context` definitions are transpiled
to functions with resolved arguments. For example:
```javascript theme={"dark"}
cube(`users`, {
// ...
measures: {
count: {
type: `count`
},
ratio: {
sql: `SUM(${CUBE}.amount) / ${count}`,
type: `number`
}
}
})
```
is transpiled to:
```javascript theme={"dark"}
cube(`users`, {
// ...
measures: {
count: {
type: `count`
},
ratio: {
sql: (CUBE, count) => `SUM(${CUBE}.amount) / ${count}`,
type: `number`
}
}
})
```
So for example if you want to pass the definition of `ratio` outside of the
cube, you would define it as:
```javascript theme={"dark"}
const measureRatioDefinition = {
sql: (CUBE, count) => `sum(${CUBE}.amount) / ${count}`,
type: `number`
}
cube(`users`, {
// ...
measures: {
count: {
type: `count`
},
ratio: measureRatioDefinition
}
})
```
[nodejs-vm]: https://nodejs.org/api/vm.html
[nodejs-require]: https://nodejs.org/api/modules.html#modules_require_id
[ref-dynamic-schemas]: /docs/data-modeling/dynamic
[self-require]: #require
[ref-schema-path]: /reference/configuration/config#schema_path
# Extending cubes
Source: https://docs.cube.dev/docs/data-modeling/extending-cubes
Use extends on cubes to inherit and merge members from a parent so shared measures, dimensions, and joins stay defined once.
The `extends` parameter, supported for [cubes][ref-cube-extends] and
[views][ref-view-extends], allows you to create a *child* cube (or a view) that reuses
all declared members of a *parent* cube (or a view). This helps build reusable data models.
Cubes declare members such as measures, dimensions, and segments. When a child cube extends
the parent cube, lists of measures, dimensions, and segments are merged.
For example, if the parent cube defines the `a` measure and the child cube defines the `b`
measure, the resulting cube will have both measures `a` and `b`.
The usual pattern is to extract common measures, dimensions, and joins into
the parent cube and then extend from it. This helps prevent code duplication
and makes code easier to maintain and refactor.
In the example below, the `base_events` cube defines the common events measures,
dimensions, and a join to the `users` cube:
```yaml title="YAML" theme={"dark"}
cubes:
- name: base_events
sql_table: events
joins:
- name: users
relationship: many_to_one
sql: "{CUBE}.user_id = {users.id}"
measures:
- name: count
type: count
dimensions:
- name: timestamp
sql: time
type: time
```
```javascript title="JavaScript" theme={"dark"}
cube(`base_events`, {
sql_table: `events`,
joins: {
users: {
relationship: `many_to_one`,
sql: `${CUBE}.user_id = ${users.id}`
}
},
measures: {
count: {
type: `count`
}
},
dimensions: {
timestamp: {
sql: `time`,
type: `time`
}
}
})
```
It’s important to use the [`CUBE` variable][ref-cube-variable] when referencing members
and columns of the cube. Not specifying the cube name or using `${base_events}` does not
work when the cube is extended.
The `product_purchases` and `page_views` cubes are extended from `base_events`
and define only the specific dimensions: `product_name` for product purchases
and `page_path` for page views.
```yaml title="YAML" theme={"dark"}
cubes:
- name: product_purchases
sql_table: product_purchases
extends: base_events
dimensions:
- name: product_name
sql: product_name
type: string
- name: page_views
sql_table: page_views
extends: base_events
dimensions:
- name: page_path
sql: page_path
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`product_purchases`, {
sql_table: `product_purchases`,
extends: base_events,
dimensions: {
product_name: {
sql: `product_name`,
type: `string`
}
}
})
cube(`page_views`, {
sql_table: `page_views`,
extends: base_events,
dimensions: {
page_path: {
sql: `page_path`,
type: `string`
}
}
})
```
## Usage with `FILTER_PARAMS`
If the parent cube is using [`FILTER_PARAMS`][ref-schema-ref-cube-filter-params]
in any `sql` parameter, then child cubes can accomodate to that in two ways.
First, the `sql` parameter can be overridden in each child cube:
```yaml title="YAML" theme={"dark"}
cubes:
- name: product_purchases
sql: |
SELECT *
FROM events
WHERE {FILTER_PARAMS.product_purchases.timestamp.filter('time')}
# ...
```
```javascript title="JavaScript" theme={"dark"}
cube(`product_purchases`, {
sql: `
SELECT *
FROM events
WHERE ${FILTER_PARAMS.product_purchases.timestamp.filter("time")}
`,
// ...
})
```
Alternatively, all filters can be put inside the parent cube and referenced
in the child cubes using `AND`. The unused filters will be rendered to `1 = 1`
in the SQL query:
```yaml title="YAML" theme={"dark"}
cubes:
- name: base_events
sql: |
SELECT *
FROM events
WHERE
{FILTER_PARAMS.base_events.timestamp.filter('time')} AND
{FILTER_PARAMS.product_purchases.timestamp.filter('time')} AND
{FILTER_PARAMS.page_views.timestamp.filter('time')}
# ...
```
```javascript title="JavaScript" theme={"dark"}
cube(`base_events`, {
sql: `
SELECT *
FROM events
WHERE
{$FILTER_PARAMS.base_events.timestamp.filter('time')} AND
{$FILTER_PARAMS.product_purchases.timestamp.filter('time')} AND
{$FILTER_PARAMS.page_views.timestamp.filter('time')}
`,
// ...
})
```
[ref-cube-extends]: /reference/data-modeling/cube#extends
[ref-view-extends]: /reference/data-modeling/view#extends
[ref-schema-ref-cube-filter-params]: /reference/data-modeling/context-variables#filter_params
[ref-cube-variable]: /docs/data-modeling/concepts/syntax#cube-variable
# Joins
Source: https://docs.cube.dev/docs/data-modeling/joins
Joins define relationships between cubes, allowing Cube to automatically generate multi-table SQL queries when views combine data from multiple cubes.
Joins define how cubes connect to each other. When a [view][ref-views]
includes members from multiple cubes, Cube uses these relationships to
automatically generate SQL `JOIN` clauses — so end-users can explore data
across tables without writing SQL.
See the [joins reference][ref-schema-ref-joins-relationship] for the full
list of parameters and configuration options.
## Relationship types
Cube supports three relationship types: `one_to_one`, `one_to_many`, and
`many_to_one`. The relationship type determines which table becomes the left
side of the `LEFT JOIN` in the generated SQL.
Consider two cubes, `orders` and `customers`. An order belongs to one
customer, but a customer can have many orders:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
joins:
- name: customers
relationship: many_to_one
sql: "{CUBE}.customer_id = {customers.id}"
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: status
sql: status
type: string
measures:
- name: count
type: count
- name: customers
sql_table: customers
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: company
sql: company
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
joins: {
customers: {
relationship: `many_to_one`,
sql: `${CUBE}.customer_id = ${customers.id}`
}
},
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true },
status: { sql: `status`, type: `string` }
},
measures: {
count: { type: `count` }
}
})
cube(`customers`, {
sql_table: `customers`,
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true },
company: { sql: `company`, type: `string` }
}
})
```
The `many_to_one` join on `orders` means: many orders belong to one customer.
When a view includes members from both cubes, Cube generates SQL with `orders`
on the left and `customers` on the right:
```sql theme={"dark"}
SELECT
"orders".status,
"customers".company,
COUNT("orders".id)
FROM orders AS "orders"
LEFT JOIN customers AS "customers"
ON "orders".customer_id = "customers".id
GROUP BY 1, 2
```
Because `orders` is on the left side of the `LEFT JOIN`, all orders are
preserved — including guest checkouts with no matching customer.
As a rule of thumb, define joins on the **fact table** (e.g., `orders`)
pointing toward the **dimension table** (e.g., `customers`) using
`many_to_one`. This ensures the fact table is always the base of the query,
preserving all its rows.
### Many-to-many relationships
A many-to-many relationship requires an associative (junction) table. For
example, `posts` and `topics` are connected through a `post_topics` table:
Model this with an associative cube, chaining the joins so they flow in one
direction (`posts → post_topics → topics`):
```yaml title="YAML" theme={"dark"}
cubes:
- name: posts
sql_table: posts
joins:
- name: post_topics
relationship: one_to_many
sql: "{CUBE}.id = {post_topics.post_id}"
- name: post_topics
sql_table: post_topics
joins:
- name: topics
relationship: many_to_one
sql: "{CUBE}.topic_id = {topics.id}"
dimensions:
- name: id
sql: "CONCAT({CUBE}.post_id, {CUBE}.topic_id)"
type: string
primary_key: true
- name: topics
sql_table: topics
dimensions:
- name: id
sql: id
type: string
primary_key: true
- name: name
sql: name
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`posts`, {
sql_table: `posts`,
joins: {
post_topics: {
relationship: `one_to_many`,
sql: `${CUBE}.id = ${post_topics.post_id}`
}
}
})
cube(`post_topics`, {
sql_table: `post_topics`,
joins: {
topics: {
relationship: `many_to_one`,
sql: `${CUBE}.topic_id = ${topics.id}`
}
},
dimensions: {
id: {
sql: `CONCAT(${CUBE}.post_id, ${CUBE}.topic_id)`,
type: `string`,
primary_key: true
}
}
})
cube(`topics`, {
sql_table: `topics`,
dimensions: {
id: { sql: `id`, type: `string`, primary_key: true },
name: { sql: `name`, type: `string` }
}
})
```
A view can then expose this through the `join_path`:
```yaml theme={"dark"}
views:
- name: posts_with_topics
cubes:
- join_path: posts
includes:
- title
- count
- join_path: posts.post_topics.topics
prefix: true
includes:
- name
```
## Direction of joins
**All joins are directed.** They flow from the source cube (where the join
is defined) to the target cube (the one referenced). Cube places the source
cube on the left side of the `LEFT JOIN` and the target on the right.
This matters because the left table preserves all its rows, while the right
table contributes matching rows or `NULL`. The direction you choose affects
which records appear in the result set.
For example, if `orders` defines a `many_to_one` join to `customers`:
* `orders` is the base → all orders are preserved, even guest checkouts
* `customers` without orders won't appear
If instead `customers` defined a `one_to_many` join to `orders`:
* `customers` is the base → all customers are preserved, even those without orders
* Guest checkout orders (with no matching customer) won't appear
### Using views to control direction
Views let you control which join path is followed via the
[`join_path`][ref-view-join-path] parameter. This is the recommended way to
handle cases where you need different join directions for different use cases:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
joins:
- name: customers
sql: "{CUBE}.customer_id = {customers.id}"
relationship: many_to_one
measures:
- name: count
type: count
- name: total_revenue
sql: revenue
type: sum
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: customers
sql_table: customers
joins:
- name: orders
sql: "{CUBE}.id = {orders.customer_id}"
relationship: one_to_many
measures:
- name: count
type: count
dimensions:
- name: id
sql: id
type: number
primary_key: true
- name: name
sql: name
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
joins: {
customers: {
sql: `${CUBE}.customer_id = ${customers.id}`,
relationship: `many_to_one`
}
},
measures: {
count: { type: `count` },
total_revenue: { sql: `revenue`, type: `sum` }
},
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true }
}
})
cube(`customers`, {
sql_table: `customers`,
joins: {
orders: {
sql: `${CUBE}.id = ${orders.customer_id}`,
relationship: `one_to_many`
}
},
measures: {
count: { type: `count` }
},
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true },
name: { sql: `name`, type: `string` }
}
})
```
Now you can create two views for two different analytical needs:
```yaml title="YAML" theme={"dark"}
views:
- name: revenue_per_customer
description: All orders with customer details. Includes guest checkouts.
cubes:
- join_path: orders
includes:
- count
- total_revenue
- join_path: orders.customers
includes:
- name
- name: customer_activity
description: All customers with their order activity. Includes customers without orders.
cubes:
- join_path: customers
includes:
- name
- count
- join_path: customers.orders
prefix: true
includes:
- count
- total_revenue
```
```javascript title="JavaScript" theme={"dark"}
view(`revenue_per_customer`, {
description: `All orders with customer details. Includes guest checkouts.`,
cubes: [
{
join_path: orders,
includes: [`count`, `total_revenue`]
},
{
join_path: orders.customers,
includes: [`name`]
}
]
})
view(`customer_activity`, {
description: `All customers with their order activity. Includes customers without orders.`,
cubes: [
{
join_path: customers,
includes: [`name`, `count`]
},
{
join_path: customers.orders,
prefix: true,
includes: [`count`, `total_revenue`]
}
]
})
```
The `revenue_per_customer` view follows the `orders → customers` path, so all
orders are preserved. The `customer_activity` view follows
`customers → orders`, so all customers are preserved.
## Diamond subgraphs
A *diamond subgraph* occurs when there's more than one join path between two
cubes — for example, `users.schools.countries` and
`users.employers.countries`. This can lead to ambiguous query generation.
Views resolve this ambiguity by specifying the exact `join_path` for each
included cube. For example, if cube `a` joins to both `b` and `c`, and both
`b` and `c` join to `d`, a view can specify which path to follow:
```yaml theme={"dark"}
views:
- name: a_with_d_via_b
cubes:
- join_path: a
includes: "*"
- join_path: a.b.d
prefix: true
includes:
- value
- name: a_with_d_via_c
cubes:
- join_path: a
includes: "*"
- join_path: a.c.d
prefix: true
includes:
- value
```
Each view follows a specific, unambiguous path through the data graph.
## Join paths in calculated members
When referencing a member of another cube in a [calculated member][ref-calculated-members],
you can use a join path to specify the exact route. This uses dot-separated
cube names:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
dimensions:
- name: customer_country
sql: "{customers.country}"
type: string
- name: shipping_country
sql: "{shipping_addresses.country}"
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
dimensions: {
customer_country: {
sql: `${customers.country}`,
type: `string`
},
shipping_country: {
sql: `${shipping_addresses.country}`,
type: `string`
}
}
})
```
## Troubleshooting
### `Can't find join path`
The error `Can't find join path to join 'cube_a', 'cube_b'` means the cubes
included in a view or query can't be connected through the defined joins.
Check that:
* Joins are defined with the correct [direction](#direction-of-joins)
* There is a continuous path from the source cube to the target cube
* You're using the [`join_path`][ref-view-join-path] parameter in views to
specify the exact path
### `Primary key is required when join is defined`
Cube uses primary keys to avoid fanouts — when rows get duplicated during
joins and aggregates are over-counted. Define a [primary key][ref-primary-key]
dimension in every cube that participates in joins.
If your data doesn't have a natural primary key, create a composite one:
```yaml theme={"dark"}
cubes:
- name: events
# ...
dimensions:
- name: composite_key
sql: CONCAT(column_a, '-', column_b, '-', column_c)
type: string
primary_key: true
```
[ref-schema-ref-joins-relationship]: /reference/data-modeling/joins
[ref-views]: /docs/data-modeling/views
[ref-view-join-path]: /reference/data-modeling/view#join_path
[ref-calculated-members]: /docs/data-modeling/measures#calculated-measures
[ref-primary-key]: /reference/data-modeling/dimensions#primary_key
[ref-visual-model]: /docs/data-modeling/visual-modeler
# Measures
Source: https://docs.cube.dev/docs/data-modeling/measures
Measures compute aggregated values across rows — counts, sums, averages, and more complex calculations like rolling windows, time shifts, and rankings.
While [dimensions][ref-dimensions-page] describe attributes of individual rows,
measures compute values across rows — sums, counts, averages, and other
aggregations. Measures can aggregate columns directly (like `sum of revenue`)
or reference other measures to create compound metrics (like `revenue / count`).
See the [measures reference][ref-measures-ref] for the full list of parameters
and configuration options.
## Defining measures
A measure specifies the SQL expression to aggregate and the aggregation type:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
sql_table: orders
measures:
- name: count
type: count
- name: total_amount
sql: amount
type: sum
- name: average_amount
sql: amount
type: avg
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
sql_table: `orders`,
measures: {
count: { type: `count` },
total_amount: { sql: `amount`, type: `sum` },
average_amount: { sql: `amount`, type: `avg` }
}
})
```
## Filtered measures
You can apply [filters][ref-filters] to a measure to create conditional
aggregations. Only rows matching the filter are included:
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
measures:
- name: count
type: count
- name: completed_count
type: count
filters:
- sql: "{CUBE}.status = 'completed'"
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
measures: {
count: { type: `count` },
completed_count: {
type: `count`,
filters: [{ sql: `${CUBE}.status = 'completed'` }]
}
}
})
```
When `completed_count` is queried, Cube generates SQL with a `CASE` expression:
```sql theme={"dark"}
SELECT
COUNT(CASE WHEN (orders.status = 'completed') THEN 1 END) AS completed_count
FROM orders
```
## Calculated measures
Calculated measures perform calculations on other measures using SQL functions
and operators. They provide a way to decompose complex metrics (e.g., ratios
or percents) into formulas involving simpler measures.
### Referencing measures in the same cube
```yaml title="YAML" theme={"dark"}
cubes:
- name: orders
# ...
measures:
- name: count
type: count
- name: completed_count
type: count
filters:
- sql: "{CUBE}.status = 'completed'"
- name: completed_ratio
sql: "1.0 * {completed_count} / NULLIF({count}, 0)"
type: number
```
```javascript title="JavaScript" theme={"dark"}
cube(`orders`, {
// ...
measures: {
count: { type: `count` },
completed_count: {
type: `count`,
filters: [{ sql: `${CUBE}.status = 'completed'` }]
},
completed_ratio: {
sql: `1.0 * ${completed_count} / NULLIF(${count}, 0)`,
type: `number`
}
}
})
```
### Referencing measures from other cubes
If cubes are [joined][ref-joins], you can reference measures across cubes.
Cube generates the necessary joins automatically:
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
# ...
joins:
- name: orders
sql: "{CUBE}.id = {orders}.user_id"
relationship: one_to_many
measures:
- name: count
type: count
- name: purchases_to_users_ratio
sql: "1.0 * {orders.purchases} / NULLIF({CUBE.count}, 0)"
type: number
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
// ...
joins: {
orders: {
sql: `${CUBE}.id = ${orders}.user_id`,
relationship: `one_to_many`
}
},
measures: {
count: { type: `count` },
purchases_to_users_ratio: {
sql: `1.0 * ${orders.purchases} / NULLIF(${CUBE.count}, 0)`,
type: `number`
}
}
})
```
## Multi-stage measures
Multi-stage measures are calculated in two or more stages, enabling
calculations on already-aggregated data. Each stage results in one or more
CTEs in the generated SQL query.
Multi-stage measures are powered by Tesseract, the [next-generation data
modeling engine][link-tesseract]. In versions before v1.7.0, it was not enabled by default.
### Rolling windows
Rolling window measures calculate metrics over a moving window of time, such
as cumulative counts or moving averages. Use the
[`rolling_window`][ref-rolling-window] parameter:
```yaml theme={"dark"}
measures:
- name: cumulative_count
type: count
rolling_window:
trailing: unbounded
- name: trailing_month_count
sql: id
type: count
rolling_window:
trailing: 1 month
```
### Period-to-date
Period-to-date measures analyze data from the start of a period to the current
date — year-to-date (YTD), quarter-to-date (QTD), or month-to-date (MTD):
```yaml theme={"dark"}
measures:
- name: revenue_ytd
sql: revenue
type: sum
rolling_window:
type: to_date
granularity: year
- name: revenue_qtd
sql: revenue
type: sum
rolling_window:
type: to_date
granularity: quarter
```
### Time shift
Time-shift measures calculate the value of another measure at a different
point in time, typically for period-over-period comparisons like
year-over-year growth. Use the [`time_shift`][ref-time-shift] parameter:
```yaml theme={"dark"}
measures:
- name: revenue
sql: revenue
type: sum
- name: revenue_prior_year
multi_stage: true
sql: "{revenue}"
type: number
time_shift:
- interval: 1 year
type: prior
```
You can combine time shift with period-to-date for comparisons like
"this year's YTD vs. last year's YTD":
```yaml theme={"dark"}
measures:
- name: revenue_ytd
sql: revenue
type: sum
rolling_window:
type: to_date
granularity: year
- name: revenue_prior_year_ytd
multi_stage: true
sql: "{revenue_ytd}"
type: number
time_shift:
- time_dimension: time
interval: 1 year
type: prior
```
Time-shift measures can also be used with [calendar cubes][ref-calendar-cubes]
to customize how time-shifting works, e.g., to shift by retail calendar
periods.
### Percent of total (fixed dimension)
Use the [`grain`][ref-grain] parameter with `keep_only` to fix the inner
aggregation to specific dimensions, enabling percent-of-total calculations:
```yaml theme={"dark"}
measures:
- name: revenue
sql: revenue
type: sum
- name: country_revenue
multi_stage: true
sql: "{revenue}"
type: sum
grain:
keep_only:
- country
- name: country_revenue_percentage
multi_stage: true
sql: "{revenue} / NULLIF({country_revenue}, 0)"
type: number
```
### Share of total (filter override)
Use the [`filter`][ref-filter] parameter to override the filters that a
multi-stage measure inherits from the query. This enables "share of total"
calculations where the denominator must ignore a filter applied by the query.
In the example below, `amount_all_statuses` uses `exclude` to drop the `status`
filter, so it always aggregates across all statuses. When the query is filtered
to a single status, `total_amount` reflects that status while
`amount_all_statuses` stays the full per-category total, and
`percent_of_total` is the share that the filtered status represents:
```yaml theme={"dark"}
measures:
- name: total_amount
sql: amount
type: sum
- name: amount_all_statuses
multi_stage: true
sql: "{total_amount}"
type: number
filter:
exclude:
- status
- name: percent_of_total
multi_stage: true
sql: "100.0 * {total_amount} / NULLIF({amount_all_statuses}, 0)"
type: number
format: percent
```
### Nested aggregates
Use the [`grain`][ref-grain] parameter with `include` to compute an aggregate
of an aggregate, e.g., the average of per-customer averages:
```yaml theme={"dark"}
measures:
- name: avg_order_value
sql: amount
type: avg
- name: avg_customer_order_value
multi_stage: true
sql: "{avg_order_value}"
type: avg
grain:
include:
- customer_id
```
When a nested aggregate combines two or more other multi-stage measures that
share the same grain, set [`grain`][ref-grain] on the **combining** measure —
not on each of its inputs. For example, to average a per-day ratio, group the
per-day components by day through the combining measure:
```yaml theme={"dark"}
measures:
- name: total_amount
sql: amount
type: sum
- name: total_count
sql: id
type: count
# Intermediate multi-stage measures — the grain is set on the combining
# measure below, so these inherit it and are joined on the shared grain.
- name: daily_amount
multi_stage: true
sql: "{total_amount}"
type: number
- name: daily_count
multi_stage: true
sql: "{total_count}"
type: number
- name: avg_daily_order_value
multi_stage: true
sql: "1.0 * {daily_amount} / NULLIF({daily_count}, 0)"
type: avg
grain:
include:
- created_at
```
`grain` fixes the inner grain of the measure it's declared on and does not
expose the added dimension to a measure built on top of it. If each input
measure declares the same `grain.include` (rather than the combining measure),
the inputs no longer carry a shared grouping key, so they are combined with a
cross join instead of being joined on that key — producing incorrect results.
Declare `grain` on the combining measure so its inputs are joined on the shared
grain.
### Ranking
Use the [`grain`][ref-grain] parameter with `exclude` to rank items within
groups:
```yaml theme={"dark"}
measures:
- name: revenue
sql: revenue
type: sum
- name: product_rank
multi_stage: true
order_by:
- sql: "{revenue}"
dir: asc
grain:
exclude:
- product
type: rank
```
`grain` replaces the standalone `group_by`, `reduce_by`, and `add_group_by`
parameters, which remain supported. See the [`grain`][ref-grain] reference for
the migration mapping.
### Conditional measures
Conditional measures depend on the value of a dimension, using the
[`case`][ref-case] parameter with [`switch` dimensions][ref-switch-dim]:
```yaml theme={"dark"}
measures:
- name: amount_in_currency
multi_stage: true
case:
switch: "{CUBE.currency}"
when:
- value: EUR
sql: "{CUBE.amount_eur}"
- value: GBP
sql: "{CUBE.amount_gbp}"
else:
sql: "{CUBE.amount_usd}"
type: number
```
## Formatting
Use the [`format`][ref-format] parameter to control how measures are displayed:
```yaml theme={"dark"}
measures:
- name: total_revenue
sql: revenue
type: sum
format: currency
- name: conversion_rate
sql: "1.0 * {completed_count} / NULLIF({count}, 0)"
type: number
format: percent
```
## Next steps
* See the [measures reference][ref-measures-ref] for all parameters
* Learn about [dimensions][ref-dimensions-page] for grouping and filtering
* Explore [pre-aggregations][ref-pre-aggs] to accelerate measure queries
* See the [period-over-period recipe][ref-pop-recipe] for advanced time
comparisons
[ref-measures-ref]: /reference/data-modeling/measures
[ref-dimensions-page]: /docs/data-modeling/dimensions
[ref-joins]: /docs/data-modeling/joins
[ref-pre-aggs]: /reference/data-modeling/pre-aggregations
[ref-type]: /reference/data-modeling/measures#type
[ref-filters]: /reference/data-modeling/measures#filters
[ref-format]: /reference/data-modeling/measures#format
[ref-rolling-window]: /reference/data-modeling/measures#rolling_window
[ref-time-shift]: /reference/data-modeling/measures#time_shift
[ref-grain]: /reference/data-modeling/measures#grain
[ref-filter]: /reference/data-modeling/measures#filter
[ref-case]: /reference/data-modeling/measures#case
[ref-switch-dim]: /reference/data-modeling/dimensions#type
[ref-calendar-cubes]: /docs/data-modeling/concepts/calendar-cubes
[ref-pop-recipe]: /recipes/data-modeling/period-over-period
[link-tesseract]: https://cube.dev/blog/introducing-next-generation-data-modeling-engine
# Multi-fact views
Source: https://docs.cube.dev/docs/data-modeling/multi-fact-views
Analyze data across multiple fact tables that share common dimensions like time or customers, without row multiplication or manual workarounds.
In many data models, you have multiple fact tables that share common
dimensions but have no direct relationship to each other. For example,
an e-commerce company tracks both orders and returns:
* **`orders`** — one row per order, with `customer_id` and `created_at`
* **`returns`** — one row per return, with `customer_id` and `created_at`
* **`customers`** — one row per customer
* **`dates`** — a date spine
Both `orders` and `returns` join to `customers` and `dates`, but they don't
join to each other:
```
customers
/ \
orders returns
\ /
dates
```
You need a report showing `orders_count`, `total_revenue`, `returns_count`,
and `total_refunds` grouped by customer and month. But joining `orders` and
`returns` directly would produce a cross product — every order matched with
every return for that customer and date — inflating all counts and sums.
## How multi-fact views solve this
In a regular [view][ref-views], there is a single **root cube** — the first
cube listed in the view's `cubes` array. All joins flow from this root, and
Cube uses it as the base table in the generated SQL.
Multi-fact views work differently. When a view includes measures from
**multiple fact tables**, Cube selects the root dynamically at query time
based on which measures are requested. Each fact table gets its own
aggregating subquery, and the results are joined on the shared dimensions.
No fanout, no manual workarounds.
Multi-fact views are powered by Tesseract, the [next-generation data modeling
engine][link-tesseract]. In versions before v1.7.0, it was not enabled by default.
## How to model it
### 1. Define the cubes
Each fact table becomes a cube with explicit joins to the shared dimension
tables:
```yaml title="YAML" theme={"dark"}
cubes:
- name: customers
sql_table: customers
dimensions:
- name: id
type: number
sql: id
primary_key: true
- name: name
type: string
sql: name
- name: city
type: string
sql: city
- name: dates
sql_table: dates
dimensions:
- name: date
type: time
sql: date
primary_key: true
- name: orders
sql_table: orders
joins:
- name: customers
relationship: many_to_one
sql: "{orders}.customer_id = {customers.id}"
- name: dates
relationship: many_to_one
sql: "DATE_TRUNC('day', {orders}.created_at) = {dates.date}"
dimensions:
- name: id
type: number
sql: id
primary_key: true
- name: status
type: string
sql: status
measures:
- name: count
type: count
- name: total_amount
type: sum
sql: amount
- name: returns
sql_table: returns
joins:
- name: customers
relationship: many_to_one
sql: "{returns}.customer_id = {customers.id}"
- name: dates
relationship: many_to_one
sql: "DATE_TRUNC('day', {returns}.created_at) = {dates.date}"
dimensions:
- name: id
type: number
sql: id
primary_key: true
measures:
- name: count
type: count
- name: total_refund
type: sum
sql: refund_amount
```
```javascript title="JavaScript" theme={"dark"}
cube(`customers`, {
sql_table: `customers`,
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true },
name: { sql: `name`, type: `string` },
city: { sql: `city`, type: `string` }
}
})
cube(`dates`, {
sql_table: `dates`,
dimensions: {
date: { sql: `date`, type: `time`, primary_key: true }
}
})
cube(`orders`, {
sql_table: `orders`,
joins: {
customers: {
relationship: `many_to_one`,
sql: `${orders}.customer_id = ${customers.id}`
},
dates: {
relationship: `many_to_one`,
sql: `DATE_TRUNC('day', ${orders}.created_at) = ${dates.date}`
}
},
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true },
status: { sql: `status`, type: `string` }
},
measures: {
count: { type: `count` },
total_amount: { sql: `amount`, type: `sum` }
}
})
cube(`returns`, {
sql_table: `returns`,
joins: {
customers: {
relationship: `many_to_one`,
sql: `${returns}.customer_id = ${customers.id}`
},
dates: {
relationship: `many_to_one`,
sql: `DATE_TRUNC('day', ${returns}.created_at) = ${dates.date}`
}
},
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true }
},
measures: {
count: { type: `count` },
total_refund: { sql: `refund_amount`, type: `sum` }
}
})
```
The critical detail: both `orders` and `returns` declare direct joins to
`customers` and `dates`. This tells Cube that these dimension tables are shared
between the two facts.
### 2. Create a view
The view brings both fact tables and the shared dimension tables together.
Dimension tables are included at root-level join paths (not nested under a
specific fact), which makes their dimensions common to both facts. Use
`prefix` to disambiguate identically named members across fact cubes:
```yaml title="YAML" theme={"dark"}
views:
- name: customer_overview
cubes:
- join_path: orders
prefix: true
includes:
- count
- total_amount
- join_path: returns
prefix: true
includes:
- count
- total_refund
- join_path: customers
includes:
- name
- city
- join_path: dates
includes:
- date
```
```javascript title="JavaScript" theme={"dark"}
view(`customer_overview`, {
cubes: [
{
join_path: orders,
prefix: true,
includes: [`count`, `total_amount`]
},
{
join_path: returns,
prefix: true,
includes: [`count`, `total_refund`]
},
{
join_path: customers,
includes: [`name`, `city`]
},
{
join_path: dates,
includes: [`date`]
}
]
})
```
When you query `orders_count`, `orders_total_amount`, `returns_count`, and
`returns_total_refund` grouped by `name`, `city`, and `date`, Cube detects
the two separate fact roots and automatically executes a multi-fact query.
## What Cube does under the hood
Cube executes the query in three stages:
### 1. Separate aggregating subqueries
Each fact table gets its own independent subquery that joins only the tables
it needs, applies relevant filters, and aggregates by the common dimensions:
* **Subquery 1** (orders): joins `orders` → `customers` and `orders` → `dates`,
computes `COUNT(*)` and `SUM(amount)`, grouped by `name`, `city`, `date`
* **Subquery 2** (returns): joins `returns` → `customers` and `returns` → `dates`,
computes `COUNT(*)` and `SUM(refund_amount)`, grouped by `name`, `city`, `date`
### 2. Join on common dimensions
The subquery results are joined with `FULL JOIN` on all common dimension
columns (`name`, `city`, `date`). This preserves rows that exist in only one
fact table — a customer who placed orders but never returned anything still
appears in the results.
### 3. Final result
The combined result shows measures from each fact table side by side:
| name | city | date | orders\_count | orders\_total\_amount | returns\_count | returns\_total\_refund |
| ------- | -------- | ---------- | ------------- | --------------------- | -------------- | ---------------------- |
| Alice | New York | 2025-01-15 | 2 | 200.00 | 0 | NULL |
| Alice | New York | 2025-02-10 | 2 | 225.00 | 1 | 100.00 |
| Bob | Seattle | 2025-01-20 | 3 | 550.00 | 2 | 130.00 |
| Charlie | New York | 2025-02-05 | 0 | NULL | 2 | 100.00 |
| Diana | Boston | 2025-03-01 | 1 | 400.00 | 0 | NULL |
Charlie has no orders and Diana has no returns — both are still included
with `NULL` values for the missing fact table.
## Combining facts in one measure
Putting measures from two facts side by side is often not the goal — you want a
single metric derived from both, such as revenue per order where revenue and
order count come from different fact tables. Neither cube can define it, because
neither can reference the other's measures.
Define it as a [measure of the view][ref-view-measures] instead, and mark it
[`multi_stage`][ref-multi-stage]:
```yaml title="YAML" theme={"dark"}
views:
- name: customer_overview
cubes:
- join_path: orders
prefix: true
includes:
- count
- total_amount
- join_path: returns
prefix: true
includes:
- total_refund
- join_path: customers
includes:
- name
- city
- join_path: dates
includes:
- date
measures:
- name: refund_rate
type: number
multi_stage: true
sql: "{CUBE.returns_total_refund} / NULLIF({CUBE.orders_total_amount}, 0)"
```
```javascript title="JavaScript" theme={"dark"}
view(`customer_overview`, {
cubes: [
{
join_path: orders,
prefix: true,
includes: [`count`, `total_amount`]
},
{
join_path: returns,
prefix: true,
includes: [`total_refund`]
},
{
join_path: customers,
includes: [`name`, `city`]
},
{
join_path: dates,
includes: [`date`]
}
],
measures: {
refund_rate: {
type: `number`,
multi_stage: true,
sql: `${CUBE.returns_total_refund} / NULLIF(${CUBE.orders_total_amount}, 0)`
}
}
})
```
`multi_stage` is what makes this work. It defers the expression to a stage that
runs *after* the per-fact subqueries have been aggregated and joined, so the
division happens once per row of the combined result:
```sql theme={"dark"}
-- one aggregating subquery per fact, at the query's grain
SUM(orders.amount) GROUP BY city
SUM(returns.refund) GROUP BY city
-- final stage, once the two are joined on city
total_refund / NULLIF(total_amount, 0)
```
Without `multi_stage`, the same expression is planned as an ordinary calculated
measure. Cube then looks for a single join tree covering both fact cubes, finds
none — the facts only meet through the shared dimensions — and the query fails
with `Can't find join path to join …`, naming both facts. If you see that error
on a measure that spans facts, `multi_stage` is what's missing.
The measure is queried like any other, on its own or next to its components, and
grouped by any of the shared dimensions.
## Joining views in the SQL API
You don't have to define a dedicated multi-fact view to get multi-fact
behavior. The [SQL API][ref-sql-api] produces the same query when you **join
two or more views on a dimension they share** and group by that dimension.
Suppose `orders_view` and `returns_view` are two separate views that each
expose the customer's `name` (both backed by the same underlying
`customers.name` member). Joining them on `name` and grouping by it triggers a
multi-fact query:
```sql theme={"dark"}
SELECT
o.name,
MEASURE(o.total_amount),
MEASURE(r.total_refund)
FROM orders_view o
LEFT JOIN returns_view r ON r.name = o.name
GROUP BY 1
```
Cube recognizes that both `name` columns resolve to the same cube member,
merges the two view scans into a single multi-fact query, and runs it with the
separate-subquery-then-join strategy described
[above](#what-cube-does-under-the-hood).
This rewrite applies only when:
* Both sides of the join condition resolve to the **same underlying cube
member** (a shared dimension), and the join key is composed only of
dimensions.
* The query is **grouped by the join key** — every grouped dimension is the
shared join key. Ungrouped joins (such as `SELECT *`) and queries that group
by a different dimension are not merged and fall back to standard join
handling.
### Joining three or more views
The rewrite is not limited to two views. Chained joins on the same shared key
are merged into a single multi-fact query, with each view contributing its own
aggregating subquery:
```sql theme={"dark"}
SELECT
o.name,
MEASURE(o.total_amount),
MEASURE(r.total_refund),
MEASURE(p.total_paid)
FROM orders_view o
FULL JOIN returns_view r ON r.name = o.name
FULL JOIN payments_view p ON p.name = o.name
GROUP BY 1
```
### Joining on a time dimension
A common multi-fact pattern joins facts on a shared time dimension and groups by
a truncated grain. **Join on `DATE_TRUNC` at the same granularity you group by:**
```sql theme={"dark"}
SELECT DATE_TRUNC('day', o.created_at), MEASURE(o.total_amount), MEASURE(r.total_refund)
FROM orders_view o
JOIN returns_view r ON DATE_TRUNC('day', r.created_at) = DATE_TRUNC('day', o.created_at)
GROUP BY 1
```
The grouped column is emitted as a time dimension with its granularity. A join
written on `DATE_TRUNC` is an `INNER` join (the SQL planner expresses it as a
filtered cross join), so both sides must share a key; both truncated columns
must resolve to the same underlying time member at the same granularity.
The join-key granularity must match the `GROUP BY` granularity, because the
facts are stitched together at the grain you group by. This has two
consequences:
* Joining on `DATE_TRUNC('month', …)` while grouping by `DATE_TRUNC('day', …)`
is not merged (it would silently stitch at day grain, diverging from the
month-grain join).
* Joining on the **raw** time column (`ON r.created_at = o.created_at`, an
exact-timestamp join) while grouping by `DATE_TRUNC('day', …)` is likewise not
merged — the row-grain join doesn't match the day-grain group-by. Truncate the
join key to the grain you group by instead.
In both cases the query falls back to standard join handling.
You can also combine a `DATE_TRUNC` equality with a plain dimension equality in
the same join (a composite key), and group by both:
```sql theme={"dark"}
SELECT DATE_TRUNC('day', o.created_at), o.name, MEASURE(o.total_amount), MEASURE(r.total_refund)
FROM orders_view o
JOIN returns_view r
ON DATE_TRUNC('day', r.created_at) = DATE_TRUNC('day', o.created_at)
AND r.name = o.name
GROUP BY 1, 2
```
### Filtering the join
Filters on top of the join are pushed into the merged query:
* A `WHERE` clause is pushed into the merged scan, becoming a filter on the
member the predicate refers to.
* A predicate in the `ON` clause that the planner can attach to a single side
(for example, a condition on the optional side of a `LEFT JOIN`) becomes a
filter on that fact. Predicates that the SQL planner can't push to one side
of an outer join (such as a left-table condition in a `LEFT JOIN ON`) aren't
supported by the planner and will raise an error.
Pushing the predicate in is only the first step: the merged query is then
planned like any other multi-fact query, so the member it filters on must be
[shared by all facts](#filters-and-segments).
### Join type
The facts are stitched together with a `FULL JOIN` on the shared key, and the
`JOIN` type in your SQL controls which rows are kept:
| SQL join | Result |
| ------------------- | ------------------------------------------------------------------------- |
| `FULL [OUTER] JOIN` | every key from either view (default multi-fact behavior) |
| `INNER JOIN` | only keys present in **both** views |
| `LEFT JOIN` | every key from the left view; right-side measures are `NULL` when missing |
| `RIGHT JOIN` | every key from the right view; left-side measures are `NULL` when missing |
## Common patterns
### Time as the shared dimension
The most common multi-fact pattern uses time as the shared dimension.
For example, you might have `page_views`, `signups`, and `purchases` that all
have timestamps but no direct relationship. By joining each to a shared
`dates` cube, you can analyze conversion funnels — page views vs. signups
vs. purchases by day — without any row multiplication.
### More than two fact tables
Multi-fact queries are not limited to two fact tables. If a view includes
three or more facts, each gets its own aggregating subquery, and all results
are joined on the common dimensions.
### Facts that don't share all dimensions
Every root fact table must be joinable to the **same set of common dimension
tables**. If a fact table doesn't naturally have a foreign key for one of the
common dimensions, you can create a synthetic join:
```yaml title="YAML" theme={"dark"}
cubes:
- name: refunds
sql: >
SELECT *, NULL AS customer_id FROM refunds
joins:
- name: customers
relationship: many_to_one
sql: "{refunds}.customer_id = {customers.id}"
- name: dates
relationship: many_to_one
sql: "DATE_TRUNC('day', {refunds}.created_at) = {dates.date}"
dimensions:
- name: id
type: number
sql: id
primary_key: true
measures:
- name: count
type: count
- name: total_amount
type: sum
sql: amount
```
```javascript title="JavaScript" theme={"dark"}
cube(`refunds`, {
sql: `SELECT *, NULL AS customer_id FROM refunds`,
joins: {
customers: {
relationship: `many_to_one`,
sql: `${refunds}.customer_id = ${customers.id}`
},
dates: {
relationship: `many_to_one`,
sql: `DATE_TRUNC('day', ${refunds}.created_at) = ${dates.date}`
}
},
dimensions: {
id: { sql: `id`, type: `number`, primary_key: true }
},
measures: {
count: { type: `count` },
total_amount: { sql: `amount`, type: `sum` }
}
})
```
The `NULL AS customer_id` makes the join syntactically valid. Refund rows
won't match a specific customer, but the subquery can still participate in
the multi-fact join on the full set of common dimensions.
## Filters and segments
**Common dimension filters** (like `city = 'New York'` or `date > '2025-01-01'`)
are applied to every subquery, ensuring consistent filtering across all facts.
**Measure filters** (like `orders_count > 1`) are applied as `HAVING`
conditions after the subqueries are joined.
**Fact-specific filters and [segments][ref-segments]** — anything that belongs to
one fact table rather than a shared dimension — can't be used in a multi-fact
query, however it is written: a `WHERE` clause in the SQL API, a filter in the
REST (JSON) API, or a segment. Every grouped dimension, filter and segment has
to be reachable from all facts, so a query that carries one fails with
`Can't find join path to join …`. The same members are fine as soon as only
that fact's measures are requested, since the query is no longer multi-fact.
To narrow one fact inside a multi-fact query, put the condition in the measure's
own [`filters`][ref-measure-filters] on its cube. It travels with the measure
into that fact's subquery and leaves the others alone:
```yaml theme={"dark"}
measures:
- name: completed_amount
sql: amount
type: sum
filters:
- sql: "{CUBE}.status = 'completed'"
```
## Join path requirements
* Each fact cube must declare **direct joins** to all shared dimension tables
* Dimension tables should be included in the view at **root-level join paths**,
not nested under a specific fact (e.g., `customers`, not `orders.customers`)
* Use `prefix` on fact cubes to disambiguate identically named members
* Everything a multi-fact query groups or filters by must be shared by all facts
[ref-views]: /docs/data-modeling/views
[ref-view-ref]: /reference/data-modeling/view
[ref-segments]: /reference/data-modeling/segments
[ref-measure-filters]: /reference/data-modeling/measures#filters
[ref-multi-stage]: /reference/data-modeling/measures#multi_stage
[ref-view-measures]: /reference/data-modeling/view#measures
[ref-sql-api]: /reference/core-data-apis/sql-api
[link-tesseract]: https://cube.dev/blog/introducing-tesseract
# Getting started
Source: https://docs.cube.dev/docs/data-modeling/overview
Build a reusable semantic layer that provides the shared context for AI agents, BI dashboards, and embedded analytics — turning warehouse tables into governed metrics and dimensions.
Let’s use a users table with the following columns as an example:
| id | paying | city | company\_name |
| -- | ------ | ------------- | ------------- |
| 1 | true | San Francisco | Pied Piper |
| 2 | true | Palo Alto | Raviga |
| 3 | true | Redwood | Aviato |
| 4 | false | Mountain View | Bream-Hall |
| 5 | false | Santa Cruz | Hooli |
We can start with a set of simple questions about users we want to answer:
* How many users do we have?
* How many paying users?
* What is the percentage of paying users out of the total?
* How many users, paying or not, are from different cities and companies?
We don’t need to write SQL queries for every question, since the data model
allows building well-organized and reusable SQL.
## 1. Creating a Cube
In Cube, [cubes][ref-schema-cube] are used to organize tables and connections
between tables. Usually one cube is created for each table in the database,
such as `users`, `orders`, `products`, etc. In the `sql_table` parameter of the
cube we define a base table for this cube. In our case, the base table is simply
our `users` table.
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
sql_table: users
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
sql_table: `users`
})
```
## 2. Adding Measures and Dimensions
Once the base table is defined, the next step is to add
[measures][ref-schema-measures] and [dimensions][ref-schema-dimensions] to the
cube.
**Measures** are referred to as quantitative data, such as number of units sold,
number of unique visits, profit, and so on.
**Dimensions** are referred to as categorical data, such as state, gender,
product name, or units of time (e.g., day, week, month).
Let's go ahead and create our first measure and two dimensions:
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
sql_table: users
measures:
- name: count
sql: id
type: count
dimensions:
- name: city
sql: city
type: string
- name: company_name
sql: company_name
type: string
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
sql_table: `users`,
measures: {
count: {
sql: `id`,
type: `count`
}
},
dimensions: {
city: {
sql: `city`,
type: `string`
},
company_name: {
sql: `company_name`,
type: `string`
}
}
})
```
Let's break down the above code snippet piece-by-piece. After defining the base
table for the cube (with the `sql_table` property), we create a `count` measure
in the `measures` block. The `count` [type][ref-schema-types-formats] and sql
`id` means that when this measure will be requested via an API, Cube will
generate and execute the following SQL:
```sql theme={"dark"}
SELECT COUNT(id) AS count
FROM users;
```
When we apply a city dimension to the measure to see "Where are users based?",
Cube will generate SQL with a `GROUP BY` clause:
```sql theme={"dark"}
SELECT city, COUNT(id) AS count
FROM users
GROUP BY 1;
```
You can add as many dimensions as you want to your query when you perform
grouping.
## 3. Adding Filters to Measures
Now let's answer the next question – "How many paying users do we have?". To
accomplish this, we will introduce **measure filters**:
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
measures:
- name: count
sql: id
type: count
- name: paying_count
sql: id
type: count
filters:
- sql: "{CUBE}.paying = 'true'"
# ...
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
measures: {
count: {
sql: `id`,
type: `count`
},
paying_count: {
sql: `id`,
type: `count`,
filters: [{ sql: `${CUBE}.paying = 'true'` }]
}
},
// ...
})
```
It is best practice to prefix references to table columns with the name of the
cube or with the `CUBE` constant when referencing the current cube's column.
That's it! Now we have the `paying_count` measure, which shows only our paying
users. When this measure is requested, Cube will generate the following SQL:
```sql theme={"dark"}
SELECT
COUNT(
CASE WHEN (users.paying = 'true') THEN users.id END
) AS paying_count
FROM users
```
Since the `filters` property is an array, you can apply as many filters as
required. `paying_count` can be used with dimensions the same way as a simple
`count`. We can group `paying_count` by `city` and `companyName` simply by
adding these dimensions alongside measures in the requested query.
## 4. Using Calculated Measures
To answer "What is the percentage of paying users out of the total?", we need to
calculate the paying users ratio, which is basically `paying_count / count`.
Cube makes it extremely easy to perform this kind of calculation by defining a
[calculated measure][ref-calculated-measures]. Let's add a new measure to our cube
called `paying_percentage`:
```yaml title="YAML" theme={"dark"}
cubes:
- name: users
measures:
- name: count
sql: id
type: count
- name: paying_count
sql: id
type: count
filters:
- sql: "{CUBE}.paying = 'true'"
- name: paying_percentage
sql: "1.0 * {paying_count} / {count}"
type: number
format: percent
# ...
```
```javascript title="JavaScript" theme={"dark"}
cube(`users`, {
measures: {
count: {
sql: `id`,
type: `count`
},
paying_count: {
sql: `id`,
type: `count`,
filters: [{ sql: `${CUBE}.paying = 'true'` }]
},
paying_percentage: {
sql: `1.0 * ${paying_count} / ${count}`,
type: `number`,
format: `percent`
}
},
// ...
})
```
Here we defined a calculated measure `paying_percentage`, which divides
`paying_count` by `count`. This example shows how you can reference measures
inside other measure definitions. When you request the `paying_percentage`
measure via an API, the following SQL will be generated:
```sql theme={"dark"}
SELECT
1.0 * COUNT(
CASE WHEN (users.paying = 'true') THEN users.id END
) / COUNT(users.id) AS paying_percentage
FROM users
```
As with other measures, `paying_percentage` can be used with dimensions.
## 5. Creating a View
[Views][ref-views] sit on top of cubes and create a facade of your whole data
model, with which data consumers can interact. They are useful for defining
metrics, managing governance, and controlling which part of the data model is
exposed to end-users.
Let's create a view that exposes our users data:
```yaml title="YAML" theme={"dark"}
views:
- name: users_view
cubes:
- join_path: users
includes:
- "*"
```
```javascript title="JavaScript" theme={"dark"}
view(`users_view`, {
cubes: [
{
join_path: users,
includes: `*`
}
]
})
```
End-users query data through views in Cube. This gives you a layer of
abstraction that makes it easier to manage changes to the underlying data
model.
## 6. Next Steps
1. [Explore][ref-explore] your data model
2. Use [Workbooks][ref-workbooks] to save your analysis and present it as a dashboard
[ref-backend-restapi]: /reference/core-data-apis/rest-api/reference
[ref-schema-cube]: /reference/data-modeling/cube
[ref-schema-measures]: /reference/data-modeling/measures
[ref-schema-dimensions]: /reference/data-modeling/dimensions
[ref-schema-types-formats]: /reference/data-modeling/measures#type
[ref-backend-query-format]: /reference/core-data-apis/rest-api/query-format
[ref-demo-deployment]: /admin/deployment#demo-deployments
[ref-apis]: /reference
[ref-calculated-measures]: /docs/data-modeling/measures#calculated-measures
[ref-views]: /reference/data-modeling/view
[ref-explore]: /docs/explore-analyze/explore
[ref-workbooks]: /docs/explore-analyze/workbooks
# View groups
Source: https://docs.cube.dev/docs/data-modeling/view-groups
View groups organize views into named collections by domain or purpose, helping downstream consumers — including AI agents and embedded analytics — navigate large data models.
When a data model contains many [views][ref-views], view groups help organize
them into named collections by domain or purpose — for example, `sales`,
`finance`, or `people`. View groups are exposed through the
[`/v1/meta`][ref-meta-endpoint] API, making it easier for downstream tools,
AI agents, and embedded analytics to present a navigable catalog.
See the [view group reference][ref-view-group-ref] for the full list of
parameters and configuration options.
## Defining a view group
A view group is a top-level entity, defined alongside views. At minimum it
needs a `name`; adding a `title` makes it easier to navigate in downstream
tools.
```yaml title="YAML" theme={"dark"}
view_groups:
- name: sales
title: Sales
```
```javascript title="JavaScript" theme={"dark"}
view_group(`sales`, {
title: `Sales`
})
```
## Assigning views to a group
To assign a view to a group, list its name on the group via the
[`includes`][ref-view-group-includes] parameter. This keeps the full
membership in one place, which makes it easy to review a group at a glance.
```yaml title="YAML" theme={"dark"}
view_groups:
- name: sales
title: Sales
includes:
- orders_overview
- revenue
```
```javascript title="JavaScript" theme={"dark"}
view_group(`sales`, {
title: `Sales`,
includes: [`orders_overview`, `revenue`]
})
```
A view can belong to more than one group — list it under the `includes`
parameter of every group it should appear in.
## Nesting
View groups can be nested, similar to [nested folders][ref-view-nesting]. Add a
nested view group — with its own `name`, `title`, `description`, and
`includes` — directly inside a parent group's `includes`.
```yaml title="YAML" theme={"dark"}
view_groups:
- name: sales
title: Sales
includes:
- orders_overview
- revenue
- name: enterprise_sales
title: Enterprise Sales
includes:
- enterprise_deals
```
```javascript title="JavaScript" theme={"dark"}
view_group(`sales`, {
title: `Sales`,
includes: [
`orders_overview`,
`revenue`,
{
name: `enterprise_sales`,
title: `Enterprise Sales`,
includes: [`enterprise_deals`]
}
]
})
```
## Where view groups live in the model
By [convention][ref-syntax], view groups are typically defined alongside
views in the `model/views` folder — for example, in a dedicated
`view_groups.yml` file. They behave like any other top-level data model
entity and can be split across multiple files as your model grows.
## Next steps
* See the [view group reference][ref-view-group-ref] for the full list of
parameters
* Learn about [views][ref-views] and how they curate cubes for downstream
consumers
* Explore [AI context][ref-ai-context] to improve AI query accuracy
[ref-views]: /docs/data-modeling/views
[ref-view-nesting]: /reference/data-modeling/view#nesting
[ref-syntax]: /docs/data-modeling/concepts/syntax
[ref-ai-context]: /docs/data-modeling/ai-context
[ref-view-group-ref]: /reference/data-modeling/view-group
[ref-view-group-includes]: /reference/data-modeling/view-group#includes
[ref-meta-endpoint]: /reference/core-data-apis/rest-api/reference
# Views
Source: https://docs.cube.dev/docs/data-modeling/views
Views are curated datasets that sit on top of cubes and create a user-friendly facade of your data model for downstream consumers, AI agents, and embedded analytics.
Views sit on top of the data graph of [cubes][ref-cubes] and create a facade
of your whole data model with which data consumers can interact. They bring
together relevant measures, dimensions, and join paths into a logical
structure that matches how business users think about their data.
See the [view reference][ref-view-reference] for the full list of
parameters and configuration options.
## Why views matter
Views are the primary interface between your data model and your users.
While cubes model the raw relationships and logic in your warehouse, views
reshape that model into business-friendly datasets for easier exploration.
Views shield end-users from complex database schemas, table
relationships, and raw SQL. Business users can pick fields from
a curated dataset in [Explore][ref-explore] or
[Workbooks][ref-workbooks] without needing to understand the joins
or cube structure underneath.
For example, an analyst could pick `product`, `total_amount`, and
`users_city` from an `orders` view without thinking about the underlying
join path from `base_orders` through `line_items` to `products`.
[AI agents][ref-ai-context] query your data model through views.
By curating which members are included and providing descriptive
metadata via `description` and `meta.ai_context`, you control the
context AI uses to generate accurate queries. Well-designed views
with clear naming and descriptions lead to significantly better
AI results.
Views give you fine-grained control over what users can see.
Each view can be scoped with [access policies][ref-access-policies]
to enforce row-level and member-level security. You can also set
`public: false` to hide internal views or use
[COMPILE\_CONTEXT][ref-compile-context] for dynamic visibility
based on the security context.
In complex data models, the same pair of cubes might be reachable
through multiple join paths. Views eliminate this ambiguity by
specifying the exact `join_path` for each included cube, ensuring
queries always follow the intended path.
Views are a natural fit for [embedded analytics][ref-embedding].
Different customer tiers can get access to different views,
allowing you to tailor the analytics experience to your
monetization strategy without duplicating cubes.
## How views work
Views do **not** define their own members. Instead, they reference cubes by
specific join paths and selectively include measures, dimensions, hierarchies,
and segments from those cubes.
```yaml title="YAML" theme={"dark"}
views:
- name: orders
cubes:
- join_path: base_orders
includes:
- status
- created_date
- total_amount
- count
- average_order_value
- join_path: base_orders.line_items.products
includes:
- name: name
alias: product
- join_path: base_orders.users
prefix: true
includes: "*"
excludes:
- company
```
```javascript title="JavaScript" theme={"dark"}
view(`orders`, {
cubes: [
{
join_path: base_orders,
includes: [
`status`,
`created_date`,
`total_amount`,
`count`,
`average_order_value`
]
},
{
join_path: base_orders.line_items.products,
includes: [
{
name: `name`,
alias: `product`
}
]
},
{
join_path: base_orders.users,
prefix: true,
includes: `*`,
excludes: [`company`]
}
]
})
```
In this example, the `orders` view pulls in members from three cubes
along their join paths. End-users see a flat list of fields — `status`,
`created_date`, `product`, `users_city`, etc. — without being exposed to
the underlying cube structure.
## Designing effective views
### Build for your audience
Design views around how your business users think about data, not around
how your database is structured. Group related fields into views that align
with departments or use cases — for example, `sales_overview`,
`customer_360`, or `product_analytics`.
A single cube can be included in multiple views. For example, a `users`
cube might appear in both a `customer_360` view and a `sales_overview`
view, with different fields exposed in each.
### Favor focused views
Smaller, focused views are easier to navigate and lead to better AI
results. Rather than one massive view with hundreds of fields, create
several purpose-built views:
* Views are easier for business users to understand when they're
scoped to a specific domain
* AI agents perform better with focused context
* Simpler views translate to simpler SQL queries with fewer joins
### Curate with metadata
Help your users understand what a view is for and how to use it:
* Set a clear [`description`][ref-view-description] to explain the
view's purpose
* Use [`title`][ref-view-title] for user-friendly display names
* Add [`meta.ai_context`][ref-ai-context] to guide AI agents
* Organize fields into [`folders`][ref-view-folders] for logical
grouping
```yaml title="YAML" theme={"dark"}
views:
- name: sales_overview
description: >
Revenue and order metrics for the sales team.
Includes order status, product details, and customer segments.
meta:
ai_context: >
Use this view for questions about sales performance,
revenue trends, and order analysis. The total_revenue
measure includes only completed orders.
cubes:
- join_path: orders
includes:
- status
- total_revenue
- count
- created_date
- join_path: orders.customers
prefix: true
includes:
- segment
- region
folders:
- name: Order Metrics
includes:
- total_revenue
- count
- status
- name: Customer Info
includes:
- customers_segment
- customers_region
```
```javascript title="JavaScript" theme={"dark"}
view(`sales_overview`, {
description: `Revenue and order metrics for the sales team.
Includes order status, product details, and customer segments.`,
meta: {
ai_context: `Use this view for questions about sales performance,
revenue trends, and order analysis. The total_revenue
measure includes only completed orders.`
},
cubes: [
{
join_path: orders,
includes: [
`status`,
`total_revenue`,
`count`,
`created_date`
]
},
{
join_path: orders.customers,
prefix: true,
includes: [
`segment`,
`region`
]
}
],
folders: [
{
name: `Order Metrics`,
includes: [
`total_revenue`,
`count`,
`status`
]
},
{
name: `Customer Info`,
includes: [
`customers_segment`,
`customers_region`
]
}
]
})
```
### Keep shared logic in cubes
Views are a curation layer. All business logic — SQL definitions, measure
calculations, join relationships — should live in cubes. Views should only
control which members are exposed, how they're named, and how they're
organized. This keeps your model [DRY][wiki-dry] and makes maintenance
straightforward.
### Define a metric on a view when it spans cubes
The exception is a metric whose parts live in different cubes. A view can define
its own [measures][ref-view-measures] and [dimensions][ref-view-dimensions] as
long as their `sql` only combines members the view already includes — a member
that reads a column instead is rejected at compile time:
```yaml title="YAML" theme={"dark"}
views:
- name: orders_overview
# cubes: … includes orders.total_amount and line_items.count
measures:
- name: average_line_value
type: number
multi_stage: true
sql: "{CUBE.total_amount} / NULLIF({CUBE.count}, 0)"
```
```javascript title="JavaScript" theme={"dark"}
view(`orders_overview`, {
// cubes: … includes orders.total_amount and line_items.count
measures: {
average_line_value: {
type: `number`,
multi_stage: true,
sql: `${CUBE.total_amount} / NULLIF(${CUBE.count}, 0)`
}
}
})
```
[`multi_stage`][ref-multi-stage] matters whenever the parts come from different
cubes: it aggregates each of them before combining, instead of evaluating the
expression inside one joined scan where a `one_to_many` join between them would
inflate the numerator. If the cubes don't join to each other at all, see
[multi-fact views][ref-multi-fact-views]. The full example, with the `cubes`
block, is on the [view reference][ref-view-measures].
A cube can also own such a metric, by referencing the other cube's member
directly — that keeps it defined once for every view that includes it, at the
cost of one cube naming another. The [average order value
recipe][ref-recipe-aov] compares the two placements.
### Control visibility
Not every view should be publicly accessible. Use [`public`][ref-view-public]
to hide views that are meant for internal use or are still in development:
```yaml title="YAML" theme={"dark"}
views:
- name: internal_diagnostics
public: false
cubes:
- join_path: system_metrics
includes: "*"
```
```javascript title="JavaScript" theme={"dark"}
view(`internal_diagnostics`, {
public: false,
cubes: [
{
join_path: system_metrics,
includes: `*`
}
]
})
```
For dynamic visibility based on user roles, use `COMPILE_CONTEXT`:
```yaml title="YAML" theme={"dark"}
views:
- name: arr
description: Annual Recurring Revenue
public: COMPILE_CONTEXT.security_context.is_finance
cubes:
- join_path: revenue
includes:
- arr
- date
```
```javascript title="JavaScript" theme={"dark"}
view(`arr`, {
description: `Annual Recurring Revenue`,
public: COMPILE_CONTEXT.security_context.is_finance,
cubes: [
{
join_path: revenue,
includes: [`arr`, `date`]
}
]
})
```
## Organizing members with folders
When a view includes many fields, [folders][ref-view-folders] help organize
them into logical groups. Cube supports both flat and nested folder
structures:
```yaml title="YAML" theme={"dark"}
views:
- name: customers
cubes:
- join_path: users
includes: "*"
- join_path: users.orders
prefix: true
includes:
- status
- price
- count
folders:
- name: Personal Details
includes:
- name
- gender
- created_at
- name: Order Analytics
includes:
- orders_status
- orders_price
- orders_count
```
```javascript title="JavaScript" theme={"dark"}
view(`customers`, {
cubes: [
{
join_path: `users`,
includes: `*`
},
{
join_path: `users.orders`,
prefix: true,
includes: [`status`, `price`, `count`]
}
],
folders: [
{
name: `Personal Details`,
includes: [`name`, `gender`, `created_at`]
},
{
name: `Order Analytics`,
includes: [
`orders_status`,
`orders_price`,
`orders_count`
]
}
]
})
```
Folders are displayed in supported [visualization tools][ref-viz-tools].
Check [APIs & Integrations][ref-apis-support] for details on folder
support. For tools that don't support nested folders, the structure is
automatically flattened.
## Grouping views with view groups
When a data model contains many views, [view groups][ref-view-groups] help
organize them into named collections by domain or purpose — for example,
`sales`, `finance`, or `people`. They're exposed through the
[`/v1/meta`][ref-meta-endpoint] API so downstream tools, AI agents, and
embedded analytics can present a navigable catalog.
See [View groups][ref-view-groups] for the full guide and the
[view group reference][ref-view-group-ref] for the complete list of
parameters.
## Next steps
* See the [view reference][ref-view-reference] for the full list of
parameters
* Learn about [view groups][ref-view-groups] to organize views into
named collections
* Learn about [access policies][ref-access-policies] to govern view access
* Explore [AI context][ref-ai-context] to improve AI query accuracy
* Use the [Semantic Model IDE][ref-ide] to develop views interactively
[ref-cubes]: /docs/data-modeling/cubes
[ref-view-reference]: /reference/data-modeling/view
[ref-view-description]: /reference/data-modeling/view#description
[ref-view-title]: /reference/data-modeling/view#title
[ref-view-public]: /reference/data-modeling/view#public
[ref-view-measures]: /reference/data-modeling/view#measures
[ref-view-dimensions]: /reference/data-modeling/view#dimensions
[ref-multi-stage]: /reference/data-modeling/measures#multi_stage
[ref-multi-fact-views]: /docs/data-modeling/multi-fact-views
[ref-recipe-aov]: /recipes/data-modeling/average-order-value#where-to-put-the-measure
[ref-view-folders]: /reference/data-modeling/view#folders
[ref-access-policies]: /reference/data-modeling/data-access-policies
[ref-ai-context]: /docs/data-modeling/ai-context
[ref-compile-context]: /docs/data-modeling/access-control/context
[ref-explore]: /docs/explore-analyze/explore
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-embedding]: /embedding
[ref-ide]: /docs/data-modeling/data-model-ide
[ref-viz-tools]: /admin/connect-to-data/visualization-tools
[ref-apis-support]: /reference#data-modeling
[ref-view-groups]: /docs/data-modeling/view-groups
[ref-view-group-ref]: /reference/data-modeling/view-group
[ref-meta-endpoint]: /reference/core-data-apis/rest-api/reference
[wiki-dry]: https://en.wikipedia.org/wiki/Don%27t_repeat_yourself
# Visual Modeler
Source: https://docs.cube.dev/docs/data-modeling/visual-modeler
Walks through Cube Cloud’s browser-based Visual Modeler for no-code modeling, including setup, workflow, and known limitations versus the code editor.
Visual Modeler editor provides the visual, no-code experience for building and enhancing
the [data model][ref-data-modeling] of your semantic layer from within your web browser.
Unlike the code-first [data model][ref-data-model] editor, Visual Modeler allows non-technical
users to participate in data modeling without writing code. However, it does not support
some advanced data modeling features listed as known [limitations](#limitations).
Available on [Premium and above plans](https://cube.dev/pricing).
Visual Modeler extends and replaces the [Data Graph](https://cube.dev/blog/introducing-data-graph)
feature.
## Prerequisites
Visual Modeler integrates with the [development mode][ref-dev-mode],
[environments][ref-environments], and [continuous deployment][ref-continuous-deployment]
in Cube Cloud.
Once you start editing the data model in Visual Modeler editor, you will automatically
enter the development mode, so you can safely make changes and test them in
[Playground][ref-playground] without affecting your production deployment.
When you are ready to deploy your changes, it is recommended to create a pull request,
which should be reviewed before merging to the production environment.
## Working with cubes
You can see [cubes][ref-cubes] on the **Cubes** tab of Visual Modeler.
Cubes are represented as rectangles on the canvas with [joins][ref-joins] between them as
an [entity-relationship diagram][wiki-erd] (ERD). You can pan and zoom the canvas and
search cubes using the input box.
### Adding cubes
To add a cube from scratch, click the **+ Add Cube** button.
You can also generate the data model for a new cube based on the database schema.
To do so, click the **+ Generate Cube** button. It will bring up the
modal window where you can select a table from your data source:
### Editing cubes
To edit a cube, click it on the canvas. Then, in the sidebar, click the **Edit**
button. This will bring up the modal window where you can edit the cube's details:
Yu can also click **+ Add → Dimension**, **+ Add → Measure**, or
**+ Add → Access Policy** to quickly add respective members to the cube.
### Working with joins
You can add or edit joins via the **Relationships** tab while you're
[editing a cube](#editing-cubes).
You can also add joins visually by dragging a line on the canvas from the side of one
dimension to another:
## Working with views
You can see views on the **Views** tab of Visual Modeler.
Views are represented as rectangles on the canvas. Cubes that are involved in a view
and provide members for it, are shown to the left of the view.
### Adding views
To add a view, click the **+ Add View** button. It will bring up a dialog where you
can choose the base cube, set the view details such as name and title, and add join paths.
### Editing views
To edit a views, click it on the canvas. Then, in the sidebar, click the **Edit**
button.
## YAML mode
All modal windows mentioned above provide the option to use the *YAML mode* which can be
entered by clicking the **Edit with YAML** button.
In this mode, you can see and edit the data model represented using the [YAML syntax][ref-yaml-syntax].
That allows you to work around known [limitations](#limitations) and use all data modeling
features while staying in Visual Modeler editor.
## Limitations
Visual Modeler does not intend to support all available data modeling features.
You always have options to use the [YAML mode](#yaml-mode) in Visual Modeler or switch the
code-first [data model editor][ref-data-model].
Currently, Visual Modeler does not support the following features in its UI:
* Editing of [meta information][ref-meta] for cubes, views, and their members.
* Editing [segments][ref-segments], [hierarchies][ref-hierarchies], and [refresh keys][ref-refresh-keys] in cubes.
* Editing [folders][ref-folders] in views.
* Adding [subquery dimensions][ref-sub-query] and time dimensions with [custom granularities][ref-granularities] in cubes.
* Adding [drill members][ref-drill-members] and [rolling window][ref-rolling-window] measures in cubes.
* Adding [pre-aggregations][ref-pre-aggregations].
* Removing cubes and views.
Cubes and views using these data modeling features will still be displayed in
Visual Modeler. However, those features can only be edited via the [YAML mode](#yaml-mode)
or in the code-forst data model editor.
Additionally, Visual Modeler only allows editing of cubes and views that are
defined using the [YAML syntax][ref-yaml-syntax] (not JavaScript). Also, it does not
allow editing of [dynamic data models][ref-dynamic-data-models] or models which use
[Jinja][ref-jinja].
[ref-data-modeling]: /docs/data-modeling/overview
[ref-data-model]: /docs/data-modeling/data-model-ide
[ref-meta]: /reference/data-modeling/cube#meta
[ref-segments]: /reference/data-modeling/segments
[ref-hierarchies]: /reference/data-modeling/hierarchies
[ref-folders]: /reference/data-modeling/view#folders
[ref-refresh-keys]: /reference/data-modeling/cube#refresh_key
[ref-pre-aggregations]: /reference/data-modeling/pre-aggregations
[ref-sub-query]: /reference/data-modeling/dimensions#sub_query
[ref-granularities]: /reference/data-modeling/dimensions#granularities
[ref-drill-members]: /reference/data-modeling/measures#drill_members
[ref-rolling-window]: /reference/data-modeling/measures#rolling_window
[ref-dynamic-data-models]: /docs/data-modeling/dynamic
[ref-jinja]: /docs/data-modeling/dynamic/jinja
[ref-continuous-deployment]: /admin/deployment/continuous-deployment
[ref-yaml-syntax]: /docs/data-modeling/concepts/syntax#model-syntax
[ref-dev-mode]: /docs/data-modeling/dev-mode
[ref-environments]: /admin/deployment/environments
[ref-playground]: /docs/explore-analyze/playground
[ref-cubes]: /reference/data-modeling/cube
[ref-joins]: /reference/data-modeling/joins
[wiki-erd]: https://en.wikipedia.org/wiki/Entity–relationship_model
# Analytics Chat
Source: https://docs.cube.dev/docs/explore-analyze/analytics-chat
Conversational analytics interface for asking plain-language questions and getting trusted, AI-powered insights from your semantic layer.
Analytics Chat is Cube's conversational analytics experience — ask questions in plain language and get trusted, AI-powered insights without writing queries or building visualizations.
## How it works
The AI agent interprets your questions, generates queries against your semantic model, and presents findings in natural language. The agent can create multiple queries to answer complex questions and perform ad hoc analysis while maintaining full data governance through your semantic layer.
## Key features
* **Natural language queries** – Ask questions in plain language without knowing SQL
* **Multi-query reasoning** – The AI creates and synthesizes multiple queries for complex questions
* **Python analysis** – For work SQL can't express — forecasting, cohort analysis, statistical tests — the agent can run [Python](/docs/explore-analyze/workbooks/python-analysis) and render the result inline, then save it as a re-runnable workbook report or exploration (in preview)
* **Semantic model integration** – All queries run against your semantic model with proper access control and security, honoring the active [security context](/docs/explore-analyze/workbooks/querying-data#applying-a-security-context)—including an override applied by a developer or admin
* **Queued messages** – Send follow-up messages while the agent is still processing
* **Save your results** – Ask the agent to save a result as a [report](/docs/explore-analyze/workbooks) inside a workbook, or as a standalone [exploration](/docs/explore-analyze/explore#saving-explorations) when you don't want to create a workbook
## Discover available fields
You can ask the AI agent what's available in your semantic model before
diving into analysis. This is useful when you're new to a deployment or
exploring an unfamiliar view.
Try prompts like:
* "What fields are available?"
* "What measures and dimensions can I query in the Orders view?"
* "Describe the fields in the Customers view and what each one means."
* "What does the `lifetime_value` measure represent?"
The agent uses the descriptions and [AI context](/docs/data-modeling/ai-context)
defined on your views, measures, and dimensions to answer. Well-documented
semantic models produce better answers — see
[AI context best practices](/docs/data-modeling/ai-context#best-practices)
for guidance on writing descriptions the agent can use.
## Queued messages
You can send follow-up messages while the AI agent is still processing a
previous request. These messages are queued and processed in order once the
agent completes its current task.
This allows you to:
* **Refine your question** before the agent finishes, if you realize you want
to adjust the scope or add constraints
* **Queue multiple questions** to ask a series of related questions without
waiting for each response
* **Provide additional context** to give the agent more information that it
can incorporate into its next response
## Sharing
Click **Share** in the chat header to grant view access to other members
of your account. You can pick individual users, a user group, or flip
**General access** to **Organization** to make the chat visible to
everyone. Use the **Copy link** button next to **Share** to grab a direct
URL once access is set up. Recipients still need to have access to the
deployment and the data the agent used; otherwise they'll see a "Chat
not available" page.
Recipients open the chat at the same URL as the owner and can scroll
through the full conversation, expand the agent's reasoning and tool
calls, and explore the charts and tables it produced — but they can't
send new messages. Only the owner can continue the thread, and the
header shows a "Shared by …" label so viewers always know whose
conversation they're reading.
If the owner asks more questions later, those messages — and the
agent's replies — show up for viewers on the next load.
## Embedding
Analytics Chat can be [embedded](/embedding) into your applications for customer-facing analytics, internal tools, or white-label solutions.
## Learn more
* [Workbooks](/docs/explore-analyze/workbooks) – Build and organize reports with AI assistance
* [Python analysis](/docs/explore-analyze/workbooks/python-analysis) – Save a Python analysis from chat as a re-runnable workbook report or exploration
* [Embedding](/embedding) – Embed analytics in your applications
* [Cube Agentic Analytics announcement](https://cube.dev/blog/cube-agentic-analytics) – Learn about the vision behind agentic analytics
# Apps
Source: https://docs.cube.dev/docs/explore-analyze/apps
Build custom, interactive data experiences backed by governed data from your Cube semantic layer.
Apps are currently in preview, and their capabilities and authoring experience may
still change. Reach out to the [Cube support team](/admin/account-billing/support)
to activate this feature for your account.
Apps let you go beyond a dashboard's grid of widgets. An app is a custom,
interactive interface with its own layout, components, and controls, powered by
live data from reports in a Cube workbook.
Use an app when you need a guided workflow, a custom visual presentation, or an
interactive data tool that does not fit a standard dashboard layout.
## Create an app with the Cube agent
Open a workbook, switch to **Dashboard Builder**, and describe the app you want in
the agent chat. Include the business question, the data you want to show, and how
people should interact with it. For example:
> Build a revenue performance app with headline metrics, a monthly trend chart,
> and a searchable status filter. Use the company's brand colors.
The agent creates the workbook reports, writes the app, and installs reusable
components for common elements such as charts, tables, KPI cards, tabs, and
searchable controls. Continue prompting the agent to adjust the layout, behavior,
or styling.
## Preview and edit the app
Use **Preview** to interact with the running app. Select **Code** to inspect or
edit its files directly. After making code changes, click **Run** to apply them
and refresh the preview.
App components are copied into the app's source, so you can customize their code
and styling for that specific experience.
## Governed data access
An app reads the results of reports in its workbook. Cube runs those reports using
the viewer's existing permissions and passes the results to the app. The app does
not receive credentials for your data source or Cube APIs, and it cannot query
your warehouse directly.
Apps run inside a sandboxed frame, isolated from the Cube interface. The app's
data access remains governed by the semantic model and the permissions already
configured in Cube.
## Preview limitations
During preview:
* Apps are created and viewed from the workbook's Dashboard Builder.
* Publishing, embedding, scheduled refreshes, and deliveries are not yet
supported for Apps.
* Apps can read workbook report results but cannot write data back to a source.
* Available components and APIs may change as the feature develops.
# Area
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/area
Show volume or part-to-whole trends over time with a filled region beneath the line.
Area charts add a filled region beneath the line. The fill communicates magnitude and cumulative volume, making area charts especially effective for stacked part-to-whole breakdowns over time.
## Variants
### Stacked
Each series is stacked on top of the previous one. The top edge shows the cumulative total; each filled band shows the individual contribution. Use this for part-to-whole breakdowns over time (e.g. revenue by product category over months).
### Percentage stacked (Stack %)
Each series is normalized to 100% at every X-axis value, showing proportional contribution over time. Use when relative share matters more than absolute volume.
### Overlaid
All series share the same baseline. Use only when series are clearly separated in value and you want to compare trajectories rather than totals — with many series, fills will occlude each other.
## Stacking
Set stacking behavior in the **Color & stacking** section of the Style tab. The **Stack**, **Stack %**, and **Overlay** options are the relevant modes for area charts.
See [Color & stacking](/docs/explore-analyze/charts/configuration/color-and-stacking) for palette options and stacked segment sorting.
## Axis behavior
Area charts use a temporal X axis for time dimensions — continuous and time-aware. For non-time fields, the axis is ordinal.
# Bar
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/bar
Compare values across categories with grouped, stacked, percentage, and horizontal variants.
Bar charts compare values across discrete categories. Use them when the magnitude of individual values matters and direct comparison between categories is the goal.
## Variants
### Basic
One dimension on the X axis, one measure on the Y axis. Each category gets a single bar.
### Grouped
Each series renders as a separate cluster of bars side-by-side at every X-axis value. Best for comparing absolute values across multiple series at each category.
### Stacked
Series are stacked on top of each other at each X-axis value. Use when you want to show both individual contributions and the total at a glance.
### Percentage stacked (Stack %)
Each stack is normalized to 100%, showing each series as a proportion of the total per X-axis value. Use when relative distribution matters more than absolute values.
### Horizontal
All three vertical variants are also available horizontally. Horizontal bars work well for long category labels, many categories, or when a left-to-right reading direction feels more natural.
### Composite (bar + line)
Assign one series to the right Y axis and set its mark type to **Line** in [series configuration](/docs/explore-analyze/charts/configuration/series-configuration). This creates a dual-axis chart — useful for overlaying a rate on top of volume data (e.g. order count as bars, revenue per order as a line).
## Stacking
Stacking behavior is set in the **Color & stacking** section of the Style tab:
| Option | Behavior |
| ------------- | -------------------------------------------------------------- |
| **Automatic** | Cube picks the best option based on your data |
| **Stack** | Series at the same X value are stacked |
| **Group** | Series at the same X value are placed side-by-side |
| **Overlay** | Series are drawn on top of each other (rarely useful for bars) |
| **Stack %** | Series are stacked and normalized to 100% |
Stacking can also be set independently per Y-axis series — enabling grouped clusters of stacked sub-groups.
## Data labels
Bar charts support totals labels on stacked charts — showing the aggregate value above each full stack. Enable **Data Labels** in the Fields tab, then configure:
* **Format** — number format applied to the label value
* **Font size** — size of the label text
* **Position** — Outside end, Inside end, Inside center, or Inside base
## Axis behavior
Bar charts use an ordinal X axis — each bar is labeled independently, the axis is not time-aware, and categories not present in the result set are not plotted. With time fields, bars are ordered ascending automatically. For other data types, order follows the results table sort.
Tooltips on stacked bars are scoped to the segment under the cursor, not the full stack total.
# Boxplot
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/boxplot
Display the statistical distribution of a measure across categories using box-and-whisker plots.
Boxplots summarize the distribution of a numeric measure per category, showing the median, interquartile range, and outliers in a single mark. Use them to compare spread and skew across groups rather than just averages.
## Variant
### Box-and-whisker
One dimension on the X axis groups the data into categories. One measure on the Y axis provides the raw values. The chart computes the distribution statistics internally.
## Reading a boxplot
| Element | Description |
| --------------- | --------------------------------------------- |
| **Box** | Interquartile range — 25th to 75th percentile |
| **Center line** | Median (50th percentile) |
| **Whiskers** | Extend to min/max within 1.5× the IQR |
| **Points** | Values beyond the whiskers (outliers) |
## Data structure
Boxplots require **raw, row-level data** — each row is one observation. Do not pre-aggregate before using a boxplot; the chart engine computes the distribution itself.
* **X axis** — the grouping dimension (e.g. product category)
* **Y axis** — the measure to distribute (e.g. sale price)
# Funnel
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/funnel
Show how a value narrows from stage to stage, such as signup → activation → purchase.
Funnel charts draw one bar per stage, centered on a shared midline, with a connector between each adjacent pair of stages. Each bar's width is proportional to its stage's value, so the figure narrows wherever values drop. Best for stage-to-stage conversion: how many users, orders, or events survive each step of a process.
## Fields
A funnel needs exactly two fields, picked on the Fields tab:
* **Stage** — a dimension (a time dimension works too, bucketed by its grain). Each distinct value becomes one stage.
* **Value** — a measure. It is rolled up per stage, so a result with extra columns still draws one bar per stage.
Stages keep the order your query returned them in. To reorder stages, adjust the sort in the results table — a stage that grows partway down is real data (re-entry, late-arriving events), so the chart never sorts it away for you.
A stage whose value is zero still draws, as a thin line in its stage color, so a gap in the data reads as a zero rather than as a missing stage.
## Direction
The **Direction** row at the top of the Style tab decides which way the stages progress:
* **Vertical** — stages run top to bottom, each bar spanning the width. The default.
* **Horizontal** — stages run left to right, each bar spanning the height. Useful when stage names are long, or alongside other left-to-right charts.
The taper, the connectors and every label carry over unchanged; only the axes swap.
Direction is greyed out until the chart has both a stage field and a value field.
## Conversion percentages
The connectors between stages carry the funnel's percentages. Two independent toggles in the **Data labels** section of the Style tab control them:
* **First** — each stage as a share of the first stage, drawn as `36% of first`. On by default.
* **Prev** — each stage as a share of the stage before it, drawn as `20% of previous`.
Turning both on stacks the two lines on each connector; turning both off leaves the connectors unlabeled. Whichever you select also appears in the tooltip, alongside the stage and value.
## Data labels
Each bar carries its stage name and value, both on by default and toggled in the **Data labels** section of the Style tab. The value's number format has its own picker, and one font size governs every label the funnel draws — the bar labels and the connector percentages alike.
Labels always sit at the bar's center. When a bar is too narrow to hold its label, the text switches from white to a dark color so it stays readable against the page background.
## Color
Each stage takes its own color from the palette selected in the **Palette** dropdown on the Style tab, in palette order. See [Palettes](/docs/explore-analyze/charts/configuration/color-and-stacking#palettes) for the built-in palettes and how to define a custom one.
## Legend
A funnel draws no legend by default — every bar already carries its stage name, so a legend would repeat it. Turn one on in the [**Legend**](/docs/explore-analyze/charts/configuration/color-and-stacking#legend) section of the Style tab, which also sets its placement. It names each stage in its palette color.
# Heatmap
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/heatmap
Visualize a measure across two categorical dimensions as a color-coded grid.
Heatmaps render data as a grid where each cell's color encodes a measure value at the intersection of two dimensions. Use them to spot patterns, dense clusters, and outliers across two categorical axes at once.
## Variant
### Color-scaled grid
One dimension on the X axis, one on the Y axis, and one measure mapped to the **Color** channel. Each cell's color encodes that measure's value for its row/column pair.
## Color palette
A continuous gradient palette is applied by default — low values map to the lighter end, high values to the darker end. Change or reverse the palette in the **Color** section of the Style tab.
For custom palettes, see [Color & stacking](/docs/explore-analyze/charts/configuration/color-and-stacking).
## Data structure
A heatmap requires exactly:
* One dimension on the **X axis**
* One dimension on the **Y axis**
* One measure on the **Color** channel
Cells with no data in the result set are left blank.
# HTML
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/html
Build fully custom layouts using HTML templates with Handlebars and JavaScript.
HTML charts let you render any layout that Vega-Lite cannot produce — custom scorecards, branded cards, rich tables, or any structure built with HTML, CSS, and JavaScript. Query results are available inside the template via [Handlebars](https://handlebarsjs.com/) expressions.
## Variant
### HTML template
A free-form HTML document with Handlebars expressions. Write any markup and bind query result data directly in the template.
## Template structure
Query results are available in the template context under `result`.
**Access the first row** with `result._first`:
```html theme={"dark"}
{{result._first.order_items.status}}
{{result._first.order_items.count}} orders
```
**Iterate over all rows** with `{{#each result.data}}`:
```html theme={"dark"}
```
## Using JavaScript
Standard `
```
# Chart types
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/index
Overview of all built-in chart types available in Cube workbooks.
Cube includes a library of built-in chart types covering the most common visualization patterns. Each type is designed for specific data shapes — choose the one that best fits your query structure and the question you're answering.
* [Bar](/docs/explore-analyze/charts/chart-types/bar)
* [Line](/docs/explore-analyze/charts/chart-types/line)
* [Area](/docs/explore-analyze/charts/chart-types/area)
* [Scatter](/docs/explore-analyze/charts/chart-types/scatter)
* [Pie & donut](/docs/explore-analyze/charts/chart-types/pie)
* [Funnel](/docs/explore-analyze/charts/chart-types/funnel)
* [Sankey](/docs/explore-analyze/charts/chart-types/sankey)
* [Heatmap](/docs/explore-analyze/charts/chart-types/heatmap)
* [Boxplot](/docs/explore-analyze/charts/chart-types/boxplot)
* [Table](/docs/explore-analyze/charts/chart-types/table)
* [KPI](/docs/explore-analyze/charts/chart-types/kpi)
* [Map](/docs/explore-analyze/charts/chart-types/map)
* [HTML](/docs/explore-analyze/charts/chart-types/html)
For configuration options that apply across chart types — axes, color, series settings, tooltips — see [Configure charts](/docs/explore-analyze/charts/configuration).
## Recommended chart type
When your query returns a result, Cube outlines the one chart type that best fits it and labels it
**Recommended**. It is a suggestion: every other available type stays selectable, and the outline
yields as soon as you pick a type in that picker.
Cube never applies a chart type for you. The recommendation is always only outlined, wherever the
picker appears, and choosing a type is always your action.
### What Cube looks at
The recommendation reads the shape of your query and its result: how many measures, dimensions and
time dimensions you selected, how many distinct values each dimension has, and how long the
category labels are. It is computed in your browser and is not saved — the suggestion lasts for the
current session only.
Where a threshold is involved, it comes from
[Draco](https://idl.cs.washington.edu/papers/draco/) (Moritz et al., *Formalizing Visualization
Design Knowledge as Constraints*, IEEE VIS 2018), which encodes established visualization research
as rules: a category axis becomes crowded past 12 values, a color or stacking channel saturates
past 10 series, and no axis reads past 30 values.
One number is Cube's own rather than Draco's: a mean label length past **16 characters** counts as
long.
No threshold is applied to the number of rows. A few rules read the row count directly — a KPI
needs exactly one row, a heatmap needs its grid at least half filled — but none of them treats "too
many rows" as a reason to recommend a different chart, because the row count follows the
granularity you chose rather than the shape of your query.
### The conditions
Cube first checks whether the query disqualifies every chart, and recommends the
[table](/docs/explore-analyze/charts/chart-types/table) if it does:
| Recommendation | When |
| -------------- | ----------------------------------------------------------------------------------------------------- |
| **Table** | The query has no measure, two time dimensions, three or more dimensions, or a category past 30 values |
A query meeting any of these has no chart encoding that reads well, so the table is the honest
answer. [Maps](/docs/explore-analyze/charts/chart-types/map) are the exception and are still
recommended over the table — a map has no category axis to crowd, so a large number of plotted
places is normal rather than unreadable.
Otherwise Cube checks these in order and takes the first that matches:
| Recommendation | When |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Map** | The query contains latitude and longitude |
| **KPI** | One measure, no dimensions, and a single row |
| **Line** | A measure over a time dimension, optionally split by one category of up to 10 values |
| **Scatter** | Two or more measures with no time dimension, across one category or several rows — the shape for correlating them |
| **Bar (stacked)** | One measure split by two categories, the first up to 12 values and the second up to 10 |
| **Heatmap** | One measure split by two categories, both past those bounds — the first past 12 values and the second past 10 — whose combinations fill at least half the grid |
| **Bar** | One measure across up to 12 categories whose labels average 16 characters or fewer |
| **Bar (horizontal)** | The same, where the labels are longer or the categories run past 12 |
The two [bar](/docs/explore-analyze/charts/chart-types/bar) rows are variants of the same chart
type, not separate ones.
Pie, funnel, sankey, area, boxplot and HTML are never recommended — pick them yourself. Pie's shape,
one measure across a few categories, is the one Bar already answers, and a bar chart compares those
values more accurately. A funnel additionally asserts that its stages are sequential steps of one
process, which the shape of a query does not reveal. All six stay available in the picker like any
other chart type.
### When nothing is recommended
Cube suppresses the recommendation rather than guessing. You will see no outline when:
* no rule matches the query cleanly;
* a measure over time is split into 11 to 30 series — too many for a color channel, too few to give
up on;
* a value it needs is still unknown, such as the number of distinct values in a dimension;
* the result came back empty, so there is nothing to draw;
* the result was truncated by a row limit, so the counts it would read are incomplete;
* you have picked a type by hand in the picker you are looking at. The outline returns the next
time you open it — including on the type you applied, since Cube marks the best fit whether or
not it is already on screen.
# KPI
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/kpi
Highlight key metrics with composable blocks — numbers, comparisons, progress bars, sparklines, text, and HTML.
KPI tiles highlight important metrics and key performance indicators. Rather than a single fixed layout, KPIs are **row-based and composable** — you build the tile by stacking blocks of different types, each configured independently.
## Getting started
The KPI visualization starts with a **Number** block that compares the first and second rows of your result for the first numeric column.
To add a new block, click the **+** button inside the visualization and choose the block type. To edit a block, click it in the visualization — the configuration panel updates to show that block's settings.
Reorder blocks by dragging the handle on the left side of each block. Delete a block with the trash icon in its configuration panel.
## Block types
### Number
Displays a single value from your query result.
| Setting | Description |
| --------------------- | ------------------------------------------- |
| **Field** | The measure or dimension to display |
| **Row** | Which row of the result to read from |
| **Format** | Number formatting (currency, percent, etc.) |
| **Font size / color** | Text appearance |
| **Alignment** | Left, center, or right |
### Comparison
Displays a current value alongside a previous value with a calculated difference. Use this to show period-over-period change or target vs. actual.
| Setting | Description |
| ----------------------------------- | ------------------------------------------------------------------------------------------------ |
| **Current field / row** | The primary value |
| **Previous field / row** | The value to compare against — can be a different column or a different row from the same column |
| **Difference format** | Absolute, percentage, or both |
| **Positive color / Negative color** | Colors applied based on whether the change is positive or negative |
To compare to a static goal, add a column to your query that always returns the same number (e.g. a calculated field with a constant), then select it as the **Previous** field.
### Progress bar
Shows a value relative to a target as a progress bar or circle. Useful for goal tracking.
| Setting | Description |
| ---------------------- | ------------------------------------- |
| **Value field / row** | The current progress value |
| **Target field / row** | The goal or maximum value |
| **Style** | Bar or circle |
| **Color** | Fill color for the progress indicator |
To compare against a static number like a quarterly goal, add a calculated field that always returns that number and use it as the **Target** field.
### Sparkline
Renders a compact trend chart — either a bar or line — within the KPI tile. Use this to show the trend behind the headline number.
| Setting | Description |
| -------------------- | --------------------------------------------------------------- |
| **Series measure** | The measure to plot |
| **Series dimension** | The dimension to use as the X axis (typically a time dimension) |
| **Chart type** | Bar or line |
| **Height / Width** | Dimensions of the sparkline in the tile |
| **Max points** | Limits the number of data points rendered |
| **Colors** | Line/fill colors |
### Text
A Markdown-enabled text field for titles, headings, or descriptions within the tile. Supports standard Markdown formatting.
Use text blocks for short labels and headings inside the KPI. For complex layouts with Markdown, use the [HTML block](#html) or consider the [Custom visualization](/docs/explore-analyze/charts/custom) instead.
### HTML
A free-form HTML block rendered inside the KPI tile. Use this for advanced custom layouts that go beyond what the other block types support.
## Layout controls
KPI tiles have a layout panel that controls how blocks are arranged:
| Setting | Description |
| ------------- | --------------------------------------------------------------- |
| **Direction** | Row (blocks side-by-side) or Column (blocks stacked vertically) |
| **Alignment** | How blocks align on the cross axis |
| **Justify** | How blocks are distributed along the main axis |
| **Gap** | Spacing between blocks |
| **Wrap** | Whether blocks wrap to a new line when the tile is narrow |
## Converting to raw Markdown
If you need customization beyond what the block builder supports, you can convert the KPI visualization directly to Markdown by clicking on the Markdown visualization option.
Converting to Markdown is a one-way operation. If you do this accidentally, use the undo button to revert.
# Line
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/line
Track how a measure changes over time or across an ordered dimension.
Line charts connect data points with a continuous line. Use them when the direction and rate of change matter — trends over time, trajectories of multiple series, or any ordered dimension where continuity is meaningful.
## Variants
### Single series
One measure plotted against a time or ordered dimension. The simplest and most common use case.
### Multi-series
Multiple lines plotted on the same axes. Map a second dimension to the **Color** channel to split one measure into one line per dimension value, or add multiple measures to the Y axis to plot them as separate series.
### With reference line
A horizontal reference line marks a target, threshold, or benchmark. Added via the **Axes** section of the Style tab.
## Reference lines
Add a horizontal reference line in the **Y axis** section of the Style tab. Set the value, label, color, and line style (solid, dashed, or dotted). A labelled line also takes horizontal (start / middle / end) and vertical (above / below) placement for its text. Multiple reference lines are supported.
See [Axes](/docs/explore-analyze/charts/configuration/axes) for the full configuration reference.
## Axis behavior
Line charts use a temporal X axis for time dimensions — continuous, time-aware, and not constrained to the result set. For non-time dimensions, the axis is ordinal and follows the result set sort order. Gaps in the line indicate missing data points.
## Combining with bars
To layer a line on top of a bar chart, add a second series to the Y axis, set its mark type to **Line** in [series configuration](/docs/explore-analyze/charts/configuration/series-configuration), and assign it to the right Y axis for a dual-axis layout.
# Map
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/map
Plot geographic data as point maps (latitude/longitude) or region maps (choropleth from GeoJSON polygons).
Map charts visualize geographic data in two modes:
* **Point map** — plot rows at their latitude/longitude coordinates, optionally sized and colored by additional fields.
* **Region map** — paint polygons (countries, US states, or any custom [GeoJSON](https://en.wikipedia.org/wiki/GeoJSON)) as a choropleth driven by a measure.
## When to use
* **Point map** — locations with known coordinates (stores, shipments, sign-ups), geographic clusters and outliers.
* **Region map** — measures aggregated by country, state, postcode, sales territory, or any user-defined polygon set.
## Chart type
Pick **Point** or **Region** in the chart settings panel's **Type** row. The default is **Point**.
## Projection
| Projection | Description |
| ----------------- | ------------------------------------------------------------------- |
| **Mercator (2D)** | Standard flat map projection — best for most use cases |
| **Globe (3D)** | Spherical globe view — useful for data spanning multiple continents |
Both modes support either projection.
## Point map
A point map requires:
* **Latitude** — numeric field containing decimal latitude.
* **Longitude** — numeric field containing decimal longitude.
Optional:
* **Size** — numeric measure that scales point radius proportionally.
* **Color** — dimension or measure that colors points by category or value.
### Point color
When no **Color** field is assigned, all points render in the configurable **Default color**. When a dimension is assigned, each unique value gets a distinct color from the active palette — pick from the built-in palettes or supply a custom one, see [Color and stacking](/docs/explore-analyze/charts/configuration/color-and-stacking).
### Point size
Assign a numeric measure to the **Size** channel to scale point radius by value. The size range (minimum and maximum radius in pixels) is configurable in the settings panel.
### Clustering
Clustering groups nearby points into a single circle that expands on zoom. It is **off by default** and can be enabled in the settings panel.
## Region map
A region map requires:
* **Source** — the GeoJSON to render: `World countries`, `US states`, or `Custom`.
* **Property** — which property in each GeoJSON feature acts as the join key (e.g. `name`, `ISO3166-1-Alpha-2`, `state_code`).
* **Dimension** — which column from your query joins to **Property**.
* **Measure** — the numeric measure that drives the choropleth fill.
Built-in sources expose a fixed list of allowed properties. **Custom** lets you point at any public GeoJSON URL; the property dropdown then enumerates every key present in the feature collection.
When the chart enters Region mode, the join dimension and measure are auto-picked when an unambiguous match exists in your query.
The choropleth gradient comes from the active palette — pick from the built-in palettes or supply a custom one, see [Color and stacking](/docs/explore-analyze/charts/configuration/color-and-stacking).
### Custom GeoJSON
When **Source** is set to **Custom**, paste a URL serving a GeoJSON `FeatureCollection`. Requirements:
* Served over HTTPS with CORS enabled.
* Polygons or multipolygons (not points or lines).
* Each feature must have a `properties` object containing the join key.
The fetched GeoJSON is cached for the session.
### Unmatched regions
Regions in the GeoJSON with no matching data row are hidden by default. Toggle **Fill unmatched** in the settings panel to render them with the **Default color** at reduced opacity instead.
## Tooltip fields
The **Tooltips** section controls which fields appear when a user hovers over a point or region. See [Tooltips](/docs/explore-analyze/charts/configuration/tooltips).
## Map interaction
The map is interactive — users can pan and zoom. Viewport changes persist while you edit other settings; switching the data source or projection re-fits the camera to the new bounds.
# Pie & donut
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/pie
Show how a single measure is distributed across a small number of categories.
Pie charts divide a circle into slices proportional to each category's share of the total. Best for communicating part-to-whole proportions to a broad audience when you have a small number of clearly distinct categories (typically five or fewer).
## Variants
### Pie
Standard filled circle. Each slice's arc length is proportional to its value.
### Donut
A pie with a hollow center. Select **Donut** under **Shape** in the Style tab to switch; select **Pie** to switch back.
## Center total
A donut shows the measure's grand total in its hole by default, using the measure's number format. Use the **Show total** button in the Data labels section of the Style tab to turn it off or on again.
The button appears only for a donut, since a pie has no hollow center to fill.
The total uses the font size set in the Data labels section, the same as the slice labels. If you turn on the **Percentage** label, the total also shows `100%` on a second line, since the whole is the entire circle.
## Color and slice ordering
Slices take their colors from the palette selected in the **Palette** dropdown on the Style tab, in palette order, matched to the sort order of your query results. See [Palettes](/docs/explore-analyze/charts/configuration/color-and-stacking#palettes) for the built-in palettes and how to define a custom one. To change which slice appears first, adjust the sort in the results table.
A pie draws a legend naming each slice. Hide it or move it in the [**Legend**](/docs/explore-analyze/charts/configuration/color-and-stacking#legend) section of the Style tab.
# Sankey
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/sankey
Show flows between nodes, such as order status to year or channel to conversion.
Sankey diagrams draw one node per distinct endpoint and one ribbon per flow, with each ribbon's thickness proportional to its value. Best for showing how a quantity divides and moves between two sets of things: order statuses across years, traffic channels into outcomes, or one step of a user journey into the next.
## Fields
A sankey needs exactly three fields, picked on the Fields tab:
* **Source** — a dimension (a time dimension works too, labelled by its grain). The node a flow leaves.
* **Target** — a dimension. The node a flow enters.
* **Value** — a measure. Each row's value becomes one ribbon's thickness.
Each result row is read as one flow, so a query grouped by two dimensions with one measure draws directly. **Swap** exchanges the source and target, which reverses every ribbon's direction.
Stages come from shared node names rather than from a stage field: rows `A → B` and `B → C` draw as a three-stage diagram, because `B` appears as both a target and a source. Two dimensions that share no values always draw two stages.
A row with a blank endpoint, or a value of zero or less, is not a flow and is left out — a sankey cannot draw a ribbon with no thickness.
A flow that returns to a node it already passed through is a cycle. The returned-to node is split into a numbered copy so the diagram still lays out left-to-right — `A → B → A` draws as `A → B → A (2)`. Every copy of a node shares its color, and hovering one highlights the ribbons of all of them.
A node that flows directly to itself has no valid split, and the chart names the loop and asks you to remove one hop.
## Node grouping
A wide breakdown can return more flows than a diagram can carry, and the small ones become unreadable first. Node grouping folds them into a single **Other** node per side, so the tail stays represented as an aggregate instead of vanishing. It is off until you pick a mode, in the Node grouping section of the Fields tab:
* **Minimal share** — fold every node below a percentage of its side's total. A node sitting exactly on the threshold is kept.
* **Top N nodes** — keep the N largest nodes per side. Tied nodes at the cutoff are kept together, so asking for the top 3 can keep more than 3 when third place is shared.
The two sides fold independently, each against its own total, because a node that is small as a target may be large as a source. The bucket's name is yours to set; leave it blank for **Other**. When both sides fold at once, the two buckets are named apart — **Other (source)** and **Other (target)** — so one is not read as the other.
A single node is never folded on its own — replacing one named node with a same-sized **Other** would say less than leaving it alone. A node whose own value already matches the bucket's name is also left out of the bucket, so a real category called "Other" is never merged into it.
The bucket stands for a set of members rather than one, so it has no query behind it: clicking a ribbon that touches it does not drill down.
## Truncation
Independently of grouping, a sankey draws at most 50 flows, keeping the largest. This is a backstop against a runaway result, not a legibility control — when it removes anything, a warning appears beside the Flows heading on the Fields tab saying how many of the total it is showing, pointing at node grouping, which is the real answer for a diagram too wide to read.
## Data labels
Each node carries its name and, optionally, its value. Both are independent buttons in the Data labels section of the Style tab: **Name** is on by default, **Value** is off. With both on, the value follows the name in parentheses; with only the value on, the node carries a bare number; with both off the diagram draws no labels.
A node's value is the larger of its inflow and its outflow, and takes the value field's own number format.
Labels sit to the right of their node, except in the final stage, where they flip to the left so they stay inside the frame.
## Color
Each node takes its own color from the palette selected in the **Palette** dropdown on the Style tab. The color follows the node's name rather than its position, so filtering or re-sorting the query does not repaint the diagram. See [Palettes](/docs/explore-analyze/charts/configuration/color-and-stacking#palettes) for the built-in palettes and how to define a custom one.
The **Flows** control sets where a ribbon takes its color from:
* **Source** — the color of the node the ribbon leaves. The default.
* **Neutral** — one muted color for every ribbon, leaving the nodes to carry the palette.
* **Target** — the color of the node the ribbon enters.
A sankey draws no legend, and has no setting for one: every node is labeled in the diagram itself.
## Tooltips
Hovering a ribbon shows the result columns named in **Tooltip fields** on the Fields tab, which starts as every column in the query except the three the diagram already draws — source, target and value. Add them back like any other field if you want them repeated. Hovering a node shows its name and value.
See [Tooltips](/docs/explore-analyze/charts/configuration/tooltips) for adding, removing and reordering fields.
## Drill down
Clicking a ribbon drills into the row behind that flow. Ribbons touching a grouped **Other** node are not clickable, since that node covers several members at once.
# Scatter
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/scatter
Plot two measures against each other to find correlations, clusters, and outliers.
Scatter charts plot individual data points using two measures as X and Y coordinates. Use them to explore relationships between measures, identify clusters, and spot outliers across a population of entities.
## Variants
### Basic scatter
Two measures on the X and Y axes. Each row in your result set becomes one point.
### With color grouping
Map a dimension to the **Color** channel to assign each point a color by category. Use this to compare how different groups distribute across the same two measures.
### With size encoding
Map a third numeric measure to the **Size** channel to scale each point's radius by value. Use this to encode a third variable without adding a new axis.
## Size encoding
Assign a measure to the **Size** channel in the Fields tab. Points scale proportionally to the measure value. Configure the minimum and maximum radius in the Style tab.
## Tooltips
All mapped fields appear in the tooltip by default. Add or remove fields in the **Tooltips** section of the Fields tab.
See [Tooltips](/docs/explore-analyze/charts/configuration/tooltips) for details.
# Table
Source: https://docs.cube.dev/docs/explore-analyze/charts/chart-types/table
Display query results as a configurable table with pivots, conditional formatting, and custom styling.
The table visualization presents query results in a structured grid. Unlike the raw results table in the query panel, the table visualization is designed for sharing — it supports pivots, conditional formatting, custom styling, and a per-cell menu for links, copy, and drill-down.
## Showing and hiding columns
Right-click a column header in the table visualization to hide it. Hidden columns can be restored from the **Fields** section of the configuration panel. You can also hide columns from the column options menu (the three-dot menu on each column header).
## Reordering columns
Drag a column header to reorder columns. Dimensions and measures can be interspersed freely when no pivot is active.
## Sorting
Sorting is set on the query, not on the table — use the column header menus or the **Sort** button. See [Sorting](/docs/explore-analyze/workbooks/querying-data#sorting).
Pivot columns are the exception: they can also be sorted from the table itself, as described under [Pivots](#pivots).
## Pivots
Pivoting a dimension turns its unique values into columns, creating a cross-tab view. Pivots are
configured in the **Pivot** section of the **Fields** tab, below the **Fields** section that lists
the columns. The **Pivot** section holds:
* **Rows** — dimensions listed down the left.
* **Columns** — dimensions whose values become column headers. Drag a dimension here to pivot it.
* **Values** — the measures filling the cells.
* A **Measures** placeholder that decides where the measure names go. Drag it to **Columns** to run
them across the top, or to **Rows** to run them down the side.
Pivot behavior:
* Pivot columns can be sorted by clicking the column header, including the totals column.
* Pivot columns can also be sorted by row values — click a row number to sort by that row, including the totals row.
* Pivoted columns can be hidden, but hiding is indexed to the specific value (e.g. hiding `status: returned`), not position.
### Pivot presets
Open the menu in the corner of the Pivot section to apply a preset — a one-shot arrangement of the three drop zones that you can rearrange further afterwards:
* **Table (default)** — dimensions to **Rows**, measures to **Values**, measure names across the
top; the standard cross-tab.
* **Vertical records** — the transpose: dimensions to **Columns**, measures to **Values**, and the
**Measures** placeholder to **Rows**, so the measure names run down the side and a single record
reads as a list of labelled values. This produces the values-as-rows layout. On a query with no
measures, **Values** stays empty and the table shows its pivot headers alone.
**Reset**, beside the menu, restores the **Table (default)** arrangement — the two write the same
layout. It is disabled while the pivot already matches it.
## Row grouping
Group rows by their leading dimensions into collapsible headers, so each grouped
dimension's value is shown once for the whole group instead of repeated on every row.
Turn it on with the **Group rows** toggle in the **Values** section of the **Style**
tab. The toggle is disabled — with a tooltip explaining why — unless the table has at
least two row dimensions and no column pivot. You can also ask [Analytics
Chat](/docs/explore-analyze/analytics-chat) to group the table's rows.
With grouping on, every row dimension except the innermost becomes a group level: rows
nest one header per level, kept contiguous by that dimension's first-appearance order,
with the header showing the dimension's value once instead of on every row underneath.
Turning it on also adds a **Row grouping** section to the **Style** tab, with:
* **Expand all groups** / **Collapse all groups** — set every group's expand state at once.
* **Subtotals** (Σ) — show an aggregated value for every measure on each group header row.
See [Per-group subtotals](#per-group-subtotals).
* Text color, background color, and text formatting for the group header rows.
Each group-level column's card in the **Columns** section gets a **Default state**
control (**Expanded** or **Collapsed**) that sets whether new groups on that level start
open or closed. It only sets the default — a viewer's own expand/collapse of an
individual group isn't affected by changing it afterwards.
### Per-group subtotals
With grouping on, the **Subtotals** (Σ) toggle in the **Row grouping** section adds an
aggregated value for every measure to each group header row, at every nesting level. Not
to be confused with [pivot subtotals](#subtotals), which add a **Total for ‹value›**
column per pivot group rather than a value per group header row.
Each subtotal is queried at its own group's grain rather than summed from the rows below
it. That distinction matters for any measure that doesn't add up: an `avg` shows the true
average over the group instead of an average of averages, and a `count_distinct` counts a
value once even when it appears in several child groups. As a result a subtotal can
legitimately differ from the sum of the rows beneath it — for a distinct count it is
usually smaller.
It also means the value stays correct on a collapsed group, where there are no visible
rows to add up at all.
Measures built on window functions (`RUNNING_TOTAL`, `OFFSET`, or a SQL `OVER` clause)
can't be recomputed at another grain, so their cells stay empty on group rows — the same
exclusion that [row, column, and pivot
totals](/docs/explore-analyze/workbooks/querying-data#subtotals) make. Period-comparison
columns are also left empty. The toggle is disabled unless grouping is active and the
values are placed in columns, since the [values-as-rows layout](#pivot-presets) has no
per-measure column for a subtotal to sit under.
## Column field options
Configure individual columns in the **Columns** section of the **Style** tab (or via the dropdown arrow on a field in the **Fields** section):
| Option | Description |
| ---------------------- | ------------------------------------------------------------------------------------------------ |
| **Label** | Override the column header text |
| **Alignment** | Left, center, or right. Defaults to the column's data type — numbers right, everything else left |
| **Vertical alignment** | Top, middle, or bottom — where the value sits when the row is taller than one line |
| **Word wrap** | Allow cell content to wrap to multiple lines |
| **Width** | **Flex** (a weight) or **Fix** (pixels) — see [Column width](#column-width) |
| **Hide** | Show or hide the column |
When a cell's value is too long for its column and word wrap is off, the value is truncated with an ellipsis. Hover over a truncated cell to see its full value in a tooltip.
## Column width
Each column's width is set independently on its card in the **Columns** section of the **Style** tab, using the **Flex** / **Fix** pair and the number next to it. A column is in one of two modes:
* **Flex** (the default) — the column shares the table's available width with the other flexible columns, in proportion to its **weight**, entered next to the mode as `×1`. A column with weight `2` is twice as wide as a column with weight `1`. New columns start at weight `1`, so by default all columns share the width equally and the table stretches to fill its tile.
* **Fix** — the column is locked to an exact pixel width, entered as `px`, and no longer participates in the weight-based distribution. Fixed columns keep their width regardless of the tile size; if the fixed columns don't fill the tile, the remaining space is left blank. Pixel widths start at 10 and step in tens.
You can mix the two: give a label column a fixed width and let the metric columns flex to share the rest.
### Setting widths
* **In the panel** — pick **Flex** or **Fix** for a column and enter its weight or pixel value. Switching to **Fix** reuses the column's last fixed width if it had one, and otherwise starts from the width it's currently drawn at.
* **By dragging** — drag the border between two column headers on the table. The change is saved into the column's current mode (a flexible column keeps flexing at its new relative size; a fixed column updates its pixel width).
* **Fit** — the **Fit** button on a column's card sizes it to its content and switches it to **Fix** at that width.
Fit measures the rows currently rendered on screen. If a wider value is further down a long, scrolled table, fit again after scrolling to it.
### Bulk controls
At the top of the **Columns** section, three buttons apply a width mode to every column at once:
* **Flex** — converts every column to flexible, deriving each weight from its current width (so the layout doesn't jump). Enabled only when at least one column is fixed.
* **Fix** — freezes every column at its current rendered pixel width. Enabled only when at least one column is flexible.
* **Fit** — sizes every column to its content and fixes it at that width, in a single action. Like the per-column **Fit**, it measures the rows currently rendered on screen.
Under a column pivot, a width set on a measure applies to every column generated for that measure.
## Showing columns as links, bars, sparklines, or images
By default, a column displays its value. It can also be drawn as an inline link, an inline bar, a sparkline, or an image — from the column's card in the **Columns** section of the **Style** tab.
Each card has a **Value** toggle, a **Bars** / **Sparkline** pair, and separate **Link**
and **Image** buttons:
* **Value** shows the formatted value in the cell. It **composes** with a bar or a sparkline rather than excluding it — a cell can show a bar and its value together.
* **Bars** and **Sparkline** are mutually exclusive: picking one replaces the other, and clicking the active one again clears it.
* **Link** and **Image** are independent of each other, so a column can be both — an image with a declared link renders as a clickable thumbnail. Switching one off leaves the other alone.
* With nothing active, **Value** is forced on and its toggle disabled — so a plain column always shows its value.
* Each button is offered only where it applies: bars and sparklines on **numeric** columns, **Link** only on a column whose dimension declares links or uses `format: link` (see [Inline links](#inline-links)), and **Image** where the dimension declares `format: imageUrl` or the column's values are image URLs (see [Images](#images)). In link mode the value *is* the link text, so **Value** stays forced on there too; image mode draws the thumbnail instead of the value, so **Value** is forced off and its toggle disabled. With **Link** and **Image** both on, image wins: the thumbnail replaces the link text and becomes the click target. Only bars and sparklines leave the toggle free, because only they draw something the value can sit beside.
### Inline links
A column renders its values as clickable links when the data model says so — links are
declared in the semantic layer, not authored in the chart.
There are two ways to get one:
* **[`links`](/docs/data-modeling/dimensions#links) on the dimension.** The link marked
`primary` renders inline on the cell value. Every link on the dimension — primary or not
— stays available from the [cell menu](#cell-menu).
* **[`format: link`](/reference/data-modeling/dimensions#format) on the dimension**, when the
value already *is* a URL. Use the object form to show a label instead of the raw URL.
For a column whose dimension declares links, selecting the **Link** button reveals a picker for
which of the declared links renders inline (inert when the dimension declares only one).
Clicking the link text opens it; clicking elsewhere in the cell selects it and opens the
cell menu as usual.
### Inline bars
Display a numeric column as a proportional in-cell bar by selecting the **Bars** button on the column's card. Each bar's length reflects the value's magnitude within the column's range.
| Option | Description |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Show value** | The shared **Value** toggle on the column's card. Turn it off for a bar-only cell. |
| **Positive bar color** | Fill color for non-negative bars |
| **Negative bar color** | Fill color for negative bars (used when the column contains both positive and negative values) |
| **Bar scale** | **Auto** anchors each bar at zero and scales to the column's largest value, so even the smallest value still shows a bar; **Manual** scales against the bounds you set |
| **Lower / Higher bound** | The minimum and maximum (in the column's raw units) used when **Bar scale** is **Manual**. Pre-filled to match the Auto range, so switching modes doesn't shift the bars. |
When a column contains both positive and negative values, bars are drawn in both directions from a centered zero baseline — positive values to the right, negative to the left — using the positive and negative bar colors. Values outside the scale are clamped to a full or empty bar, but the displayed number is always the true value.
### Sparklines
Display a numeric column as a **sparkline** — a mini trend chart in each cell that plots the measure across a time dimension. Select the **Sparkline** button on the column's card in the **Columns** section of the **Style** tab.
A sparkline needs a time dimension to use as its horizontal axis. When you switch a column to **Sparkline** and pick its horizontal axis time dimension, that dimension is **removed from the table query** (if it was there): the table shows one row per remaining dimension, correctly aggregated, while the sparkline plots the measure's value across the time dimension. Values are always correct for any measure type, including counts of distinct values, averages, and custom measures.
Removing the dimension is a one-way change — turning the sparkline back off doesn't restore it to the query. Add it again yourself if you want it back.
The time dimensions on offer come from the **semantic view the query is built on**, not from the query itself — so the **Sparkline** button is disabled whenever that view has no time dimension, and equally for a query not built on a view at all, however many time dimensions the query selects.
| Option | Description |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Horizontal axis** | The time dimension plotted along the sparkline. Picked for you when the view offers only one; choose it here when there are several. |
| **Granularity** | The time bucket for the horizontal axis. Defaults to the dimension's granularity in the query if present, otherwise month. |
| **Type** | **Area** (filled line, the default), **Line**, or **Bar** (mini columns). |
| **Line color** / **Area color** | Stroke color for line and bar; fill color for area. |
| **Line width** | Stroke width in pixels (Line and Area types). |
| **Show value** | The shared **Value** toggle on the column's card — shows the measure's value for the row, the same number the cell would show without a sparkline, as a headline next to it. Turn it off for a chart-only cell. Its setting carries over when you switch between inline bars and sparkline. |
A row needs at least two data points to draw a sparkline; cells with fewer fall back to the formatted value.
Choose a reasonable granularity for the data's time span. An overly fine granularity (e.g. by the second over several years) makes the background queries return many rows and can compromise the chart's performance.
Internally, sparklines are powered by additional queries grouped by time dimension and granularity: measures sharing the same dimension and granularity are fetched together, so a chart with several sparklines runs at most one extra query per distinct dimension and granularity combination.
### Images
A column whose values are image URLs can render them as inline thumbnails.
A dimension declared with
[`format: imageUrl`](/reference/data-modeling/dimensions#format) renders as thumbnails
straight away, with no chart configuration — the data model is enough. Any other string
column offers an **Image** button on the column's card once its values look like image
URLs, so you can render an unmodelled column of URLs without changing the data model.
| Option | Description |
| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| **Image shape** | **Square** or **Circle**. |
| **Image fit** | **Fit whole image** shows the whole image inside the box; **Fill and crop** fills the box, cropping the overflow. |
| **Image height** | Thumbnail height in pixels, 20 by default and up to 100, in steps of ten. The row grows to fit it and doesn't shift as images load. |
Cells that can't render an image degrade rather than break: an empty value, a total or
subtotal row, and a URL whose scheme isn't `https` or `data:image/*` all render as plain
text, with no request issued. A permitted URL that fails to load leaves a neutral
placeholder of the same size, so the row never reflows.
Copying a cell copies the URL, not the image — the same value that lands in an export.
A thumbnail is clickable only where the dimension also declares
[`links`](/docs/data-modeling/dimensions#links); the `imageUrl` format makes a value an
image source, not a destination. **Image** and **Link** are independent buttons, so a
column that declares both can render a clickable thumbnail — or you can switch either off
on its own.
Images at `https` URLs are fetched directly by the viewer's browser, so they must be
publicly reachable — Cube does not proxy or cache them. `data:image/*` values are
rendered inline and need no network access.
## Cell menu
Left-clicking a table cell opens a context menu with any [`links` defined on the
dimension](/docs/data-modeling/dimensions#links) — drill into another dashboard
(`dashboard:`) or open an external URL (`url:`) — plus **Copy value** and, where
they apply, **Drill down** options.
**Drill down** works in two ways:
* On a **measure** cell whose data model defines drill members, it opens the
underlying rows behind that value.
* On a **time-dimension** cell shown at a granularity, it opens a submenu of
finer granularities. Picking one re-buckets the **same measures** to that
granularity, scoped to the period you clicked — drilling `2024` to **Quarter**
shows the four quarters of 2024. You can keep drilling, then close the view to
return to the original table.
Because a week can straddle two months or quarters, **Week** is never a drill
target from a coarser granularity — a week cell drills only to **Day** and finer.
## Selecting and copying cells
You can copy a single cell or a rectangular block of cells straight out of the
table — copied values paste cleanly into a spreadsheet or text editor.
### Copy a single cell
Click the cell to select it, then press **Cmd/Ctrl+C** or choose **Copy value**
from the [cell menu](#cell-menu). Either way copies that cell's raw
(unformatted) value.
### Select and copy a range
* **Click and drag** across the cells to sweep out a rectangle, or **Shift+click**
a second cell to extend the selection from the first.
* Press **Cmd/Ctrl+C** to copy the selected cells. When more than one cell is
selected, the cell menu also offers **Copy values** and **Copy values with headers** — the
latter prepends a row of column titles.
Copied cells use the raw (unformatted) values. A range is joined by tabs across
columns and newlines across rows, so it lands in a spreadsheet as a matching
grid; a single cell copies as just its value.
The same single-cell and range selection works in the **Results** table on the query
panel, not just the table visualization.
## Style options
The **Style** tab controls the visual appearance of the table. It is organized into sections — **Headers**, **Values**, **Totals**, **Columns**, **Borders**, and **Pagination** — each with its own **Reset** button in the section header that appears once you've changed anything in that section.
### Headers, Values, and Totals formatting
Each of these sections has a compact formatting toolbar that sets, for that part of the table:
* **Colors** — text and background color (the paired color chip).
* **Text format** — **bold**, **italic**, and **underline** toggles.
* **Alignment** — left, center, or right.
* **Vertical alignment** — top, middle, or bottom (**Values** only; headers and totals are always a single line).
* **Overflow** — truncate or wrap long content.
When you don't pick a color, the element uses the theme's default, which adapts to light and dark mode automatically.
Alignment is set independently in each of the three sections, and their defaults differ: headers align left, totals align right, and values follow the column's data type — numbers right, everything else left. A column's own alignment, set in the **Columns** section, overrides all three.
The **Values** alignment control shows left until you set it, even where numeric columns are drawn right-aligned. Choosing left there is still meaningful: it pins numeric columns to the left instead of letting them follow their data type.
Vertical alignment is always available, but it only has a visible effect once a row is taller than one line — set **Overflow** to wrap, or turn on **Word wrap** for an individual column, and the shorter cells beside a wrapped value will sit at the top, middle, or bottom as you choose. Values are vertically centred unless you choose otherwise. A per-column setting overrides the table-wide one.
The **Values** section also holds these table-wide row toggles, shown as icon buttons:
* **Row numbers** — show or hide the row number column.
* **Hover row** — highlight the row under the pointer.
* **Row banding** — alternate the row background with a secondary color (its color chip sits next to the toggle).
The **Totals** section holds the **Column totals** and **Row totals** toggles (see [Totals](#totals)).
### Borders
The **Borders** section of the **Style** tab controls which table borders are drawn. Five borders are configurable, each with its own toggle, color, and width (1–10 pixels):
| Border | Description | Default |
| ---------------------- | --------------------------------------------------------------------- | ------- |
| **Horizontal borders** | Lines between value rows | On |
| **Vertical borders** | Lines between columns, running through the header, values, and totals | Off |
| **Header border** | The line under the header row | On |
| **Totals border** | The line above the totals row | On |
| **Outer border** | The frame around the whole table | Off |
Each border is a row in the section: click the **icon button** to toggle the border on or off, and use the **color** swatch and **width** input next to it to style it. Color and width are only available while the border is on. When you don't pick a color, the border uses the theme's default border color, which adapts to light and dark mode automatically.
Next to the first border's width input is a **link** toggle that syncs all widths: while it's on, editing that width applies the same value to every border, so the whole table shares one line weight. It turns on automatically when all borders already share a width, and off as soon as they differ (for example after applying the Accounting preset).
Because every border has its own color and width, styles like a financial report — a thick line under the header, above the totals, and around the table, with thin lines between value rows — are a few clicks away. Or one: see presets below.
#### Border presets
Open the menu in the corner of the Borders section to apply a preset — a one-shot configuration of all five borders that you can tweak further afterwards:
* **Rows (default)** — horizontal row lines with header and totals separators; the standard look.
* **Full grid** — all borders on; a spreadsheet look.
* **Minimal** — dark lines under the header and above the totals, nothing else.
* **Accounting** — thin light row lines with a thick dark header separator, totals separator, and outer frame; a financial-statement look.
* **None** — no borders at all.
Use **Reset** in the section header to clear all border settings.
### Pagination
The **Pagination** section holds a single toggle, off unless you turn it on. With it off, the table renders every row in one scrollable grid; with it on, rows are split into pages of **50**, with page controls below the table. The page size is fixed and can't be changed.
Totals are pinned to the bottom of the table rather than being part of the paged rows, so a totals row stays visible on every page.
## Conditional formatting
The **Formatting** tab applies cell or row styling based on conditions you define. A **rule** picks the field to evaluate, and holds one or more **cases** — each pairing a condition with the styling to apply when it is met. Configure:
1. The field to evaluate
2. The condition of each case — an operator and, for most operators, a value to compare against (see [Conditions](#conditions))
3. The styling that case applies (background color, text color, bold/italic/underline, and an [indicator icon](#indicator-icons))
Use **Add case** to give a rule more than one condition — a three-tier indicator, for example, is three cases in one rule. When several could apply to the same cell, the first matching case in the first matching rule wins.
To make a background transparent, open the color picker and clear the hex value.
### Conditions
The operators offered depend on the evaluated field's data type. **is null** and **is not null** are available for every type and take no value.
**Text**
| Operator | Input |
| ------------------------------ | ------- |
| **is**, **is not** | A value |
| **contains**, **not contains** | A value |
| **starts with**, **ends with** | A value |
| **is null**, **is not null** | — |
**Numbers**
| Operator | Input |
| ------------------------------------------- | -------------------------- |
| **is**, **is not** | A value |
| **greater than**, **greater than or equal** | A value |
| **less than**, **less than or equal** | A value |
| **between** | A lower and an upper bound |
| **is null**, **is not null** | — |
**Dates**
| Operator | Input |
| -------------------------------------------------------------- | ----------------------- |
| **is** | A date |
| **before date**, **after date** | A date |
| **between** | A start and an end date |
| **in the last N days** | A number of days |
| **this week**, **this month**, **this quarter**, **this year** | — |
| **is null**, **is not null** | — |
**Booleans**
| Operator | Input |
| ---------------------------- | ----- |
| **is true**, **is false** | — |
| **is null**, **is not null** | — |
### Which cells a rule styles
Each rule's **Format** control sets how far its styling reaches:
* **Source column** — only the column the rule evaluates. The default.
* **Entire row** — every cell in a matching row.
* **Select columns** — a chosen set of columns, picked from a checklist in the same popover.
**Select columns** starts with the rule's own field checked. Clearing every column would leave a rule that styles nothing, so it falls back to **Source column** instead.
Under a column pivot, a row spans several values of the same measure (one per pivot combination), so a row-spanning rule (**Entire row** or **Select columns**) styles each measure cell by its own combination's value. Cells with no single value to test — row headers, row dimensions, and any rule whose source dimension has been pivoted into columns — are left unstyled.
### Rule presets
Picking an entry from the menu next to **Add rule** adds a ready-made rule you can then adjust. It holds three groups:
* **Indicators** — cases that mark values with an [icon](#indicator-icons) instead of recoloring them.
* **Conditional formats** — cases that tint the cell: **Above threshold**, **Below threshold**, **2-tier traffic light** (green at or above the threshold, red on the rest), and **3-tier traffic light** (green above the threshold, amber exactly at it, red on the rest).
* **Color scales** — gradient rules, described under [color scale](#color-scale).
Both indicator and conditional-format presets compare against a single threshold, which starts at `0` — adjust it to the value you're comparing against.
### Indicator icons
A conditional-formatting case can also draw an **icon** beside the cell value — an arrow up, an arrow down, a check mark, a cross, a dash, or a dot. This is how you show a goal comparison: set the condition to the threshold you care about, and the icon marks the cells that meet it.
Each case carries its own icon settings, on the row below its text formatting. Pick a shape from the **Icon** dropdown to turn the icon on; the **No icon** option turns it back off. Once a shape is chosen you can also set:
* **Icon color** — starts from the case's text color and is independent from then on, so changing the text color later leaves the icon as you set it.
* **Icon placement** — **right** of the value (the default) or **left**.
The icon is drawn next to the value, never instead of it. On a column displayed as [inline bars](#inline-bars) or [sparklines](#sparklines), the icon sits beside the bar or the trend line, and it still appears when that column's value is hidden.
Each case draws one icon, so a three-tier indicator — arrow up above target, dash within tolerance, arrow down below — takes three cases. Totals rows are never given an icon.
To add a ready-made indicator rule, open the [preset menu](#rule-presets) and pick from **Indicators**:
* **Achieved or failed** — a green check mark at or above the threshold, a red cross on every other value.
* **Above or below** — a green arrow up at or above the threshold, a red arrow down on every other value.
* **Standout or not** — a dot on the values strictly above the threshold, a muted dash on the rest.
All three are two-case, with their icons already set; add the third case yourself if you want a three-tier indicator.
### Color scale
A **color scale** (heat map) tints each cell of a numeric column with a gradient based on where its value falls in the column's range. It is a rule type in the **Formatting** tab, alongside conditional formatting — switch a rule between **Conditional** and **Scale** with the **Type** selector inside the rule.
To add one quickly, open the [preset menu](#rule-presets) and pick from **Color scales**:
* **Outstanding values** — a two-color scale that highlights the highest values.
* **Divergent values** — a three-color scale that distinguishes low, middle, and high values.
* **Traffic light gradient** — a red–yellow–green three-color scale.
A color scale requires a **numeric source column**. For non-numeric columns (strings, dates, booleans), the **Scale** type is disabled — use conditional formatting instead.
Configure the gradient with three stops, laid out top-to-bottom to mirror the column:
* **Start** (low end) — anchored at the column **Minimum** by default, or a custom **Number** or **Percentile**.
* **Center** (optional middle) — **Disabled** by default, which produces a two-color gradient. Enable it for a three-color gradient anchored at the **Midpoint**, **Average**, **Median**, or a custom **Number** / **Percentile**.
* **End** (high end) — anchored at the column **Maximum** by default, or a custom **Number** or **Percentile**.
Each stop has its own color. The anchor determines *which value* in the column the stop's color is pinned to:
* **Minimum** / **Maximum** — the lowest / highest value in the column.
* **Midpoint** — the halfway point of the range, i.e. `(minimum + maximum) / 2`. Independent of how the values are distributed.
* **Average** — the mean of all values (sensitive to outliers).
* **Median** — the middle value when sorted (robust to outliers).
* **Number** — a fixed value you enter.
* **Percentile** — a value at the given percentile (e.g. `90` = the 90th percentile).
For the computed anchors (Minimum, Maximum, Midpoint, Average, Median), the dropdown previews the resolved value from your data.
Use **Reverse color scale** to swap the Start and End colors. **Treat nulls as zero** is on by default, coloring null/blank cells as zero; turn it off to leave them uncolored.
The scale is normalized **per column** — each targeted column uses its own value range, taken from the column's data cells only. With [totals](#totals) enabled, the row totals and [pivot subtotal](#subtotals) columns are still colored by the scale of the measure they aggregate, but they don't widen its range: a total past the column's highest stop is clamped to that stop's color. Cells in the column totals row are never colored by a scale, since totals rows don't take formatting rules.
## Totals
Enable column totals and row totals from the **Totals** section of the **Style** tab. Totals rows and columns are styled separately from body cells.
### Subtotals
When the table is pivoted by two or more dimensions, you can also enable **subtotals** — a bold **Total for ‹value›** column appended after each pivot group's columns, at every nesting level. Not to be confused with [per-group subtotals](#per-group-subtotals), which add a value to each row group's header row rather than a column per pivot group. Subtotals combine freely with column and row totals, work in workbooks and on dashboards, and — like the other totals — carry over from the results table when you switch the chart type to Table. See [Subtotals](/docs/explore-analyze/workbooks/querying-data#subtotals) for details and limitations.
# Axes
Source: https://docs.cube.dev/docs/explore-analyze/charts/configuration/axes
Configure axis titles, grid lines, label formatting, dual Y-axis, and reference lines.
The **Axes** section in the Style tab controls the appearance and behavior of the X axis and Y axis (or left and right Y axes in dual-axis charts). Most axis settings are available for all Vega-based chart types (bar, line, area, scatter, heatmap).
## X axis
| Setting | Description |
| ---------------- | ------------------------------------------------------ |
| **Show axis** | Toggle the X axis on or off |
| **Title** | Custom axis label — leave blank to hide the axis title |
| **Grid lines** | Toggle grid lines perpendicular to the X axis |
| **Labels** | Toggle axis tick labels |
| **Label angle** | Rotate axis labels (useful for long category names) |
| **Label format** | Number or date format applied to axis labels |
## Y axis (left)
| Setting | Description |
| ---------------- | ---------------------------------------------------------------- |
| **Show axis** | Toggle the Y axis on or off |
| **Title** | Custom axis label |
| **Grid lines** | Toggle horizontal grid lines |
| **Labels** | Toggle axis tick labels |
| **Label format** | Number format applied to axis labels (e.g. `$,.0f` for currency) |
| **Min / Max** | Override the axis scale minimum and maximum values |
## Dual Y axis (right axis)
Add a second Y axis on the right side of the chart to plot a series on a different scale. This is useful for combining measures with different units or magnitudes — for example, showing order count on the left axis and average order value on the right.
To use the right axis:
1. In the **Series configuration** for a specific series, change the **Y axis** assignment from **Left** to **Right**.
2. The right axis settings appear in the Style tab — configure its title, labels, and scale independently from the left axis.
## Reference lines
Add horizontal reference lines to mark a target, threshold, or benchmark value. Reference lines appear at a fixed Y value across the full width of the chart.
To add a reference line:
1. In the **Y axis** section of the Style tab, click **Add reference line**.
2. Set the **Value** — the Y position of the line.
3. Optionally set a **Label** that appears next to the line.
4. Configure **Color** and **Style** (solid, dashed, or dotted).
You can add multiple reference lines to the same axis. Each is configured independently.
### Reference line labels
A reference line with no label renders as a bare rule — the reader has to already know what it
means. Type text into **Label** to name it, for example `Target 18k`.
Once a label is set, two alignment controls appear next to it:
| Control | Options | Description |
| -------------- | -------------------- | --------------------------------------- |
| **Horizontal** | Start / Middle / End | Where along the line the text sits |
| **Vertical** | Above / Below | Which side of the line the text sits on |
New labels start at **End** and **Above** — the right end of the line, clear of the bars. The label
takes the line's color, so it always reads as part of that line. Clearing the label removes the text
and leaves the line itself untouched.
A label on a line near the top or bottom edge of the plot flips to the other side of the line
automatically, so the text stays inside the chart area.
Reference lines are only available on the left Y axis. For right-axis reference lines, use a [custom Vega-Lite spec](/docs/explore-analyze/charts/custom).
# Color & stacking
Source: https://docs.cube.dev/docs/explore-analyze/charts/configuration/color-and-stacking
Configure color palettes, stacking behavior, stacked segment sorting, and legend placement for charts.
The **Color** section in the Style tab controls two related things: how series are colored, and how they are stacked or grouped relative to each other. These settings apply to charts that color by a Color channel: bar, line, area, scatter, and heatmap. Chart types that color some other way, such as pie, funnel and sankey, pick their palette from a **Palette** dropdown of their own — but the [palettes](#palettes) below are the same ones they offer.
## Color faceting
Mapping a field to the **Color** channel in the Fields tab creates one series per unique value in that field, each assigned a color from the active palette. How this field is typed affects the available options:
| Data type | Palette options |
| -------------------------- | -------------------------------------------------------- |
| **String / nominal** | Discrete palettes (one color per category value) |
| **Date / temporal** | Discrete (nominal) or continuous (temporal gradient) |
| **Numeric / quantitative** | Discrete (nominal) or continuous (quantitative gradient) |
For strings, only discrete bucketing is available. For dates and numbers, you can choose whether each value gets its own distinct color (discrete) or a gradient is applied across the range (continuous).
If a date or numeric field only offers discrete options, try casting the field to the appropriate type.
## Palettes
### Discrete palettes
For discrete data, colors from the palette are applied in palette order, matching the sort order of the query results. Change the palette order by adjusting the sort in your query.
Enable **Reverse colors** to reverse the order colors are applied.
### Continuous (gradient) palettes
For continuous data, the gradient palette maps to the range of values in your results. Select a different gradient from the palette dropdown to change the color treatment.
### Custom palettes
If none of the provided palettes fit your needs, you can build a custom palette for an individual chart:
1. Open the **Color** section of the Style tab.
2. Select the last option in the palette menu: **Custom palette**.
3. Click **Customize** to open the palette editor.
4. Add, remove, and edit individual hex colors.
5. Use **Open hex code editor** to toggle a bulk editor and paste a comma-separated list of hex values.
To reuse a custom palette across charts, copy the hex codes and paste them into the custom palette editor of another chart.
## Series color controls
When no Color channel is assigned (single measure, no color-by), each series gets an individual color picker in the **Series** section of the Style tab. Click the color swatch next to a series to change its color.
## Stacking options
The stacking behavior controls how multiple series at the same X-axis value are positioned relative to each other. The available modes:
| Option | Behavior |
| ------------- | ------------------------------------------------------------------- |
| **Automatic** | Cube selects the best option based on chart type and data |
| **Stack** | Series are stacked on top of each other |
| **Group** | Series are placed side-by-side (bar charts only) |
| **Overlay** | Series are drawn on top of each other from the same baseline |
| **Stack %** | Series are stacked and normalized to 100% — tooltip shows raw value |
### Per-axis stacking
Stacking can also be set independently per Y-axis series in the Y-axis series configuration. This enables creating grouped clusters of stacked sub-groups — for example, two stacked grouplets placed side-by-side.
## Stacked segment sorting
When stacking is active, control how segments within each stack are ordered using **Sort stack by**:
| Option | Description |
| ------------ | ---------------------------------------------------------------------------- |
| **Label** | Alphabetical by dimension value (default) |
| **Value** | By measure value within each individual stack (bar/column only) |
| **Sum** | By total sum across all stacks — useful for ordering by overall contribution |
| **Unsorted** | Preserves the original data order from the query results |
The **Value** sorting option is only available for bar and column charts. For area and line charts, use **Sum** or **Label**.
## Legend
The legend appears when a Color channel is assigned. A chart type without a Color channel decides for itself: a pie carries a legend regardless, a funnel starts without one, and a sankey has none to configure. Where a legend is offered, it uses the same **Legend** section:
| Option | Description |
| ------------ | --------------------------- |
| **Position** | Right, left, top, or bottom |
| **Hidden** | Remove the legend entirely |
# Data labels
Source: https://docs.cube.dev/docs/explore-analyze/charts/configuration/data-labels
Display totals directly on stacked bar charts with configurable positioning and formatting.
Data labels show the aggregate total value above each full stack on stacked bar charts. They let users read exact values without cross-referencing the Y axis.
## Enabling data labels
Open the **Fields** tab and enable the **Data Labels** toggle.
## Label settings
| Setting | Description |
| ------------- | ----------------------------------------------------------------------------------------- |
| **Format** | Number format applied to the label value (e.g. `,.0f` for integers, `$,.2f` for currency) |
| **Font size** | Size of the label text |
| **Position** | Outside end, Inside end, Inside center, or Inside base |
# Configure charts
Source: https://docs.cube.dev/docs/explore-analyze/charts/configuration/index
Reference for chart configuration options — series mapping, color, axes, tooltips, and data labels.
The chart configuration panel has two tabs — **Fields** and **Style** — available for every chart type. Configuration options are divided into per-channel field assignments and per-chart style controls.
| Page | What it covers |
| --------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- |
| [Series mapping](/docs/explore-analyze/charts/configuration/series-mapping) | Assigning query columns to chart channels (X, Y, color, size, tooltip) |
| [Series configuration](/docs/explore-analyze/charts/configuration/series-configuration) | Per-series mark type, color, and individual display options |
| [Color & stacking](/docs/explore-analyze/charts/configuration/color-and-stacking) | Color palettes, stacking mode, stacked segment sorting, and legend placement |
| [Small multiples](/docs/explore-analyze/charts/configuration/small-multiples) | Splitting a chart into a grid of panels, one per value of a dimension |
| [Axes](/docs/explore-analyze/charts/configuration/axes) | Axis titles, grid lines, label formatting, dual Y-axis, and reference lines |
| [Tooltips](/docs/explore-analyze/charts/configuration/tooltips) | Which fields appear on hover and how they are formatted |
| [Data labels](/docs/explore-analyze/charts/configuration/data-labels) | Labels shown on chart marks — positioning, formatting, and stacked totals |
# Series configuration
Source: https://docs.cube.dev/docs/explore-analyze/charts/configuration/series-configuration
Configure individual series — mark type, color, and display options — independently from the global chart settings.
Series configuration lets you override chart settings on a per-series basis. Access it from the **Fields** tab by expanding an individual series in the Y-axis section.
## Mark type override
Each series can use a different mark type, independent of the global chart mark. This enables composite charts — for example, plotting one measure as a bar and another as a line on the same chart.
Available mark types per series:
* **Bar**
* **Line**
* **Area**
* **Scatter** (point)
To create a bar + line chart:
1. Add two measures to the Y axis.
2. Expand the second series in the Fields tab.
3. Set its mark type to **Line**.
4. Optionally assign it to the **Right Y axis** in the same panel (see [Axes](/docs/explore-analyze/charts/configuration/axes)).
## Series color
When no **Color** channel is assigned (single-color charts with no color-by dimension), each series has an individual color picker. Click the color swatch to open the picker and set a custom color for that series.
This setting has no effect when a Color channel is active — in that case, colors are managed by the palette in [Color & stacking](/docs/explore-analyze/charts/configuration/color-and-stacking).
## Y axis assignment
For charts with a dual Y axis, assign each series to either the **Left** or **Right** Y axis. This controls which axis scale the series uses.
See [Axes](/docs/explore-analyze/charts/configuration/axes) for configuring the right axis title, labels, and scale.
## Data labels per series
Enable data labels for an individual series without enabling them globally. Configure label position, font, and format independently for each series.
See [Data labels](/docs/explore-analyze/charts/configuration/data-labels) for the full configuration reference.
## Stacking override
On charts with multiple Y-axis series, you can set stacking behavior per-axis rather than globally. This lets you create grouped clusters where each cluster is internally stacked — expand the Y-axis settings for a specific series to access this option.
# Series mapping
Source: https://docs.cube.dev/docs/explore-analyze/charts/configuration/series-mapping
Assign query columns to chart channels — X axis, Y axis, color, size, and tooltips.
Series mapping controls which query columns are assigned to which chart channels. This is done in the **Fields** tab of the chart configuration panel.
## Chart channels
| Channel | Description | Applicable chart types |
| -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------ |
| **X** | The horizontal axis dimension or time field | Bar, line, area, scatter, heatmap, boxplot |
| **Y** | The vertical axis measure | Bar, line, area, scatter, boxplot |
| **Color** | Creates one series per unique value; controls stacking behavior | All Vega-based types |
| **Size** | Scales point radius by a numeric measure | Scatter, map |
| **Split by** | Repeats the chart once per unique value, as a grid of panels — see [small multiples](/docs/explore-analyze/charts/configuration/small-multiples) | Bar, line, area, scatter |
| **Tooltip** | Fields shown on hover | All types |
| **Theta** (pie) | The measure that determines slice size | Pie |
| **Source** / **Target** (sankey) | The dimensions holding the node each flow leaves and enters | Sankey |
| **Value** (sankey) | The measure that determines ribbon thickness | Sankey |
## Assigning fields
Drag a field from the **Available fields** list at the bottom of the Fields tab into the target channel slot. You can also drag an already-assigned field between channels.
Fields can appear in more than one channel simultaneously — drag from **Available fields** to add a field to a second channel without removing it from the first. For example, you can assign the same dimension to both the X axis and the tooltip.
## Multiple measures on Y
To plot multiple measures as separate series, drag additional measures into the **Y** channel. Each measure renders as its own series, with independent color and style settings in [series configuration](/docs/explore-analyze/charts/configuration/series-configuration).
## Removing a field
Drag a field out of its channel slot back to **Available fields**, or click the **×** on the field token to remove it.
# Small multiples
Source: https://docs.cube.dev/docs/explore-analyze/charts/configuration/small-multiples
Split a chart into a grid of panels, one per value of a dimension, for side-by-side comparison across segments.
**Small multiples** repeat one chart once per value of a dimension, laying the copies out in a grid. Instead of a single plot with a dozen overlapping series, you get a dozen small panels that can be read and compared at a glance.
It is a setting on an existing chart, not a chart type of its own: a bar, line, area, or scatter chart stays that chart and simply gets drawn per panel.
## Splitting a chart into panels
The **Small multiples** section appears in the Fields tab for bar, line, area, and scatter charts.
Pick a dimension in **Split by** and the chart is replaced by one panel per value of that dimension. Clearing the picker returns the chart to a single plot.
Only dimensions are offered. Splitting by a measure is not supported — a measure has no discrete values to make panels from.
## Options
These options appear once a **Split by** dimension is chosen.
| Option | What it does |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Grid** | Columns × rows, up to 5 × 5. Both are preselected from the number of distinct values in the split dimension, so a four-value dimension opens as a 2 × 2 grid. |
| **Axis scales** | Whether every panel is drawn against the same scale (**Shared**) or each scales to its own data (**Independent**). |
| **Sort panels by** | Orders the panels by the dimension's own values (**Value**) or by a measure (**Measure**). |
| **Sort order** | **Ascending** or **Descending**. |
### Shared vs. independent axis scales
Every panel is labeled with its own X and Y axes either way. What the setting changes is the scale those axes are drawn against.
**Shared** is the default, and it is usually what makes small multiples worth using: every panel is drawn against the same scale, so the panels are directly comparable and a taller bar really does mean a bigger number.
Switch to **Independent** when the question is about the *shape* of each series rather than its magnitude — "what does each warehouse's weekly pattern look like", where one high-volume warehouse would otherwise flatten every other panel to a straight line.
### How many panels are drawn
The grid bounds the render: a chart split into 3 × 2 draws at most six panels, so a high-cardinality dimension can never produce hundreds of unreadable ones. The largest grid is 5 × 5, or twenty-five panels.
When a dimension has more values than the grid has tiles, the chart draws the first ones in the current sort order. The underlying query is unaffected, so no data is lost — a bigger grid simply shows more of it.
Sorting interacts with this: with **Sort panels by** set to a measure and a descending order, the grid keeps the top panels by that measure, which is usually more useful than the first ones alphabetically.
## Interaction with other settings
Chart settings apply to all panels at once — axis titles and formats, colors and palettes, data labels, reference lines, and tooltips are configured once and take effect in every panel. A reference line is drawn once inside each panel, against that panel's own scale when the scales are independent.
Data labels and reference lines have to be switched on *after* the split, not before — see the limitations below.
A legend is shared across the grid rather than repeated per panel. Clicking a legend entry hides that series in every panel.
## Limitations
* **One split dimension.** One dimension fills the grid, panel by panel. Splitting by two dimensions at once — one down the rows and another across the columns — is not supported.
* **Cartesian charts only.** Bar, line, area, and scatter. Pie, sankey, table, KPI, heatmap, boxplot, map, and HTML charts cannot be split.
* **Panel labels are not configurable.** Each panel is labeled with its dimension value; the font, size, and color are fixed.
* **The split is enabled on a single-view chart.** A chart that already carries data labels, a reference line, or a second Y axis series cannot be split — turn the split on first. The order is the only constraint: once a chart is split, data labels and reference lines can be added freely and are drawn in every panel.
# Tooltips
Source: https://docs.cube.dev/docs/explore-analyze/charts/configuration/tooltips
Control which fields appear when a user hovers over a chart mark.
Tooltips appear when a user hovers over a data point in a chart. By default, Cube automatically populates the tooltip with all fields mapped to chart channels. You can customize which fields are shown and in what order.
## Default behavior
When a chart is first created, all fields assigned to the chart (X, Y, color, size) are included in the tooltip automatically. The tooltip is enabled by default on all Vega-based chart types (bar, line, area, scatter, heatmap, boxplot), on map charts, and on sankey charts.
[Sankey](/docs/explore-analyze/charts/chart-types/sankey) is the exception to the first rule: it seeds the tooltip with every column *except* the source, target and value it already draws.
## Configuring tooltip fields
Open the **Fields** tab of the chart configuration panel. The **Tooltip** section shows the list of fields currently included in the tooltip.
| Action | How |
| ------------------ | ------------------------------------------------------------ |
| **Add a field** | Drag a field from **Available fields** into the Tooltip slot |
| **Remove a field** | Click the **×** on a field token in the Tooltip list |
| **Reorder fields** | Drag field tokens within the Tooltip list |
## Disabling tooltips
To turn off tooltips entirely, remove all fields from the Tooltip slot. When the Tooltip slot is empty, no tooltip is shown on hover.
## Tooltip behavior on stacked charts
On stacked bar charts, the tooltip is scoped to the individual stack segment under the cursor — it shows the value for that specific series at that X position, not the total stack value. To display the stack total, use [data labels](/docs/explore-analyze/charts/configuration/data-labels) with the **Simple totals** option.
# Custom visualizations
Source: https://docs.cube.dev/docs/explore-analyze/charts/custom
Edit raw Vega-Lite specs to build visualizations beyond what the chart builder supports.
All standard Cube chart types are built on [Vega-Lite v5](https://vega.github.io/vega-lite/). The visual builder covers the most common patterns, but for anything beyond that — multiple layers with transforms, complex dual-axis rules, custom aggregations, or annotations — you can edit the raw spec directly.
For fully freeform layouts using HTML and CSS, see [HTML charts](/docs/explore-analyze/charts/chart-types/html).
## When the spec editor appears
The chart panel automatically falls back to the **spec editor** when the current Vega-Lite spec has features the visual builder cannot represent:
* Top-level `transform` clauses
* More than two non-text layers
* Certain dual-axis + rule mark combinations
In these cases, the visual builder is replaced by a code editor showing the full Vega-Lite JSON.
## Opening the spec editor manually
Click the **Edit spec** button in the chart panel (visible on Vega-based charts) to open the spec editor. This lets you inspect and modify the generated spec even when the visual builder is active.
Changes take effect immediately in the preview.
## Spec structure
Cube generates standard Vega-Lite v5 specs. The key fields:
```json theme={"dark"}
{
"$schema": "https://vega.github.io/schema/vega-lite/v5.json",
"mark": { "type": "bar" },
"encoding": {
"x": { "field": "order_items.status", "type": "nominal" },
"y": { "field": "order_items.count", "type": "quantitative" },
"color": { "field": "order_items.status", "type": "nominal" }
}
}
```
Any valid Vega-Lite v5 spec is supported. Refer to the [Vega-Lite documentation](https://vega.github.io/vega-lite/docs/) for the full spec reference.
## AI-generated specs
The AI agent can generate a full Vega-Lite spec from a natural language description. After running a query, prompt the agent — for example: *"Show this as a heatmap of order count by status and month"* — and the agent will write the complete spec.
AI-generated specs produce valid Vega-Lite v5 JSON and can be further refined in the spec editor.
# Charts
Source: https://docs.cube.dev/docs/explore-analyze/charts/index
Visualize query results as charts, tables, KPIs, and maps directly inside Cube workbooks.
Every tab in a Cube workbook has a chart panel. Once your query returns results, switch to the chart panel to pick a visualization type and configure how the data is displayed — from a simple bar chart to a composable KPI tile or a fully custom Vega-Lite spec.
## Chart types
Cube has built-in support for bar, line, area, pie, scatter, heatmap, KPI, map, boxplot, and more. See the [chart types overview](/docs/explore-analyze/charts/chart-types) for a full list and guidance on which type fits which data shape.
### Selecting a chart type
Chart type icons at the top of the chart panel let you quickly switch between common layouts — grouped vs. stacked bars, line vs. area, table vs. chart view — without opening the full configuration panel.
When you run a query on a new tab, Cube shows the chart type picker rather than picking for you. It outlines the type that best fits your query and labels it **Recommended** — see [Recommended chart type](/docs/explore-analyze/charts/chart-types#recommended-chart-type) — but choosing one is always your action. Any manual configuration you apply is preserved when you change query fields.
### Resetting chart settings
Sections of the configuration panel — fields, series, pivot, column widths — each carry their own **Reset** control, which clears that section back to its defaults. Resetting does not change the chart type you picked.
## Generate charts with AI
Describe the visualization you want in plain language — for example, *"show this as a stacked bar grouped by status"* — and the AI agent will build and configure it. You can also let the agent auto-suggest the best chart for your query results. This is a separate mechanism from the **Recommended** outline in the chart type picker: the agent writes a Vega-Lite spec and applies it, while the recommendation only marks a built-in type for you to pick.
Charts generated by the AI are written as [Vega-Lite v5](https://vega.github.io/vega-lite/) specs and can be further edited in the [custom visualization](/docs/explore-analyze/charts/custom) spec editor.
## Configure charts
Adjust the appearance of any chart through the configuration panel:
* **[Color and stacking](/docs/explore-analyze/charts/configuration/color-and-stacking)** — Color palettes, series coloring, and stacking behavior
* **[Series configuration](/docs/explore-analyze/charts/configuration/series-configuration)** — Per-series mark type, color, and display options
* **[Series mapping](/docs/explore-analyze/charts/configuration/series-mapping)** — Assign query fields to chart channels (X, Y, color, size)
* **[Axes](/docs/explore-analyze/charts/configuration/axes)** — Axis titles, grid lines, scale, and reference lines
* **[Tooltips](/docs/explore-analyze/charts/configuration/tooltips)** — Fields shown on hover
* **[Data labels](/docs/explore-analyze/charts/configuration/data-labels)** — Values displayed directly on chart marks
## Custom visualizations
When the built-in chart types don't cover your use case, write a raw [Vega-Lite spec](/docs/explore-analyze/charts/custom) or use an HTML template with Handlebars. Both options give you complete control over the rendered output.
# Dashboard Agent
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/dashboard-agent
Ask questions about a dashboard and its charts — and change its filters and time granularity — using a conversational AI assistant docked alongside the dashboard.
The Dashboard Agent is a conversational assistant docked as a panel on a
[published or embedded dashboard][ref-dashboards]. Ask it questions in plain
language about what's on the dashboard — and it answers using **live queries
against your semantic model**, not just the numbers already rendered on screen.
Because it queries the model directly, the agent can go beyond what a chart
shows. It can drill into a single chart, re-aggregate a metric, break a number
down by a dimension, pull a longer time range than the chart displays, and
summarize trends across the dashboard.
It can also **apply filters and change the time granularity** on the dashboard
you're viewing — just ask in plain language (for example, *"filter to last 30
days"* or *"group by month"*). These changes affect only your own session and
are never saved to the dashboard — see
[Changing filters and time granularity](#changing-filters-and-time-granularity).
There are **two different agent surfaces** with different capabilities:
* **Dashboard Agent** (this page) — for viewers of a published or embedded
dashboard: question-and-answer about what's on screen, plus changing your
own filters and time granularity.
* **[Workbook Agent][ref-workbook-agent]** — full authoring inside a
[workbook][ref-workbooks]: creating reports, building and editing
dashboards, and running analysis.
The published Dashboard Agent can adjust **your own view** — filters and time
granularity, for your session only — but intentionally cannot author or change
the **saved** dashboard.
## Opening the agent
Open a published dashboard and toggle the agent panel from the dashboard header
to start a conversation. Use **New Chat** to start a fresh thread.
The panel is docked to the right side of the dashboard and is **resizable** —
drag its edge to give the conversation more or less room. On narrow screens it
switches to a **compact mode** so it stays usable next to the dashboard.
## Focusing on a specific chart
Chart-level follow-ups and attached-chart chips are **rolling out**. If you
don't see the **Ask a follow up question** action on a chart's menu yet, it
hasn't reached your deployment.
To ask about one chart in particular, open that chart tile's menu (`⋮`) and
choose **Ask a follow up question**. This opens the agent with that chart
**attached as context**, so your question is scoped to that chart's data.
Attached charts appear as **removable chips** in the chat input. To drop a
chart's focus, remove its chip. You can attach more than one chart to ask about
them together.
When a chart is attached, the agent receives a read-only snapshot of it — the
underlying SQL, the data, and the chart spec — along with the dashboard's
**current filter and time-granularity state**. It uses this as context for
answering and never edits the chart's definition. You can still ask it to change
the dashboard's filters or time granularity — see
[Changing filters and time granularity](#changing-filters-and-time-granularity).
## Changing filters and time granularity
Ask the agent, in plain language, to filter the dashboard or change a time
granularity, and it applies the change to the controls in front of you:
* *"Filter this dashboard to orders with status completed."*
* *"Show only the Outerwear category."*
* *"Filter to the last 30 days."*
* *"Group by month."*
* *"Reset the filters."*
How it works and what to expect:
* **Only controls that exist on the dashboard can be changed.** The agent can
change a filter or time granularity only when the dashboard has a matching
[filter or time-granularity control widget][ref-controls] for that dimension.
If there's no control for what you asked about, the agent tells you so instead
of changing anything. A control whose [visibility][ref-control-visibility] is
**Disabled** also can't be changed.
* **Time granularity stays within the control's allowed granularities.** If a
time-granularity control restricts which granularities viewers can pick, the
agent only chooses from that set.
* **Changes apply to your session only.** They're applied the same way as
changing a control by hand: reflected in the dashboard's URL — so you can
share or bookmark the filtered view — but **not** saved to the dashboard. They
don't affect what other viewers see, and the author's saved defaults are left
untouched.
This works the same way on both published dashboards and
[embedded][ref-embedding] dashboards.
## Attachments
You can attach uploaded files to the conversation, and the agent can read their
contents — text, PDF, ZIP, and image files are supported.
Pasting an image directly from your clipboard into the chat may also be
supported. Confirm this in your deployment before relying on it.
## What the Dashboard Agent can and can't do
The published Dashboard Agent can answer questions and adjust your own view, but
it never changes the **saved** dashboard.
**It can:**
* Answer questions about what's on the dashboard from the dashboard's own data
* Run **read-only queries** against the semantic model to go deeper — drill
down, break a metric down by a dimension, or pull a longer time range than a
chart shows
* Apply filters and change the time granularity on the dashboard you're
viewing, for dimensions that have a matching
[control widget][ref-controls] (your session only — see
[Changing filters and time granularity](#changing-filters-and-time-granularity))
* Explain how a metric or dimension is defined
* Look up the values available for a dimension
* Read the contents of [attachments](#attachments) you add to the chat
**It can't:**
* Save its filter or time-granularity changes to the dashboard — they apply to
your session only and never change what other viewers see
* Edit the dashboard, or add/remove widgets
* Create reports or workbooks
* Change the data model
If you need any of those, use the [Workbook Agent][ref-workbook-agent] instead.
## Limitations
* **Filter and time-granularity changes need a matching control.** The agent can
change a filter or time granularity only for dimensions that have a
corresponding [control widget][ref-controls] on the dashboard, and the change
applies to your session only — it is never saved to the dashboard. If you ask
it to change something with no control on the dashboard, it explains the
current state instead of applying a change.
* **Field switchers aren't driven by the agent.** The agent can read and change
[filters and time granularities][ref-controls]; a
[field switcher][ref-field-switcher] is not among the controls it operates, so
ask it about the data instead and change the field yourself in the control.
* **No authoring on published dashboards.** The published Dashboard Agent
cannot create reports or build dashboards, edit the saved dashboard, or change
the data model. Those live in the [Workbook Agent][ref-workbook-agent].
* **Report search covers published reports only.** When the agent searches for
existing reports, it sees published reports.
[ref-dashboards]: /docs/explore-analyze/dashboards
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-workbook-agent]: /docs/explore-analyze/workbooks/workbook-agent
[ref-controls]: /docs/explore-analyze/dashboards/widgets/controls
[ref-field-switcher]: /docs/explore-analyze/dashboards/widgets/controls#field-switcher
[ref-control-visibility]: /docs/explore-analyze/dashboards/widgets/controls#visibility
[ref-embedding]: /embedding/iframe/dashboards
# Dashboards as code
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/dashboards-as-code
Manage workbooks, dashboards, and reports as code with idempotent REST endpoints keyed by portable identifiers, so a CI/CD pipeline can apply the same definitions across deployments.
**Dashboards as code** lets you manage the reporting assets in a deployment —
[workbooks][ref-workbooks], their [dashboards][ref-dashboards], and the
[reports][ref-reports] the dashboard widgets render — from source control instead
of only through the UI. You keep each asset's definition in Git and apply it to a
deployment with the [Cube Cloud REST API][ref-api], the same way you might manage
Superset assets with `preset-cli` or infrastructure with Terraform.
Two **idempotent upsert** endpoints make this possible. Instead of tracking the
per-deployment numeric id that a `POST` returns, you address each asset by a
**portable identifier you choose** and re-apply its definition as often as you
like:
| Endpoint | Keyed by | Upserts |
| -------------------------------------------------------------------------------------- | -------------------------------- | ------------------------------------ |
| [`PUT /deployments/{deploymentId}/workbooks/by-slug/{slug}`][ref-upsert-workbook] | a deployment-scoped **slug** | a workbook (and its dashboard draft) |
| [`PUT /deployments/{deploymentId}/reports/by-public-id/{publicId}`][ref-upsert-report] | an account-unique **`publicId`** | a report |
Because the identifier is stable and lives in your repository, applying the same
definition twice is a no-op, and applying it to a second deployment (staging →
production) reproduces the same assets there.
This page covers the REST primitives available today. They are the building
blocks for an as-code workflow you assemble in your own pipeline — Cube does not
yet ship a single bundle export/apply command that wraps them.
## How the pieces fit
Three assets are involved, each with its own identity:
* A **report** is a saved query plus its visualization. Its portable identity is
a **`publicId`**: a 12-character alphanumeric (`[0-9A-Za-z]`) id that is unique
across your account. You mint it when you author the report and keep it fixed
for the report's lifetime.
* A **workbook** is the container that holds a dashboard. Its portable identity
is a **slug**: a human-readable, deployment-scoped id (the same slug a data
model targets with `links: [{ dashboard: }]` for drill-in).
* A **dashboard** is the layout — which widgets sit where. It is stored on its
workbook as `meta.dashboardDraft` and is made visible by **publishing** the
workbook. Each chart widget references a report.
The identifiers you control (`publicId`, `slug`) are what make a definition
portable. The numeric ids that `POST` responses return are per-deployment and are
resolved at apply time — you never store them in Git.
## Authenticating
These are public REST endpoints. Authenticate with a deployment API key exactly
as for the rest of the [REST API][ref-api] — see [Authentication][ref-auth] for
how to create a key and pass it. The examples below assume:
```bash theme={"dark"}
export CUBE_API_URL="https://"
export CUBE_API_TOKEN=""
export DEPLOYMENT_ID=""
```
## The apply flow
An as-code pipeline applies a dashboard bottom-up: reports first, then the
workbook that lays them out, then publish.
### 1. Author once, then export
The report and dashboard-draft definitions are large and are not meant to be
hand-written. Build the reports and dashboard once in the UI, then read them back
over the API and commit the results:
* [`GET /deployments/{deploymentId}/reports/{reportId}`][ref-get-report] returns a
report's definition.
* [`GET /deployments/{deploymentId}/workbooks/{workbookId}`][ref-get-workbook]
returns the workbook, including its `dashboardDraft`.
Assign each report a `publicId` and the workbook a `slug` of your choosing, store
those alongside the exported definitions in your repository, and treat that as the
source of truth.
### 2. Upsert each report
For every report, [upsert it by `publicId`][ref-upsert-report]. If a report with
that `publicId` already exists in the deployment it is updated with the fields you
send (same semantics as [`PUT /reports/{reportId}`][ref-update-report]);
otherwise it is created with that `publicId`.
```bash theme={"dark"}
curl -X PUT \
"$CUBE_API_URL/api/v1/deployments/$DEPLOYMENT_ID/reports/by-public-id/revqZ1x8Kp0a" \
-H "Authorization: $CUBE_API_TOKEN" \
-H "Content-Type: application/json" \
-d @report-revenue-by-month.json
```
The path `publicId` is the report's identity; the request body is the report
definition you exported (its query in `sqlQuery` / `jsonQuery`, pivot in
`pivotItems`, and visualization config in `meta`). Keep track of the numeric
`id` each response returns — the dashboard draft references reports by that
per-deployment id.
### 3. Upsert the workbook and its dashboard
[Upsert the workbook by `slug`][ref-upsert-workbook], carrying the dashboard
layout in `meta.dashboardDraft`. Only the fields you send are changed, and `meta`
is **merged** into the existing metadata rather than replacing it. The
`dashboardDraft` is validated the same way the builder validates it.
```bash theme={"dark"}
curl -X PUT \
"$CUBE_API_URL/api/v1/deployments/$DEPLOYMENT_ID/workbooks/by-slug/revenue-overview" \
-H "Authorization: $CUBE_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Revenue Overview",
"meta": { "dashboardDraft": { "...": "the exported dashboard config" } }
}'
```
Because each chart widget inside `dashboardDraft` points at a report by its
per-deployment numeric id, rewrite those references to the ids returned in
step 2 before applying the workbook to a **new** deployment. Re-applying to
the same deployment needs no rewriting — the ids are stable there.
### 4. Publish
Upserting the workbook writes the dashboard **draft**. Publish it to make it
visible to viewers with [`POST /workbooks/{workbookId}/publish`][ref-publish],
using the workbook id returned in step 3. Publishing is itself idempotent per
workbook, so it is safe to run on every apply.
## Idempotency and conflicts
Re-applying an unchanged definition is a no-op — that is the property that makes
these endpoints safe to run on every pipeline execution. When something does go
wrong, both upserts fail with a `409` rather than guessing, and the report upsert
distinguishes three cases by a `code` field in the response body so your pipeline
can react correctly:
| Endpoint | `code` | Meaning | What to do |
| ----------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------- |
| workbook & report | `upsert_branch_changed` | A concurrent writer created or deleted the asset between the access check and the write, so the request would have applied under the wrong permission check. | **Retry.** Transient; happens only under concurrent applies of the same key. |
| report | *(none)* | The `publicId` already belongs to a report in a **different** deployment. `publicId` is unique across the account. | **Permanent.** Use a different `publicId`. |
| report | `ambiguous_legacy_id` | The id matches more than one legacy report (see below), so it can't identify one. | **Permanent.** Give the intended report a `publicId` of your own (see below), then key on that. |
The upserts serialize per key (per slug, per `publicId`), so two pipeline runs
applying the same bundle at once can't create a duplicate — the loser gets a
retryable `upsert_branch_changed` instead.
## Choosing and adopting `publicId`s
A report's `publicId` is **write-once**: you can assign one to a report that
doesn't have one yet, but a report's existing `publicId` can never be changed,
because clients may already have stored it. You can supply a `publicId`:
* **On create** — pass it in the body to [`POST /reports`][ref-create-report], or
just call the [upsert endpoint][ref-upsert-report] with the id in the path.
* **On an existing report** — assign one with
[`PUT /reports/{reportId}`][ref-update-report]. This is how you bring a report
that was authored in the UI under as-code management.
Pick any distinct 12-character `[0-9A-Za-z]` id. The auto-generated placeholder
ids shown for reports that don't have a stable id yet are a reserved, non-unique
shape and are rejected with `400` — you must choose your own.
### Reports created before stable ids
Reports created before `publicId` existed don't store one; the API **synthesizes**
one from the report's internal id so every report has an id on the wire. These
synthesized ids are **not unique** — several reports can share one. The upsert
endpoint resolves a synthesized id only when it is unambiguous, adopting it as the
report's real `publicId` at that point; if it matches more than one report it
returns the `ambiguous_legacy_id` conflict above. For anything you manage as code,
don't rely on a synthesized id — assign a `publicId` you chose and key on that.
## Reference
* [Create or update a workbook by slug][ref-upsert-workbook]
* [Create or update a report by publicId][ref-upsert-report]
* [Create a report][ref-create-report] · [Update a report][ref-update-report]
* [Publish dashboard][ref-publish]
* [Building dashboards in the UI][ref-dashboards]
[ref-dashboards]: /docs/explore-analyze/dashboards
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-reports]: /docs/explore-analyze/workbooks/querying-data
[ref-api]: /api-reference/introduction
[ref-auth]: /api-reference/authentication
[ref-upsert-workbook]: /api-reference/workbooks/create-or-update-a-workbook-by-slug
[ref-upsert-report]: /api-reference/reports/create-or-update-a-report-by-publicid
[ref-create-report]: /api-reference/reports/create-a-report
[ref-update-report]: /api-reference/reports/update-a-report
[ref-get-report]: /api-reference/reports/get-report
[ref-get-workbook]: /api-reference/workbooks/get-workbook
[ref-publish]: /api-reference/workbooks/publish-dashboard
# Dashboards
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/index
Curate and share polished views of your workbook reports with stakeholders as interactive dashboards.
Dashboards let you select and organize reports from your workbooks into polished, shareable views for your team and stakeholders. Transform your exploratory analysis into production-ready deliverables by choosing which insights to highlight and present.
## Use cases
Dashboards enable you to:
* **Monitor key metrics** – Display critical business indicators that refresh with live data
* **Distribute analysis** – Package your findings into reports for regular distribution
* **Build data tools** – Develop specialized applications powered by your semantic layer
* **Present insights** – Showcase important discoveries in a polished, accessible format
## How it works
In the dashboard builder inside your [workbook][ref-workbooks], select the reports you want to include and arrange them on the canvas alongside other [widgets][ref-widgets] to tell your data story, then publish the dashboard. This gives stakeholders direct access to the insights that matter most, without the complexity of the underlying analysis.
Prefer to manage dashboards from source control? You can also apply dashboards,
workbooks, and reports to a deployment through the REST API and keep their
definitions in Git — see [Dashboards as code](/docs/explore-analyze/dashboards/dashboards-as-code).
## Data freshness
Each widget shows a [freshness](/docs/explore-analyze/workbooks/querying-data#result-freshness-and-provenance) leaf indicating how recently its data was refreshed. The dashboard's own leaf reflects its least-recently-refreshed widget, so you can see at a glance whether everything on the dashboard is up to date.
## Linking between dashboards
Dashboards can link to one another. When a table widget shows a dimension that
has [links][ref-dimension-links] defined in the data model, left-clicking a cell
opens a menu with those links: a **drill-in** link (`dashboard:`) navigates to
another dashboard in the same deployment, filtered by the values in the clicked
row, and an **external** link (`url:`) opens a URL. This is how you build
overview → detail flows — click a row in a summary dashboard to jump straight to
a focused dashboard scoped to that row.
Links are declared once on the dimension in the data model (not configured
per-dashboard), so every table that shows that dimension — in dashboards,
workbooks, embedded dashboards, and Explore — offers the same links. A drill-in
link targets a dashboard by its [slug](#dashboard-slug); see [Dimensions →
Links][ref-dimension-links] to define them.
## Dashboard slug
A dashboard can have a **slug** — a short, stable, human-readable identifier
(e.g. `orders-detail`) used to link to it from the data model. Data-model
[dimension links][ref-dimension-links] with a `dashboard:` value reference a
dashboard by its slug to drill into it.
To set it, open the dashboard in the builder, open the **options sidebar**, and
fill the **Slug** field. Slugs are **unique per deployment** and are resolved
within the current deployment, so the same model works across environments. The
dashboard's own URL is unaffected (it stays `publicId`-based).
**There is no required order between the model and the dashboard.** You can
reference a slug from a `dashboard:` link in the data model before any dashboard
uses it, or set a dashboard's slug before any link references it — neither
breaks. A link whose slug doesn't (yet) resolve is simply skipped in the cell
menu, and starts working as soon as a dashboard in the deployment claims that
slug. (If you later clear or change a slug that links still reference, the
options sidebar warns you to update those links.)
## Download as PNG, PDF, or CSV
Open a published dashboard, click the **More actions** (`⋯`) button in the
header, and choose **Download as PNG** or **Download as PDF**. The file is
named after the dashboard's title.
To export a **single chart** instead of the whole dashboard, hover the chart
widget, open its **⋮** menu, and choose **Download as PNG** or **Download as
PDF** — the same menu also offers **Download as CSV** for the chart's
underlying data. The image captures just that chart, named after the chart's
title.
PNG and PDF downloads are server-rendered snapshots — Cube re-opens the
dashboard (or, for a single chart, that chart on its own), waits for rendering
to finish, and captures the result. This can take up to a couple of minutes
for large dashboards. The control selections you currently have applied in your
browser are carried into a download you start yourself — filters, time
granularity switchers and [field switchers][ref-controls] alike. A scheduled
notification has no browser session, so its attachment renders the dashboard's
own defaults, unless that notification carries selections of its own. (CSV is
generated from the data already loaded in the chart and downloads immediately.)
Downloading the whole dashboard requires **Manage** permission on the workbook
that owns it; exporting a single chart requires the **Download data**
permission. Both are hidden entirely when an admin turns off the account-wide
[data download controls][ref-data-download-controls]. The same screenshot
mechanism powers PNG/PDF attachments on [notifications][ref-notifications] sent
after a [scheduled refresh][ref-scheduled-refreshes].
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-widgets]: /docs/explore-analyze/dashboards/widgets
[ref-controls]: /docs/explore-analyze/dashboards/widgets/controls
[ref-dimension-links]: /docs/data-modeling/dimensions#links
[ref-notifications]: /docs/explore-analyze/notifications
[ref-scheduled-refreshes]: /docs/explore-analyze/scheduled-refreshes
[ref-data-download-controls]: /admin/users-and-permissions/roles-and-permissions#restricting-data-downloads
# Styling
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/styling
Customize the appearance of a dashboard from the Styling panel in the Dashboard Builder.
You can customize the appearance of any dashboard — background, padding, widget cards, borders, titles, and fonts — from the **Styling** panel in the Dashboard Builder.
Styling is **per-dashboard**: settings are saved to the dashboard's configuration and apply equally to the dashboard in Cube and to its [embedded view](/embedding/iframe/dashboards) (both [private](/embedding/iframe/auth/private) and [signed](/embedding/iframe/auth/signed) embedding).
## Open the Styling panel
1. Open your dashboard
2. Click **Dashboard Builder** in the top bar
3. Click the gear icon to open the **Styling** panel
Click **Apply** to preview your changes. Click **Reset** to revert to the defaults. Publish the dashboard for the styling to take effect in embedded views.
## Options
All values accept standard CSS — colors as hex (`#1a1a2e`), `rgb()`, or named colors; lengths as CSS units (e.g., `16px`).
### Background
Settings that apply to the dashboard canvas as a whole.
| Option | Description | Example |
| ---------- | --------------------------------------- | --------- |
| Background | Dashboard background color | `#f9f9fb` |
| Padding | Outer padding around the dashboard grid | `16px` |
### Widgets
Settings that apply to every widget (chart card) on the dashboard.
| Option | Description | Example |
| ---------- | -------------------------------- | --------- |
| Background | Widget card background color | `#ffffff` |
| Padding | Inner padding within each widget | `12px` |
### Widgets → Borders
Border styling for widget cards.
| Option | Description | Example |
| ------ | ------------------------------------------------------------ | --------- |
| Color | Border color | `#E1E2E8` |
| Radius | Corner radius | `8px` |
| Style | Border style (`solid`, `dashed`, `dotted`, `double`, `none`) | `solid` |
| Width | Border thickness | `1px` |
### Widgets → Titles
Typography for widget titles.
| Option | Description | Example |
| ----------- | ----------------- | ----------- |
| Color | Title text color | `#1a1a2e` |
| Font Size | Title font size | `18px` |
| Font Weight | Title font weight | `600` |
| Font Family | Title font family | `system-ui` |
# AI summary
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/widgets/ai-summary
Generate natural-language summaries of dashboard data on demand using an AI agent.
AI summary widgets generate a natural-language summary of the data shown on the dashboard. Write a prompt — for example, *"Summarize the key trends and call out anything unusual"* — and the configured [AI agent][ref-agents] produces a Markdown narrative based on the current dashboard state, including the data behind every chart and the values of any active [controls][ref-controls].
## Adding an AI summary
In the [dashboard builder][ref-workbooks], add an AI summary from the **Add Widgets** menu in the toolbar. The widget opens with a prompt editor — write your prompt and click **Generate Summary** to produce the first response.
## Use cases
* Add an executive summary at the top of a dashboard
* Surface anomalies or notable changes since the last refresh
* Explain numbers to stakeholders who want context without diving into the charts
* Generate a TL;DR for periodic dashboard distributions
## Writing a prompt
Write the prompt as if you were briefing an analyst — describe the audience, the angle you want, and any specific things to call out. The agent has access to the full dashboard context, so you can reference charts and controls by name.
## Caching
Once generated, the summary is **cached** with the widget. Viewers loading the dashboard see the cached response immediately — they don't pay for regeneration on every page load.
## Staleness detection
The widget keeps a checksum of the dashboard state at the time of generation: the queries behind each chart, the active control values, and the chart configuration. When any of those change, the widget marks the cached summary as **stale** and shows a refresh prompt so viewers know the narrative may no longer match the data.
Click the refresh icon (or open the widget menu and choose **Refresh**) to regenerate using the saved prompt against the latest state.
## Choosing the agent
By default, AI summaries use the agent configured at the dashboard level. You can override the agent per widget when you need a particular [agent's][ref-agents] tooling, model, or guardrails for a specific summary.
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-agents]: /admin/ai
[ref-controls]: /docs/explore-analyze/dashboards/widgets/controls
# Charts
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/widgets/charts
Display reports from your workbook as visualizations on the dashboard.
Chart widgets display a report from the source [workbook][ref-workbooks] as a visualization on the dashboard. Every chart on a dashboard is backed by a tab in the workbook — when you publish a new version of the workbook, the chart updates accordingly.
## Adding a chart
In the [dashboard builder][ref-workbooks], open the **Charts** picker in the toolbar and select one or more reports from the workbook. A chart is added to the canvas for each selected report, inheriting the report's query, configured visualization, and styling.
If the picker is empty, create the report first in a workbook tab — only published reports are available to add to a dashboard.
## Interaction with controls
Charts respect the [controls][ref-controls] placed on the same dashboard — filters, time granularity switchers and field switchers. A single control can drive multiple charts at once: its value is applied to every chart whose query references the targeted dimension.
If some controls are incompatible with a chart's query, the chart skips them and shows a warning icon. See [Incompatible controls][ref-incompatible-controls] for details.
## Updating charts
Charts on a published dashboard reflect the most recent published version of the workbook. To change the query, switch the chart type, or restyle the visualization, edit the underlying report in the workbook and publish a new version of the dashboard.
**Edit in Workbook** in the chart's settings menu opens that report directly, on its own workbook tab.
### Filters carried into the workbook
The dashboard's [filters][ref-controls] come along, so the report opens showing the same slice of data the chart was showing rather than re-running over everything.
They sit in the report's filter bar next to its own filters, tinted violet to set them apart, and read the same way — member, operator, value. Hovering one says **From dashboard**. What is different is that they are not the report's:
* they apply to the results on screen, and are **not** saved to the report. Publishing the workbook again will not pin them onto the chart for everyone.
* their values cannot be changed in the workbook. To explore a different value, add your own filter on the same field, or change the dashboard's control and reopen the chart.
* they can be dropped, but not kept: **Clear filter** on the chip removes it for the rest of the session and the report re-runs without it. Reopening the chart from the dashboard brings the current values back.
To change what a viewer of the dashboard can filter by, edit the [filter control][ref-controls] there — not the report.
When the report already filters the same field, both filters apply — exactly as they do on the dashboard, so the numbers match the chart you came from. A carried filter that says precisely what the report already says is left out rather than shown twice.
Only filters are carried. A [time granularity switcher or field switcher][ref-controls] on the dashboard does not follow into the workbook, so a report opened this way shows its own granularity and its own fields.
## Title
Each chart shows the name of the underlying workbook tab as its title. To rename it, open the tab in the workbook, rename the tab, and republish the dashboard — the new name flows through to every chart backed by that tab.
Use **Hide Title** in the widget's settings menu to suppress the title on the dashboard — useful when the chart's content already makes the subject obvious, or when an adjacent [text widget][ref-text] provides its own heading. Choose **Show Title** in the same menu to bring it back.
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-controls]: /docs/explore-analyze/dashboards/widgets/controls
[ref-incompatible-controls]: /docs/explore-analyze/dashboards/widgets/controls#incompatible-controls
[ref-text]: /docs/explore-analyze/dashboards/widgets/text
# Controls
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/widgets/controls
Filter, time granularity switcher, field switcher, and parent widgets that let dashboard viewers change what's shown on the dashboard.
Controls are widgets that let dashboard viewers change what's shown without leaving the dashboard. The dashboard builder offers four control types:
* [Filter](#filter) — Narrow the data shown on the dashboard
* [Time granularity switcher](#time-granularity-switcher) — Change the granularity of time-based dimensions
* [Field switcher](#field-switcher) — Swap which dimension or measure the charts are built on
* [Parent](#parent) — Re-point several other controls at once from a single dropdown
The first three each target a member from your semantic model, and apply the viewer's choice to every [chart][ref-charts] on the dashboard whose query references that member. Filters and time granularity switchers change *how a member is queried* — which rows come back, which buckets they fall into. A field switcher goes further and changes *which member is queried at all*. A parent control works one level up: it targets no member of its own and drives *other controls* instead.
## Filter
Filter widgets let viewers narrow down the data shown on the dashboard. In the [dashboard builder][ref-workbooks], open the **Add Controls** menu in the toolbar and choose **Filter**. The new filter is added in an unconfigured state — click **Configure Filter** (or open the widget's settings menu) to pick a semantic view and a dimension.
### Operators by dimension type
The available operators depend on the type of the underlying dimension:
| Dimension type | Operators |
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
| **String** | `is`, `is not`, `contains`, `not contains`, `starts with`, `not starts with`, `ends with`, `not ends with`, `is null`, `is not null` |
| **Number** | `is`, `is not`, `greater than`, `greater than or equal`, `less than`, `less than or equal`, `is null`, `is not null` |
| **Time** | `is`, `is not`, `before date`, `before or on date`, `after date`, `after or on date`, `between`, `relative date`, `is null`, `is not null` |
### Single vs. multiple selection
Filters can allow either a single value or multiple values. Configure this when adding or editing the filter — multi-select is the default for string dimensions, while time and number dimensions default to a single value.
### Default values
You can set a default value that's applied when the dashboard loads. Defaults are useful for scoping the dashboard to "this quarter" or "the user's region" without requiring viewers to interact with the filter first.
There are two ways to set a default:
* **Static default** — pick a value (or values) directly in the filter. Every viewer sees the same default.
* **User attribute default** — resolve the default from the viewer's [user attribute][ref-user-attributes] at load time, so each viewer sees their own personalized default. [Time granularity switchers](#time-granularity-user-attribute-default), [field switchers](#field-switcher-user-attribute-default) and [parent controls](#parent-user-attribute-default) support this too.
Static defaults are configured by interacting with the filter in the dashboard builder — the value you select is saved on the widget and applied to every viewer when the dashboard loads.
#### User attribute default
Use the **User attribute default** toggle in the filter's edit sidebar to pre-fill a filter from the viewer's [user attribute][ref-user-attributes]. When the dashboard loads, Cube looks up the attribute value for the current viewer and applies it as the filter's default.
This is useful for scoping a dashboard to the viewer's own slice of the data — for example, defaulting a **Region** filter to the viewer's `region` attribute, or a **Sales rep** filter to their `email`.
To configure it:
In the dashboard builder, click the filter widget's settings menu and choose **Edit Filter**.
Scroll to the **User attribute default** switch and turn it on.
Select the [user attribute][ref-user-attributes] whose value should be used as the default. Only attributes defined in your account appear in the picker.
How the attribute value is matched to the filter:
| Attribute type | How it's applied |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| **String**, **Number** | Used as a single value. Works with single-value operators like `is` / `is not`, and is also accepted by multi-select filters as a one-item selection. |
| **String array**, **Number array** | Used as a list of values, one per array entry. Empty values are dropped. |
Empty, `null`, or unresolvable attribute values are skipped — the filter falls back to whatever static default it has, or no default if none is set.
The user attribute default only seeds the filter's *initial* value. Viewers can still change the filter unless its [visibility](#visibility) is set to **Disabled**, in which case the resolved attribute value is locked in for that viewer. Values passed via URL parameters also take precedence over user attribute defaults, so deep links continue to work.
### Faceted filters
When multiple filters target dimensions from the same semantic view, you can mark them as **faceted**. Faceted filters scope each other's value lists — selecting a value in one filter narrows the options shown in the others, so viewers only see combinations that exist in the data.
For example, on a sales dashboard with a **Country** filter and a **City** filter, marking both as faceted means selecting `United States` in the Country filter limits the City filter to U.S. cities only.
## Time granularity switcher
Time granularity switchers let viewers change the granularity of time-based dimensions on the dashboard — for example, switching a revenue chart from daily to weekly or monthly. The widget targets a single time dimension and applies the chosen [granularity][ref-granularities] to every chart that groups by that dimension.
In the [dashboard builder][ref-workbooks], open the **Add Controls** menu in the toolbar and choose **Time Granularity**. The new switcher is added in an unconfigured state — open its settings to pick a semantic view and a time dimension.
### Allowed granularities
By default, viewers can choose between **day**, **week**, **month**, **quarter**, and **year**. You can narrow this list in the widget's settings to only expose the granularities that make sense for the dashboard.
For time dimensions backed by a `TIMESTAMP` or `DATETIME` column, sub-day granularities (**second**, **minute**, **hour**) are also available. `DATE`-typed columns don't expose sub-day granularities, since they would bucket the entire day into a single point.
Custom granularities defined in the [data model][ref-granularities] aren't offered in this list yet — the switcher exposes the built-in granularities only.
### Default granularity
You can configure a default granularity that's applied when the dashboard loads. If no default is set, charts use the granularity that was saved on the underlying report — viewers can still switch granularities, but the dashboard opens with each chart at its original granularity.
#### User attribute default
The default above is one granularity for everyone. To give each viewer their own, turn on **User attribute default** in the switcher's settings and pick a [user attribute][ref-user-attributes]. When the dashboard loads, Cube reads that attribute for the current viewer and opens the control on the granularity it names.
This is how one dashboard opens at the interval each audience works in — daily for the operations team, monthly for the executives who read the same charts — from a single published dashboard.
To configure it:
In the dashboard builder, open the widget's settings menu and choose **Edit Control**.
Below **Visibility**, turn on the **User attribute default** switch.
Select the [user attribute][ref-user-attributes] to resolve. Only attributes defined in your account appear in the picker.
The attribute value is matched against the **granularity names** the switcher [allows](#allowed-granularities) — `day`, `week`, `month`, `quarter`, `year`, and `second`, `minute`, `hour` where the dimension exposes them — ignoring case and surrounding spaces, so an attribute reading `Week` resolves to `week`. The names are matched, not the labels the control displays, so one attribute works the same for viewers in every language.
| Attribute type | How it's applied |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------- |
| **String**, **Number** | Matched against the granularity names as a single value. |
| **String array**, **Number array** | The first entry that names an allowed granularity wins. The switcher is single-select, so the rest are ignored. |
A value that isn't one of the switcher's [allowed granularities](#allowed-granularities) — or is empty, `null`, or unresolvable — is ignored rather than forced, and the control falls back to the [default granularity](#default-granularity). Attributes are set per user and the allowed list per dashboard, so the two can drift apart without anyone editing either; the safe reading of an unusable value is "no opinion".
The attribute is resolved for the viewer, not baked into the dashboard. Editing the attribute's value changes what that viewer opens on the next time the dashboard loads; it never rewrites the published dashboard, so the default granularity you set in the builder stays intact for everyone else.
Viewers can still switch granularities unless the control's [visibility](#visibility) is set to **Disabled**, and their own pick outranks the attribute for the rest of the session. A granularity passed [in the URL](#sharing-the-current-selection) outranks both, so deep links continue to work. If a [parent control](#parent) drives this switcher and the option the viewer is on maps a granularity to it, that mapping decides the granularity — a mapping the author made for that arrangement is more specific than a per-viewer starting point. An option that [leaves the switcher empty](#children) has no opinion, so the attribute still seeds it; one set to **Reset to default** returns the switcher to the granularity it opens on *for that viewer* — the attribute granularity when one resolves, otherwise the saved [default granularity](#default-granularity) — and leaves it untouched when neither resolves.
## Field switcher
A field switcher lets viewers change *which* dimension or measure the charts are built on — swapping a revenue chart's breakdown from **Status** to **City**, or its measure from **Order count** to **Total revenue** — without leaving the dashboard or opening the report.
Where a [filter](#filter) narrows the rows and a [time granularity switcher](#time-granularity-switcher) rebuckets them, a field switcher replaces the member itself in the chart's query. One dashboard can then answer several questions that would otherwise need a chart each.
In the [dashboard builder][ref-workbooks], open the **Add Controls** menu in the toolbar and choose **Field Switcher**, then click **Configure Field Switcher** to set it up.
### Choosing what it switches
A field switcher works on one member kind at a time — set **Field Type** to either **Dimension** or **Measure**. The rest of the settings follow from that choice:
| Setting | What it does |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| **Field Type** | Whether this control swaps a dimension or a measure. Switching the type resets the settings below it, since the two draw from different pools. |
| **Dimension to Replace** / **Measure to Replace** | The member the charts are currently built on — the one this control stands in for. |
| **Alternatives** | The members viewers may switch to. The replaced member is always offered as well, so a viewer can get back to the original view. |
Alternatives come from the control's own semantic view, and must be the same kind as the replaced member — a dimension switcher offers dimensions, a measure switcher offers measures.
### Default option
The member a viewer starts on is set the same way a filter's static default is: by picking it in the control while you're in the dashboard builder. The selection is saved on the widget and applied to every viewer when the dashboard loads. If you never pick one, the dashboard opens on the replaced member.
Removing a member from **Alternatives** after it was serving as the default clears the default, so viewers can't start on a member the control no longer offers.
### Default granularity per option
When a dimension switcher offers **time** dimensions, each of them can carry its own granularity, set under **Default granularity per option**.
This exists because a swap otherwise inherits whatever granularity the chart already had. A viewer moving from **Created at** to a **Completed at** that only makes sense monthly would get the replaced dimension's daily buckets, and the author would have no way to say otherwise.
Each time option is either pinned to a granularity or left at **Inherit from the chart**, which is the default and the behavior of every control configured before this setting existed. Non-time options don't have the setting, and a measure switcher has no granularities to speak of.
If the dashboard also has a [time granularity switcher](#time-granularity-switcher) pointed at the swapped-in dimension, the viewer's own pick wins over the per-option granularity — a granularity a viewer actively chose outranks one the author set as a starting point. A granularity control nobody has touched does not.
### User attribute default
Like [filters](#user-attribute-default), [time granularity switchers](#time-granularity-user-attribute-default) and [parent controls](#parent-user-attribute-default), a field switcher can start each viewer on their own member. Turn on **User attribute default** in the control's settings and pick a [user attribute][ref-user-attributes]; when the dashboard loads, Cube reads that attribute for the current viewer and opens the control on the member it names.
This is how one dashboard opens on the breakdown each audience cares about — a **Breakdown** switcher opening on `region` for one team and `channel` for another, from a single published dashboard.
The attribute seeds the *selection*, exactly as the [default option](#field-switcher-default-option) does, and loses to a pick the viewer has already made. A member passed [in the URL](#sharing-the-current-selection) outranks both, so deep links keep working. If a [parent control](#parent) drives this switcher and the option the viewer is on maps a member to it, that mapping decides the member; an option that [leaves the switcher empty](#children) has no opinion, so the attribute still seeds it, and one set to **Reset to default** returns the switcher to the member it opens on *for that viewer* — the attribute member when the switcher still offers it, otherwise its [default option](#field-switcher-default-option). A value that isn't among the **Alternatives** is ignored rather than forced: the attribute is set per user and the options are set per dashboard, so the two can drift apart without anyone editing either, and the safe reading of an unusable value is "no opinion" — the control falls back to the default option.
### What the swap preserves
The swapped-in member is queried under the replaced member's output name, so everything the chart configured against that column keeps working across a switch — column formatting, sorting, pivots, and conditional formatting rules all survive, rather than resetting each time the viewer picks a different member.
### When a chart can't take the switch
Cube applies the switch in two places: to the SQL that runs, and to the query description the chart formats its results with. It applies the switch only if **both** take it — otherwise the chart would be labelled and formatted as one member while showing another's numbers, which nothing on screen would reveal.
When only one half can take it, the chart keeps rendering the member it was built on and shows a notice reading **"The Field switcher could not be applied to this chart"**. The usual reason is a query Cube can't read back as semantic members — a hand-written one, or one built with a `JOIN` or `UNION`. The rest of the dashboard still switches.
A chart whose query doesn't use the replaced member at all is a different case: it is simply out of the control's scope, exactly as it would be for a filter, and shows no notice.
A chart with a [period comparison][ref-charts] is a narrower case: the comparison can't follow a member switch, so the chart applies the switch and drops the comparison, saying so in its own notice rather than silently showing a comparison that no longer matches the data.
## Parent
A parent control is a dropdown of options you define. Picking one re-points a whole row of other controls at once — so a viewer makes a single choice instead of adjusting three or four filters by hand.
Unlike the other control types, a parent control targets no member and never touches a chart query directly. It applies values to the controls it *drives* — its **children** — and those children then apply themselves to charts exactly as if the viewer had operated each one. Filters, time granularity switchers and [field switchers](#field-switcher) can all be children; a parent control cannot be a child of another parent control.
For example, an **Analysis** parent with the options `Retail`, `Wholesale` and `Promo` can set a **Channel** filter, a **Minimum order value** filter, and a **Date range** filter to a different combination for each option. Viewers see one dropdown; you can [hide](#visibility) the children if the individual values aren't worth showing.
In the [dashboard builder][ref-workbooks], open the **Add Controls** menu in the toolbar and choose **Parent**, then click **Configure Parent** to set it up. The editor has two tabs — **Options** and **Children**.
Add the child controls to the dashboard *before* the parent control. The **Children** tab can only map controls that already exist, so a parent added to an empty dashboard has nothing to drive yet.
### Options
On the **Options** tab, type a label and click **Add** for each entry you want in the dropdown. Options appear as chips — remove one with its close button. A parent control can hold up to 50 options.
Renaming an option later doesn't disturb the values you've mapped to it, so you can reword a label without redoing the mapping — but if the control has a [user attribute default](#parent-user-attribute-default), that match is by label, so rename the attribute's values with it.
### Children
On the **Children** tab, pick a control from **Child control**, then give each of the parent's options a value for it. Each row renders *that child's own control* — a time granularity switcher's row shows its granularity picker, limited to the granularities that switcher allows; a filter's row shows that filter's operator and value inputs; a field switcher's row shows its member picker, limited to the members that switcher offers — its **Alternatives**, plus the replaced member. So the values you can offer are exactly the ones a viewer could pick in the child itself.
Repeat for each control you want the parent to drive. Every option/child pair can be in one of three states:
| State | What happens when the viewer picks that option |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **A value** | The child is set to that value. |
| **Reset to default** | The child is put back to the value it opens on *for that viewer*. When the child takes its default from a **User attribute default** — a [filter](#user-attribute-default)'s, a [time granularity switcher](#time-granularity-user-attribute-default)'s, or a [field switcher](#field-switcher-user-attribute-default)'s — that's the attribute value; otherwise it's the child's own static default: no filtering on that dimension for a filter, its [default granularity](#default-granularity) for a time granularity switcher, its [default option](#field-switcher-default-option) for a field switcher. Turn on the row's **Reset to default** switch. |
| **Left empty** | The child is left alone — it keeps whatever value the viewer already had. Use this deliberately when an option shouldn't have an opinion about a particular child. |
While the parent's settings are open, the children it drives are highlighted on the canvas, so you can see the scope of the mapping at a glance.
### One parent per child
A control can be driven by only one parent control at a time. Mapping a child that another parent already drives **moves** it rather than sharing it — the editor warns you before you save, naming the parent that currently owns it.
### Mapping status on child controls
Once a dashboard has at least one parent control, every filter, time granularity switcher and [field switcher](#field-switcher) on it shows a small indicator reporting how it's driven:
| Status | Meaning |
| ----------------- | ----------------------------------------------------------------------------------- |
| **Fully driven** | Every option of the owning parent sets this control. |
| **Partly driven** | Only some of the owning parent's options set this control; the rest leave it alone. |
| **Not driven** | No parent control maps this one. |
Click the indicator to jump straight to the **Children** tab of the parent that owns that control, with it already selected. For a control nothing drives yet, the click opens the first parent control on the dashboard — topmost, then leftmost — so you can map it.
### Default option
A parent control's default is set the same way a filter's static default is — by interacting with the control in the dashboard builder. The option you select is saved on the widget and applied to every viewer when the dashboard loads; there's no static default field in the parent's settings.
Picking in the builder also applies that option's values to the children, so their saved defaults line up with the parent's and a published dashboard opens in a consistent state.
If you never pick an option, the parent opens with nothing selected and the children use their own defaults. Deleting the option that was serving as the default clears it, and the parent goes back to opening on nothing.
#### User attribute default
The default above is one arrangement for everyone. To give each viewer their own, turn on **User attribute default** in the parent control's settings and pick a [user attribute][ref-user-attributes]. When the dashboard loads, Cube reads that attribute for the current viewer and opens the control on the option it names — and drives the children with it, exactly as if the viewer had picked that option themselves.
This is how you ship one dashboard that opens differently per audience: a **Reporting period** parent whose options are `Month` and `Quarter`, opening on whichever one the viewer's `reporting_period` attribute says, with every control behind it already set to match.
To configure it:
In the dashboard builder, click **Configure Parent** on the control.
Below the **Options** and **Children** tabs — next to **Visibility** — turn on the **User attribute default** switch.
Select the [user attribute][ref-user-attributes] to resolve. Only attributes defined in your account appear in the picker.
The attribute value is matched against the **option labels**, ignoring case and surrounding spaces — an attribute reading `quarter` selects the option labelled `Quarter`. Give the options the labels your attribute already uses, or adjust the attribute values to match.
Renaming an option leaves its child mappings intact, but the attribute match is by label — so a rename that moves a label away from the values your attribute holds silently stops it resolving, with no error. The control falls back to the option you picked as the default. Rename labels and attribute values together.
| Attribute type | How it's applied |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------ |
| **String**, **Number** | Matched against the option labels as a single value. |
| **String array**, **Number array** | The first entry that names an option wins. A parent control is single-select, so the rest are ignored. |
If the value matches no option — or is empty, `null`, or unresolvable — the control falls back to the [default option](#default-option) you picked, and the children keep the arrangement that goes with it.
The attribute is resolved for the viewer, not baked into the dashboard. Editing the attribute's value changes what that viewer opens on the next time the dashboard loads; it never rewrites the published dashboard, so the default you picked in the builder stays intact for everyone else.
Viewers can still switch to another option unless the control's [visibility](#visibility) is set to **Disabled**, and their own pick outranks the attribute for the rest of the session. Values passed [in the URL](#sharing-the-current-selection) outrank both — but only what the sharer actually picked travels: a parent control has no parameter of its own, and an untouched, attribute-resolved one puts nothing in the link, so the recipient still opens on their own attribute. When the sharer did pick an option, the link carries that option's *children's* values and the recipient opens on those.
## Sharing the current selection
On a published dashboard, the values a viewer picks in the controls are reflected in the URL, so the view they are looking at is bookmarkable and shareable. Copy the address bar, send it on, and the recipient opens the dashboard with the same filters, granularities and member choices applied.
Each control type has its own parameter:
| Control | Parameter | Example |
| ------------------------------------------------------- | -------------------------------------------------------- | ------------------------------------- |
| [Filter](#filter) | `f_.=` | `f_orders.status={"value":"shipped"}` |
| [Time granularity switcher](#time-granularity-switcher) | `tg_.=` | `tg_orders.created_at=week` |
| [Field switcher](#field-switcher) | `ms_.=` | `ms_orders.status=users_city` |
The semantic view and member are the **internal names** configured on the control — not the display titles you see in the picker. A view shown as `Orders` is usually `orders` in the parameter. The same goes for the member on the right-hand side of `ms_`: it is the internal name of the member to switch to, and it is a measure rather than a dimension when the switcher's **Field Type** is **Measure**. Granularities are lowercase and must be one of the switcher's [allowed granularities](#allowed-granularities) — `day`, `week`, `month`, `quarter`, `year`, plus `second`, `minute`, and `hour` for time dimensions that expose them.
You can also write these parameters by hand to open a dashboard in a particular state — see [Pre-set dashboard filters and granularities via URL][ref-embed-url-filters] for the embedded case, which uses the same format.
What does and doesn't travel in the link:
* **Only what the viewer chose.** Values that came from the control's own configuration — a static default, a [default granularity](#default-granularity) — are not written into the URL. Every viewer already gets those from the dashboard itself, and leaving them out means a link stays correct after the dashboard's defaults change.
* **Never a personalized default.** A value resolved from a [user attribute](#user-attribute-default) stays out of the link — whether it seeded a filter, a [time granularity switcher](#time-granularity-user-attribute-default) or a [field switcher](#field-switcher-user-attribute-default) directly, or reached one through a [parent control](#parent) opening on the viewer's own option. Sharing a dashboard never pins your attribute value onto the recipient; they see it through their own attributes.
* **Every control's pick, together.** Picking in several controls puts them all in the link, including when a [parent control](#parent) sets several children at once — filters, time granularity switchers and [field switchers](#field-switcher) alike. A parent control isn't serialized itself — the link carries the values its children ended up with, so the recipient sees the same data while the parent dropdown opens on whatever default it resolves for them, which may not be the option the sharer picked.
* **Written out on published dashboards only.** Reading these parameters works anywhere, including [embedded][ref-embed-url-filters] dashboards; it's the writing that is published-only. In the dashboard builder the URL is left to the editing session, so changing a control there doesn't rewrite it.
When a dashboard opens with these parameters, they are applied on top of whatever defaults the controls carry. A parameter is ignored when nothing on the dashboard can honor it — there is no matching control for that member, the requested granularity isn't in the switcher's [allowed granularities](#allowed-granularities), or the selected member is one the field switcher doesn't offer — its **Alternatives**, plus the replaced member.
## Visibility
Each control has a **Visibility** setting that determines how it appears on the published dashboard. The setting applies to all four control types.
| Visibility | Behavior on the published dashboard |
| --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Visible** (default) | Shown on the dashboard and viewers can change its value. |
| **Hidden** | Not shown to viewers, but the control's value is still applied to the charts it targets. Use this to scope a dashboard with a fixed value — e.g., always filter to the current quarter — without exposing the control. Hiding the *children* of a [parent control](#parent) is the usual way to present one dropdown instead of the row of controls behind it. |
| **Disabled** | Shown on the dashboard so viewers can see the active value, but they cannot change it. |
Set the visibility from the **Visibility** dropdown when editing the control. **Hidden** controls remain visible in the dashboard builder so editors can reconfigure them, but disappear from the published view.
## Interaction with charts
This section applies to filters, time granularity switchers and field switchers. A [parent control](#parent) has no member of its own and never applies to a chart itself, so it doesn't appear in any chart's [Controls mapping](#controls-mapping) — it acts only through the children it drives, and it's those children that show up here.
When a control is added to a dashboard, it's automatically wired up to every [chart][ref-charts] whose query already uses the same member. Charts that don't reference that member are left alone, so a dashboard can mix scoped and unscoped views by default. Filters and time granularity switchers always target a dimension; a [field switcher](#field-switcher) targets a dimension or a measure depending on its **Field Type**, and scopes on whichever it is set to. You can override this default per chart from its [Controls mapping](#controls-mapping) — disable the control for that chart, or remap it onto a different member.
### Incompatible controls
If controls of a certain type are incompatible with a particular chart's query, the chart skips all controls of that type and renders the data without them. Each type is skipped independently — if filters fail but a time granularity switcher works, the chart shows the granularity-adjusted data without filtering, and vice versa. A [field switcher](#field-switcher) that can't be applied says so [on the chart itself](#when-a-chart-cant-take-the-switch) rather than through the icons below.
The chart displays a warning icon to indicate the problem:
| Icon | Meaning |
| ------------------ | ---------------------------------------------------- |
| Crossed-out filter | Filters were skipped for this chart |
| Crossed-out clock | Time granularity override was skipped for this chart |
Hover over the icon for details. Click it to open the chart's [Controls mapping](#controls-mapping) and fix the issue — remap the control to a compatible dimension or disable it for this chart.
### Controls mapping
Each chart decides which controls apply to it through its **Controls mapping**. The mapping is resolved automatically in most cases and only needs manual attention when a control targets a member the chart doesn't have.
Open **Controls mapping** from a chart's settings menu to inspect or override the mapping for that chart. For each control on the dashboard you can:
* **Toggle the control on or off** for the chart, even when a mapping exists
* **Pick a different member** from the chart's semantic view to remap the control to
Three states show up in the mapping sidebar:
| Status | What it means |
| --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| **Mapped automatically** | The control's member exists on the chart's semantic view, so it's wired up without configuration. |
| **Manually mapped** | You (or an AI agent) picked a specific member for this chart. **Reset** restores the automatic mapping. |
| **Can't map automatically** | The control targets a member that doesn't exist on the chart's semantic view. The chart is unaffected by the control until you map it manually. |
Each control's picker offers only the members it can actually be resolved against, so a mapping you can stage is a mapping that will work:
| Control | What the picker offers |
| ------------------------------------------------------- | -------------------------------------------------------------------------------------------- |
| [Filter](#filter) | Dimensions on the chart's semantic view. |
| [Time granularity switcher](#time-granularity-switcher) | Time-typed dimensions only — other types can't be resolved by the time granularity pipeline. |
| [Field switcher](#field-switcher) set to **Dimension** | Dimensions. |
| [Field switcher](#field-switcher) set to **Measure** | Measures, since the control replaces a measure. |
Filter and time granularity mappings are also configurable by AI agents when they build or edit a dashboard, so an agent can wire those controls across charts that use different semantic views without you needing to revisit each chart manually. Field switchers are outside what agents author, so their mappings are yours to set.
[ref-embed-url-filters]: /embedding/iframe/dashboards#pre-set-dashboard-filters-via-url
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-charts]: /docs/explore-analyze/dashboards/widgets/charts
[ref-granularities]: /docs/data-modeling/dimensions
[ref-user-attributes]: /admin/users-and-permissions/user-attributes
# Widgets
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/widgets/index
Building blocks for dashboards — charts, text, controls, AI summaries, and layout elements that you arrange on the canvas to tell your data story.
Widgets are the building blocks of a dashboard. Each tile placed on the canvas in the [dashboard builder][ref-workbooks] is a widget — a chart, a block of text, a control that viewers interact with, an AI-generated summary, or a layout element such as a spacer, divider, or stack. Combine them to assemble polished, interactive views of your data.
## Widget types
The dashboard builder supports the following widget types:
* [Charts](/docs/explore-analyze/dashboards/widgets/charts) — Visualize reports from your workbook
* [Text](/docs/explore-analyze/dashboards/widgets/text) — Add titles, descriptions, and rich formatting in Markdown
* [Controls](/docs/explore-analyze/dashboards/widgets/controls) — Let viewers filter the data, switch the time granularity, swap which field the charts are built on, or drive several controls at once
* [AI summary](/docs/explore-analyze/dashboards/widgets/ai-summary) — Generate narrative summaries of dashboard data on demand
* [Spacer, Divider & Stack](/docs/explore-analyze/dashboards/widgets/layout) — Non-data layout elements for whitespace, section breaks, and grouping widgets
## Adding widgets
Add widgets from the toolbar at the top of the dashboard builder: pick reports from the **Charts** picker to add charts, use the **Add Widgets** menu for text, AI summaries, and layout elements, or the **Add Controls** menu for **Filter**, **Time Granularity**, **Field Switcher** and **Parent** controls.
Each item — a toolbar button, or an option inside the **Add Widgets** / **Add Controls** menus — can be added in two ways:
* **Click** it to drop the widget into the first open spot on the canvas.
* **Drag** it from the toolbar onto the canvas to place it exactly where you want. As you drag, a full-size placeholder previews the widget's footprint and the surrounding widgets reflow to open a slot; release to drop it there. Dragging is especially handy on dense dashboards, where clicking would otherwise place the new widget far down the page.
Starting a drag from a menu option closes the menu, so it doesn't cover the canvas while you place the widget.
## Arranging widgets
Drag any widget to move it, and use the handle in its bottom-right corner to resize it — the surrounding widgets shift to make room.
To work with several widgets at once, select them first:
* **Click** a widget to select it.
* **Shift-click** to add widgets to, or remove them from, the selection.
* **Drag a marquee** — press on any empty part of the canvas and drag a rectangle over the widgets you want; every widget it touches is selected.
With a selection in place:
* **Move the group** — drag any selected widget and the whole selection moves together, keeping its relative arrangement.
* **Delete the group** — press **Delete** or **Backspace** to remove every selected widget at once.
[ref-workbooks]: /docs/explore-analyze/workbooks
# Spacer, Divider & Stack
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/widgets/layout
Non-data layout elements — a spacer for whitespace, a divider line, and stack containers that group widgets — to help you structure a dashboard.
Spacers, dividers, and stacks are non-data **layout** widgets. They carry no data of their own; you place them on the canvas alongside charts, text, and controls to add whitespace, visual structure, and grouping to a dashboard.
## Spacer
A spacer is an empty, resizable box. Use it to add deliberate whitespace between widgets — for example, to separate a header row from the charts below it, or to push a widget into a particular column.
A spacer is only visible while you are editing in the [dashboard builder][ref-workbooks]. On the published dashboard (including [embedded views](/embedding/iframe/dashboards)) it renders as empty space, so it never draws a card or border for viewers.
## Divider
A divider is a horizontal separator line that breaks up the layout flow — a lightweight way to signal the boundary between sections of a dashboard. Unlike other widgets, a divider is a fixed height and **cannot be resized** vertically.
## Stack
A **horizontal stack** or **vertical stack** is a container that groups several widgets and lays them out evenly — side by side in a horizontal stack, or top to bottom in a vertical stack. Use a stack to keep a set of related widgets aligned and sized together, instead of positioning each one on the grid by hand.
Place a stack on the canvas and resize it to any width and height; its children always share that space evenly:
* In a **horizontal stack**, each child takes the full height and an equal share of the width.
* In a **vertical stack**, each child takes the full width and an equal share of the height.
Add a widget to a stack by dragging it in — from the canvas, the **Add Widgets** menu, or another stack — and remove one by dragging it back out; the remaining children redistribute to fill the space. Stacks can be nested inside one another.
Because a stack owns its children's layout, you work with the **container as a unit**: select, move, resize, or delete the stack itself and it carries its children with it. Individual widgets inside a stack are not separately selectable on the board — a marquee drawn over a stack selects the container, not its children. To rearrange children, drag one to a new position **within** the stack; to take one out, drag it onto the canvas or into another container.
## Adding a layout widget
In the [dashboard builder][ref-workbooks], open the **Add Widgets** menu in the toolbar and choose **Spacer**, **Divider**, **Horizontal Stack**, or **Vertical Stack**. The widget is added to the canvas, where you can drag it into place and resize it (a divider resizes horizontally only).
## Styling
The **spacer** and **divider** follow the dashboard's [widget styling settings](/docs/explore-analyze/dashboards/styling):
* The **divider** draws its line using the widget **border** width, style, and color.
* A **spacer** picks up the same border and background settings while you are editing, so its bounds are easy to see.
A **stack** draws no card of its own — it only groups and sizes its children, each of which keeps its own styling.
[ref-workbooks]: /docs/explore-analyze/workbooks
# Text
Source: https://docs.cube.dev/docs/explore-analyze/dashboards/widgets/text
Add titles, descriptions, and rich formatting to dashboards using Markdown.
Text widgets render Markdown content directly on the dashboard. Use them to add narrative, structure, and context around your charts.
## Adding a text widget
In the [dashboard builder][ref-workbooks], open the **Add Widgets** menu in the toolbar and choose **Text**. A new text widget is added to the canvas with an empty editor.
## Use cases
* Add a dashboard title or section headings to break long dashboards into scannable chunks
* Explain what stakeholders are looking at and how to interpret the numbers
* Link out to related dashboards, runbooks, or documentation
* Embed images, callouts, and formatted lists
## Supported syntax
Text widgets support standard Markdown — see the [Markdown Guide](https://www.markdownguide.org/basic-syntax/) for a full reference of supported syntax.
## Editing
Click the widget's settings menu and choose **Edit** to open the Markdown editor. The widget renders as you type so you can preview formatting before saving. Use **Hide Title** in the same menu to suppress the widget's title bar when the Markdown content already provides its own heading.
[ref-workbooks]: /docs/explore-analyze/workbooks
# Explore
Source: https://docs.cube.dev/docs/explore-analyze/explore
Self-serve data exploration from Analytics Chat, dashboards, or any semantic view — governed by your data model, shareable via URL.
Explore is a quick way to explore data in your semantic layer either by point and click or with an AI agent. Unlike workbooks, Explore doesn't require you to create a workbook—you can start exploring immediately from a dashboard, Analytics Chat, or any semantic view.
Explore leverages your data model definitions, ensuring that all queries use consistent metrics, respect access control policies, and benefit from pre-aggregations for fast performance.
## Using Explore
You can start exploring from three places in Cube:
### Exploring from Analytics Chat
When the AI agent returns a query result in Analytics Chat, click **Explore** to open that result in an exploration. This allows you to further analyze and visualize the data returned by the AI agent.
### Exploring from a published dashboard
Hover over a chart widget in a published dashboard, and click **Explore**. The exploration will use the same view and data that feeds into the visualization, giving you access to all measures and dimensions available in that view.
### Starting a new exploration from scratch
Navigate to the **Explore** page in the sidebar. From here, you can select any semantic view as the data source and start building visualizations from scratch.
## Explore functionality
The functionality in Explore is similar to when working with semantic views in workbooks. You can build visualizations, pivot tables, and tables by selecting measures and dimensions from your semantic views, apply filters, group and aggregate data.
Explore state is also saved in the URL, making it easy to share your exploration with other users by simply copying and sharing the link.
If you have developer or admin access, you can [apply a security context](/docs/explore-analyze/workbooks/querying-data#applying-a-security-context) to an exploration to verify what a specific end user—or an AI agent querying on their behalf—would see.
Explore results carry the same [freshness and pre-aggregation indicators](/docs/explore-analyze/workbooks/querying-data#result-freshness-and-provenance) as workbook reports.
## Saving explorations
You can save an exploration so you can return to it later without
re-building your query. An exploration is the same kind of object as a
workbook [report](/docs/explore-analyze/workbooks) — the only difference is
whether it belongs to a workbook. [Analytics Chat](/docs/explore-analyze/analytics-chat)
and the Cube MCP `createReport` tool save a standalone exploration the same
way: by leaving the workbook unset.
* **Save** — For an unsaved exploration, click **Save** to open the
**Save Exploration** dialog. Give your exploration a name and click
**Save**. The exploration is now persisted and appears in your
workspace.
* **Save (existing)** — After you modify a previously saved exploration,
click **Save** to update it in place.
* **Save as** — For a saved exploration, open the chevron menu next to
**Save** and choose **Save as...** to create a copy under a new name.
You can also rename a saved exploration at any time by clicking its title
in the header and typing a new name.
## Converting to a workbook
When you're ready to build a more permanent analysis, you can convert
your exploration to a [workbook](/docs/explore-analyze/workbooks). Click
**Convert to workbook** to create a new workbook with your exploration
as its first tab.
## Copying to clipboard
You can copy an exploration to your clipboard and paste it as a new tab
in any [workbook](/docs/explore-analyze/workbooks). Click the chevron
next to **Convert to workbook** and select **Copy to Clipboard**.
Then navigate to the target workbook and press **Cmd+V** (macOS) or
**Ctrl+V** (Windows/Linux) to paste the exploration as a new tab.
# Notifications
Source: https://docs.cube.dev/docs/explore-analyze/notifications
Notifications are available on Premium and Enterprise plans.
Notifications are available on [Premium and Enterprise plans](https://cube.dev/pricing).
Users need at least the [Explorer][ref-roles] role and Edit or Manage permission
on the workbook to set up notifications.
Notifications let you send email or Slack messages with a screenshot of a
dashboard after each [scheduled refresh][ref-scheduled-refreshes]. This is useful
for distributing regular updates to stakeholders without requiring them to log in.
They can optionally carry a short [AI-generated summary](#ai-summary) of what
changed since the previous notification.
## Creating a notification
Notifications are configured as part of a [scheduled refresh][ref-scheduled-refreshes].
When creating or editing a schedule, click **Add notification** to expand
the notification configuration card.
To send the same dashboard to another channel or recipient list, [duplicate an
existing schedule][ref-duplicate] instead of configuring the notification from
scratch — the copy carries over the source schedule's notification settings.
### Delivery channel
A schedule delivers to one channel, not both. Pick it with the **Send via**
radio buttons:
* **Email** — sends to the [recipients](#recipients) you choose: individual
workspace users, [user groups][ref-user-groups], or a mix of both.
* **Slack** — select a Slack channel to post to. Requires connecting your Slack
workspace first (one-time OAuth flow via **Connect to Slack**). Once
connected, select a channel from a searchable picker.
### Screenshot attachment
Choose the format for the dashboard screenshot attached to the notification:
* **PNG** (default)
* **PDF**
Select the format using the **Attach screenshot as** option. The same
formats are available for ad-hoc [downloads from the dashboard
header][ref-download].
### AI summary
Turn on **Include AI summary** to add a short "what changed" note to the body of
the email or Slack message, above the screenshot. It is off by default.
The note is written by the configured [AI agent][ref-agents] and is deliberately
short — an opening line naming the most important change, then two to four
bullets, each with a real number and its movement. It is meant to be read in a
few seconds to decide whether to open the dashboard at all:
```text theme={"dark"}
WHAT CHANGED
Order growth accelerated this week, driven mainly by a surge in returns
outpacing overall gains.
- Weekly orders: 416 last full week, up 20% from 346 the week before.
- Returned orders: 42 last week, up 40% week over week (from 30).
- Completed orders: 202 last week, up 13% week over week (from 179).
```
**What it compares against.** Where possible the note describes what changed
since the *previous notification* for that recipient, rather than what changed
inside the data — so a daily notification does not repeat the same sentence
every morning. On the first send, or when there is no earlier note to compare
with, it falls back to comparing recent periods (week over week, month over
month) and says so rather than inventing a comparison.
Each recipient's note is generated under **their own data access**, so it only
describes rows that recipient is allowed to see. That also means it costs one
agent run per recipient per send, which is why the option is off by default.
If the summary can't be generated — the agent is unavailable, or the run takes
too long — the notification is still delivered, without the note.
This is not the same as the [AI summary widget][ref-ai-widget], which lives on
the dashboard, takes a prompt you write, and is cached for everyone who views
it. The notification summary uses a fixed brief, is generated per recipient, and
is never shown on the dashboard.
### Removing a notification
To remove a notification from a scheduled refresh, click the X button in the
notification card header before saving.
## Email notifications
Email notifications are sent to the selected recipients after a scheduled
refresh completes. Each recipient receives a message with a link to the
dashboard and the attached screenshot.
Recipients don't have to be added by an editor. Anyone who can view the dashboard
can [subscribe themselves][ref-subscribe] to a schedule's email notifications
from the published dashboard.
### Recipients
Under the **Recipients** heading there are two separate controls:
* **Users** — a searchable picker for individual workspace users.
* **User groups** — a dropdown holding an expandable checklist of
[user groups][ref-user-groups] and their members. It appears only when your
workspace has at least one user group.
A group is expanded to its current members each time the notification is sent,
so adding or removing members takes effect on the next run without editing the
schedule. The number beside a group's name counts the members it can currently
email, so a group can show fewer than its full membership — anyone without an
email address is left out of both the count and the list.
A recipient who is selected individually and also belongs to a selected group is
emailed only once.
#### Excluding individual group members
Tick a group to notify everyone in it, then expand it and untick anyone who
should be skipped. Those people are stored as exceptions to that group on this
schedule — the group stays selected and keeps picking up new members, but the
excluded members are left out of every run.
One search box filters both levels at once, so typing a person's name narrows
the list to the groups that would notify them, with that person shown under
each.
Exceptions are saved with the rest of the form, not applied as you click. A few
details worth knowing:
* Unticking a group's **last** remaining member unticks the whole group.
* Exceptions belong to one schedule and one group. Excluding someone from a
group here does not affect the group anywhere else, or any other schedule.
* An exception is not a block on the person. If they are **also** selected under
**Users**, or have [subscribed themselves][ref-subscribe], they still get the
email — a direct recipient row and group membership are independent, and
either one alone delivers.
## Slack notifications
Slack notifications post a message with the dashboard screenshot to a single
Slack channel per schedule.
### Connecting Slack
Before you can send Slack notifications, connect your Slack workspace:
1. When configuring a notification, select **Slack** as the delivery
channel.
2. Click **Connect to Slack** to start the OAuth flow.
3. Authorize Cube to post to your Slack workspace.
This is a one-time setup. Once connected, you can select any channel your Slack
workspace has access to.
[ref-download]: /docs/explore-analyze/dashboards#download-as-png-or-pdf
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-scheduled-refreshes]: /docs/explore-analyze/scheduled-refreshes
[ref-duplicate]: /docs/explore-analyze/scheduled-refreshes#duplicating-a-schedule
[ref-user-groups]: /admin/users-and-permissions/user-groups
[ref-subscribe]: /docs/explore-analyze/scheduled-refreshes#subscribing-to-notifications
[ref-agents]: /admin/ai
[ref-ai-widget]: /docs/explore-analyze/dashboards/widgets/ai-summary
# Scheduled Refreshes
Source: https://docs.cube.dev/docs/explore-analyze/scheduled-refreshes
Scheduled refreshes are available on Premium and Enterprise plans.
Scheduled refreshes are available on [Premium and Enterprise plans](https://cube.dev/pricing).
Users need at least the [Explorer][ref-roles] role and Edit or Manage permission
on the workbook to set up scheduled refreshes. Anyone who can view the dashboard
can open the sidebar to see its schedules and [subscribe to
notifications](#subscribing-to-notifications).
Scheduled refreshes allow you to automatically refresh published dashboards on a
recurring schedule. When a scheduled refresh runs, the dashboard widgets are
queried in the background, refreshing Cube's in-memory cache. This warms up
dashboards so they load instantly when users open them, eliminating wait times.
You can also attach [notifications][ref-notifications] to scheduled refreshes to
send email or Slack messages with a dashboard screenshot after each refresh.
## Accessing scheduled refreshes
From the Dashboard Builder, click the calendar icon in the toolbar to open the
scheduled refreshes sidebar. From this sidebar, you can create new schedules,
modify existing ones, or remove schedules you no longer need.
The same calendar icon is also available on a published dashboard, so you can
reach the sidebar without opening the builder. What you can do there depends on
your permission on the workbook:
* **Edit** or **Manage** — the full sidebar: create, edit, run, and delete
schedules.
* **View only** — a read-only list of the configured schedules, where you can
[subscribe or unsubscribe](#subscribing-to-notifications) yourself from each
schedule's email notifications.
The dashboard must be published before schedules can be created.
## Creating a schedule
Click **New scheduled refresh** in the sidebar footer. A dialog opens
with the following configuration options.
### Schedule
Pick how often the refresh runs with the **Schedule** select:
| Option | Description |
| ------- | ----------------------------------------------------------------- |
| Hourly | Runs every hour at the specified minute |
| Daily | Runs once a day at the specified time |
| Weekly | Runs on a chosen day of the week at the specified time |
| Monthly | Runs on a chosen day of the month at the specified time |
| Custom | Accepts a cron expression (5-part: minute hour day month weekday) |
Time is specified in 12-hour format with an AM/PM selector. For Hourly
schedules, only the minute field is shown.
### Timezone
Select from common timezones including UTC, US timezones, London, Paris, Berlin,
Tokyo, Shanghai, Singapore, and Sydney. Defaults to your browser's timezone if
it matches a common one, otherwise UTC.
### Notifications
You can optionally attach [notifications][ref-notifications] to a scheduled
refresh. Click **Add notification** to expand the notification
configuration. See the [Notifications][ref-notifications] page for details.
Click **Save** to create or update the schedule.
## Managing schedules
The sidebar lists all existing schedules with the following information:
* **Status indicator** — current state: not run yet, last run succeeded, last
run failed, or currently running with phase details.
* **Schedule description** — human-readable text like "At 09:00 AM".
* **Enable/disable toggle** — turn schedules on or off without deleting them.
* **Next run time** — shows the next scheduled execution, or "Disabled" if the
schedule is off.
* **Notification channel** — when a [notification][ref-notifications] is
configured, the card shows its delivery channel, for example "Notification:
Email" or "Notification: Slack (#channel)".
### Actions
| Action | Description |
| --------- | ------------------------------------------------------------------------------------------------------ |
| Run now | Manually trigger the schedule immediately |
| Edit | Open the form dialog to modify the schedule |
| Duplicate | Create a new schedule pre-filled from this one (see [Duplicating a schedule](#duplicating-a-schedule)) |
| Delete | Remove the schedule permanently |
## Duplicating a schedule
To create a new schedule based on an existing one — for example, to send the
same dashboard to another Slack channel or a different set of recipients without
re-entering the **Schedule**, timezone, and notification settings — you can
duplicate it. There are two ways:
* **From the sidebar** — click the duplicate (copy) icon on a schedule card. The
form dialog opens pre-filled with its **Schedule**, timezone, and
notification configuration. Adjust anything you need, then click **Save** to
create the new schedule.
* **From the edit dialog** — while editing a schedule, click **Save as copy** in
the dialog footer. This creates a new schedule from the current form values,
including any unsaved changes, instead of updating the original.
Duplicates are independent schedules: the original is left unchanged, and the
copy keeps the source schedule's enabled or disabled state.
## Subscribing to notifications
If a schedule sends [email notifications][ref-notifications], anyone who can view
the dashboard can subscribe to it themselves — an editor does not have to add
them as a recipient. Open the scheduled refreshes sidebar from the published
dashboard and use the toggle on a schedule to subscribe or unsubscribe.
The toggle covers **every** way that schedule could reach you, in one click. If
you were added individually, unsubscribing removes you from its recipients. If
you are reached through a [user group][ref-user-groups] — or several — it also
records you as an exception to each of those groups on this schedule, so the
group keeps notifying everyone else but stops emailing you. Subscribing reverses
both: it clears those exceptions and adds you back as a recipient in your own
right. The change takes effect on the next run.
Unsubscribing affects only this schedule. It does not remove you from the user
group, and it does not change any other schedule that group receives.
The toggle appears only for schedules that send email — those with
notifications enabled and the email delivery channel. A schedule delivers to
either individual inboxes or a Slack channel, not both, so a schedule that posts
to Slack (or that has notifications turned off) has nothing to subscribe to. If an
editor switches a schedule to Slack, its individual subscriptions are cancelled
and any group exceptions it held are cleared; if it is later switched back to
email, everyone starts from the group's full membership again and can subscribe
or unsubscribe afresh.
Every notification email also includes a one-click **Unsubscribe** link in its
footer, so a recipient can stop receiving a schedule's emails without signing in
to Cube. It behaves exactly like the toggle, covering both a direct recipient row
and any [user group][ref-user-groups] delivering to you. This works for any
recipient, including those added directly by an editor and embed users who have
no Cube account of their own.
## Run phases
When a scheduled refresh runs, it progresses through these phases:
1. Initializing
2. Fetching credentials
3. Refreshing dashboard
4. Generating screenshot (if notifications are configured)
5. Generating AI summary (if the notification has [AI summary][ref-ai-summary]
turned on)
6. Sending notifications (if notifications are configured)
The sidebar shows real-time status updates during execution.
[ref-ai-summary]: /docs/explore-analyze/notifications#ai-summary
[ref-roles]: /admin/users-and-permissions/roles-and-permissions
[ref-notifications]: /docs/explore-analyze/notifications
[ref-user-groups]: /admin/users-and-permissions/user-groups
# Scheduled Tasks
Source: https://docs.cube.dev/docs/explore-analyze/scheduled-tasks
Save a natural-language agent prompt and have Cube run it automatically on a schedule, producing a chat thread for each run.
Scheduled Tasks let you run tasks on a schedule — or whenever you need them.
Save a natural-language prompt for the
[agent](/docs/explore-analyze/analytics-chat) and have Cube run it
automatically on a schedule or on demand. Each run produces a chat thread
you can open later to read the agent's answer. You can also ask the agent to create a
scheduled task for you from any of your chats.
For example, you might schedule a task to *"every weekday at 9am, summarize
yesterday's signups, flag anything anomalous, and email me the results"* and
review the summary each morning.
Scheduled Tasks are **scoped to a deployment** — each task belongs to the
deployment it was created in.
## Where to find it
In the deployment sidebar, open **Scheduled** (the clock icon, below
**Explore**). The page lists the deployment's tasks with their name,
schedule, status, and description, and is searchable.
## Anatomy of a task
A task has:
* **Name** (required) — how the task appears in the list.
* **Description** (optional) — a short summary shown in the list.
* **Instructions** (required) — the natural-language prompt the agent runs on
each execution.
* **Schedule** — how often the task runs (see below).
* **Enabled** — a toggle to activate or pause a scheduled task.
## Schedule options
Scheduled Tasks use the same schedule editor as
[dashboard scheduled refresh](/docs/explore-analyze/scheduled-refreshes),
offering these frequencies:
* **Manual** — no schedule; the task runs only when you trigger it with
**Run now**.
* **Hourly**, **Daily**, **Weekly** (with a day of the week), and **Monthly**
(with a day of the month) — each with a time picker.
* **Custom** — a raw cron expression for full control.
A **timezone** selector (defaulting to your browser's timezone) controls when
the schedule fires; it is stored per task.
## Managing tasks
From the list, you can:
* **Create** a new task. The **New task** button is a dropdown with two
options:
* **Create with agent** — opens a new Analytics Chat pre-seeded with a
message asking the agent to explain scheduled tasks and interview you
about what the task should do and when it should run. The agent then
creates the task for you (see
[Managing tasks from chat](#managing-tasks-from-chat)).
* **Set up manually** — opens the create dialog where you fill in the
task's details yourself.
* **Edit** an existing task's instructions, schedule, or details.
* **Enable / Pause** a scheduled task to control whether it runs on schedule.
* **Run now** — trigger a one-off run immediately. This works for both manual
and scheduled tasks. Triggering a run shows a notification with a **View**
link straight to the run's chat thread, and the thread appears in the
Recent Chats sidebar immediately.
* **Delete** a task. Deleting removes its schedule and stops all future runs.
## Task detail page
Clicking a task in the list opens its detail page. The header shows a
breadcrumb back to Scheduled Tasks, the task name, a status tag (**Manual**,
**Active**, or **Paused**), and the description, along with actions to
**Edit** (pencil), **Delete** (trash), and a primary **Run now** button.
The page shows:
* **History** — the task's runs, newest first. Each entry is a timestamped
link that opens that run's chat thread. Currently-executing runs
show a **Running** tag, and failed runs a **Failed** tag. The list shows
the latest 50 runs; a "Showing the latest 50 runs" note appears once the
cap is hit.
* **Instructions** — the task's prompt.
* **Repeats** — the schedule in plain language.
## Reading the output
Each run creates a chat thread containing the agent's response. Open the
thread in [Analytics Chat](/docs/explore-analyze/analytics-chat) to read the
full answer, ask follow-up questions, or
[save results to a Workbook](/docs/explore-analyze/workbooks).
Scheduled-run threads are marked in the Recent Chats sidebar with a clock
icon (hover over it to see the "Scheduled task" tooltip).
A run shows as **Running** while it executes, then completes or fails. A
failed run's thread shows a failure notice instead of an empty thread.
The task runs headlessly under the security context of the user who created
it, so it sees exactly the data that user can access.
### What the agent can do in a scheduled run
Beyond querying the semantic model, the agent in a scheduled run can:
* **Send email to workspace members** — for example, *"summarize yesterday's
signups and email the summary to me."* This requires the agent email tool
to be enabled for the workspace; recipients are restricted to workspace
members.
* **Use web search**.
* **Create and update reports, workbooks, and dashboards**, using the task
creator's permissions.
Scheduled runs do not yet have data-model access — editing the semantic
layer is currently available only in interactive chat.
## Managing tasks from chat
You can also create and manage Scheduled Tasks conversationally in
[Analytics Chat](/docs/explore-analyze/analytics-chat) — from any chat, not
just ones started with **Create with agent** (that menu option simply opens
a chat pre-seeded for this flow). Ask the agent to schedule, update, list,
or delete tasks in plain language — for example, *"schedule a daily summary
of yesterday's signups at 9am"* or *"list my scheduled tasks."*
The agent's actions render in the chat as labeled steps with a clock icon —
*"Creating scheduled task…"* / *"Created scheduled task"*, *"Listed
scheduled tasks"*, *"Updated scheduled task"*, and *"Deleted scheduled
task"*.
The agent manages task definitions: when listing tasks, it can report each
task's id, name, description, schedule, timezone, and enabled state. A
task's run history lives on its [detail page](#task-detail-page).
# Agent skills
Source: https://docs.cube.dev/docs/explore-analyze/skills
Run saved agent workflows using skill buttons, the slash menu, or automatic matching.
Skills are reusable, named workflows the agent can run on demand wherever you work with
it — [Analytics Chat](/docs/explore-analyze/analytics-chat), Workbooks, dashboards, and
the IDE. Instead of re-typing the same multi-step request every time, you pick a skill and
the agent follows the workflow your team has already defined — for example, a
**Weekly revenue report** that totals revenue, breaks it down by region, and calls out
week-over-week trends.
Skills are authored by your data team as part of the semantic model. This page covers how
to find and run them. To create skills, see [Skills](/admin/ai/skills) in the agent
configuration docs.
This page covers **Agent Skills in Cube**. These are not
[Cube agent skills](/docs/integrations/agent-skills), which run in a coding agent and
operate Cube through the CLI, or the [Cube connector](/docs/integrations/mcp-server),
which connects Claude and other MCP clients directly to Cube.
## Running a skill
There are three ways to run a skill from the chat input.
When skills are available, they appear as buttons in the chat empty state — one button
per skill, labeled with the skill's title. Click a button to start the skill.
Type `/` in the chat input to open the skill menu. The menu filters as you type and is
keyboard-navigable. When skills are available, the input placeholder reads
**"Type / for skills."**
You don't always have to pick a skill explicitly. If a free-text request matches a
skill's description, the agent recognizes the intent and runs the matching skill on
its own — asking "can I get the weekly revenue summary?" may trigger a weekly revenue
report skill.
## Adding context before you run
Picking a skill — whether from a button or the `/` menu — does **not** send the message
immediately. Instead, it inserts a styled `/skill-name` token into the chat input and
keeps your cursor there.
This lets you add specifics before running. Type any extra context after the token, then
send. On send, the message becomes the skill invocation plus whatever you added. For
example:
```text theme={"dark"}
/weekly-revenue-report for EMEA, last 6 weeks
```
runs the weekly revenue report skill, scoped to the EMEA region over the last six weeks.
Running skills is available to anyone with chat access, including **Explorer** and
**Viewer** [roles](/admin/users-and-permissions/roles-and-permissions). Authoring skills
requires data-model edit access — see [Skills](/admin/ai/skills) in the agent
configuration docs.
# Calculated fields
Source: https://docs.cube.dev/docs/explore-analyze/workbooks/calculated-fields
Create ad-hoc custom dimensions and measures with Semantic SQL in workbooks, with help from AI or the field picker.
Calculated fields are ad-hoc dimensions and measures you add only to the current
workbook report. They do not change the shared data model.
As described in [Semantic SQL](/docs/introduction#semantic-sql), Cube routes
analysis through the semantic layer instead of sending arbitrary SQL straight to
the warehouse. The runtime validates every request and applies your security
policies. Semantic SQL builds on Postgres-compatible SQL—including the
`MEASURE()` function—so you can express derived logic on top of existing
semantic definitions with both flexibility and governance.
Calculated fields are expressed as Semantic SQL and pushed down to the Cube
backend for evaluation. The semantic layer compiles them with the rest of the
query—rather than applying them only in the browser—so the same validation,
governance, and warehouse execution path apply as for any other Semantic SQL
analysis.
## Using AI to create calculated fields
You can ask the Cube AI agent to create custom calculations in natural language.
The agent can add or refine calculated fields from different parts of the
product—for example while exploring in **Analytics chat** or working in
**Workbooks**—so you are not limited to a single entry point when you want a new
metric or dimension for the analysis in front of you.
## Creating calculated fields in UI
You can also build and edit calculated fields directly in the workbook. New
fields appear in the **Query fields** section of the field picker sidebar.
### Aggregations from existing dimensions
Right-click a dimension column header and choose an aggregation to create a
calculated field automatically. Available aggregations depend on the column type:
| Column type | Available aggregations |
| --------------- | -------------------------------------- |
| Number | Count Distinct, Sum, Average, Min, Max |
| Time | Count Distinct, Min, Max |
| String, Boolean | Count Distinct |
### Calculations from existing measures
Open the menu on a measure column header and use the **Calculations** submenu
for derived calculations:
| Calculation | Description |
| ---------------------- | ------------------------------------------------------- |
| % of total | Ratio of the measure value to the total across all rows |
| % of previous | Ratio of the measure value to the previous row's value |
| % change from previous | Percentage change compared to the previous row |
| Running total | Cumulative sum of the measure across rows |
**% of previous**, **% change from previous**, and **Running total** require at
least one dimension in the query.
Which calculations are offered depends on the measure’s aggregation type:
| Aggregation type | Available calculations |
| ----------------------- | ---------------------- |
| Count, Sum | All calculations |
| Min, Max | Running total |
| Average, Count Distinct | None |
### Filtered measures
When working with query **Results**, pivot so at least one dimension is on
columns, then open the header menu on a **pivoted measure column** and choose
**Create filtered measure**. Cube adds a calculated measure that applies the
column’s slice—for example, from **Count** broken down by **Status**, you get a
measure that only aggregates rows matching that status (such as completed
orders only).
The option appears only for **native** measures on pivoted columns, not for
calculated fields. The same flow works in **Explore** when results are pivoted
the same way.
### Bins and value groups
You can also bucket an existing dimension without writing SQL. Open its menu in
the field picker sidebar and choose **Create bins…** on a number dimension, or
**Group values…** on a string or boolean one (grouping a boolean dimension is
how you rename its `true`/`false` values). Time dimensions have granularities
instead, and an already derived field cannot be bucketed again.
**Bins** take their boundaries either as a list (**Custom ranges**) or from a
**Range start** and **Range end** (**Equal intervals**), which are prefilled
from the column's own minimum and maximum. For equal intervals, choose whether
the range is split by **Number of bins** or by a fixed **Bin size**. Each
boundary opens a bucket that includes its lower bound and excludes the upper
one, and two open-ended buckets are added at the edges—so `0, 18, 25` yields
`< 0`, `[0, 18)`, `[18, 25)`, `>= 25`, and no row is dropped. **Label style**
renders a bucket as `[10, 20)`, `>= 10 and < 20`, `10 to 19` (offered only while
every boundary is a whole number), or **Custom**, which lets you type your own
label for each bucket. Turn off **Label empty values separately** to fold rows
where the dimension is `NULL` into the last bucket instead of reporting them
under their own label (`Unknown` by default).
**Value groups** collect the dimension's values into named sets: pick values, name
the group, and choose **Add group**. A value belongs to one group at a time, and
an existing group's picked values can be changed later via **Edit group**. By
default, whatever you did not pick—including empty values—falls under
**Everything else**, which defaults to `Other`; turn off **Group remaining
values** to have those rows return `NULL` instead.
Bucket labels carry their position as a prefix (`1.`, `2.`, zero-padded past nine
buckets) so that sorting the column sorts it by value rather than alphabetically,
which would put `>= 25` before `[0, 18)`. The prefix is visible in results, chart
legends, and axes.
The panel previews the Semantic SQL it generates as you build:
```sql theme={"dark"}
CASE WHEN orders_view.age IS NULL THEN 'Unknown'
WHEN orders_view.age < 0 THEN '1. < 0'
WHEN orders_view.age < 18 THEN '2. [0, 18)'
ELSE '3. >= 18' END
```
**Equal intervals** ranges are resolved into boundaries when the field is created, not
recomputed from the data. Values arriving later outside the range join the first
and last buckets instead of extending them.
To change a bucketed field, choose **Edit bins…** or **Edit groups…** from its
menu—either in the sidebar or on its column header in the results. Only fields
this panel generated offer the action; a `CASE` expression written by hand does
not.
### Editing a calculated field
Select a calculated field in the sidebar to open the editor. You can change its
**name** and **SQL expression**, then choose **Update** to apply.
# Workbooks
Source: https://docs.cube.dev/docs/explore-analyze/workbooks/index
Build reports and explore data with AI agents, organize analyses across multiple tabs, and share trusted insights with your team.
Workbooks allow you to build reports and explore data with AI agents, organize the results of your
analysis, and share trusted insights with your team.
A report inside a workbook and a standalone [exploration](/docs/explore-analyze/explore#saving-explorations)
are the same kind of object — the only difference is whether it's filed inside a workbook.
## Tabs
Workbooks can contain one or more tabs, each providing a different way to query and analyze your data. Tabs enable you to organize multiple analyses within a single workbook, making it easy to explore different aspects of your data or combine insights from different sources.
## New tab page
When you open [Explore](/docs/explore-analyze/explore) or add a new tab in a workbook, Cube shows
the **new tab page** — a single launchpad for starting a report. From here you can search across
your data model, browse starting points by type, or jump straight to pasting a query or asking an
AI agent.
### Searching and browsing
A search box sits at the top of the launchpad. Typing filters every category at once; categories
with no matches appear disabled, so you can tell at a glance which ones hold results.
The category selector lets you switch between starting-point types:
* **View groups** — browse [view groups](/docs/data-modeling/view-groups) as folders. Drilling
into a group reveals its sub-groups and views.
* **Views** — every [view](/docs/data-modeling/views) in the model. Search also matches a view's
measures and dimensions; when a view matches through a member, the row notes which one.
* **Workbooks** — drill into an existing workbook and pick one of its reports to reuse its query
as the starting point for the new report. The workbook you are currently in and empty tabs
(which have no query to reuse) appear disabled.
* **Source tables** — raw tables from your connected
[data sources](/admin/connect-to-data/data-sources), browsable by data source and schema.
Click through the list to drill down to an individual view, report, or table. Once you reach a
starting point, the data model sidebar activates so you can continue building your query there.
### Starting from a pasted query or an agent
The launchpad also offers shortcuts at the bottom:
* **Paste semantic SQL query** — opens the SQL panel so you can paste a semantic SQL query. The
dropdown adds **Paste source SQL query** (raw data-source SQL), **Paste JSON query** (a
[REST API query](/reference/core-data-apis/rest-api/query-format)), and **Paste GraphQL query** (a
[GraphQL API query](/reference/core-data-apis/graphql-api)) — the JSON and GraphQL queries are
converted to a semantic query for you.
* **Ask Cube agent** — opens the chat sidebar so you can describe the report you want in natural
language.
### Configuring the launchpad
Administrators can tailor the launchpad per deployment under **Settings → Configuration**, in the
**Workbooks and Explore** section. Two settings are available.
**Available tabs** — choose which categories appear on the launchpad. **View groups** and **Views**
are always shown and cannot be turned off; **Workbooks** and **Source tables** can each be enabled
or disabled to simplify the starting points your users see.
**Default view group** — set a default view group so the launchpad opens inside a specific
top-level view group instead of listing all of them, useful for steering users toward a curated set
of governed starting points.
Go to **Settings → Configuration** for your deployment.
In the **Workbooks and Explore** section, select the tabs to show and, optionally, enter the
**name** of the view group the launchpad should open inside. Leave the group empty to show all
view groups. Save your changes.
When a default view group is set, new Explore and workbook tabs open inside that group. You can
still navigate up to **All view groups** to browse everything. If the configured group no longer
exists, the launchpad falls back to showing all categories.
## Tab types
Workbooks support two types of tabs, each designed for different querying approaches:
### Semantic Query
[Semantic Query tabs](/docs/explore-analyze/workbooks/querying-data) allow you to query your data model through the semantic layer. This ensures all queries use consistent metrics, respect access control policies, and benefit from pre-aggregations for fast performance.
You can query the semantic layer in three ways:
* **Point and click** – Use the intuitive interface to select measures and dimensions, build visualizations, and explore your data interactively
* **AI agent** – Ask the AI agent to build queries and visualizations based on natural language questions
* **Semantic SQL** – Write Semantic SQL queries manually to have full control over your queries while still leveraging the semantic layer
### Source SQL Query
[Source SQL Query tabs](/docs/explore-analyze/workbooks/source-sql-tabs) allow you to query your connected data sources directly, bypassing the semantic layer. This gives you direct access to your raw data and enables deeper exploration beyond what's defined in your data model.
You can query source SQL in two ways:
* **Write SQL manually** – Write standard SQL queries directly against your data source for maximum flexibility
* **Ask AI** – Use the AI agent to help you research your data source, build SQL queries, and refine results in real-time
## Copying and pasting tabs
You can copy a tab from one workbook and paste it into another workbook
using the clipboard.
### Copying a tab
Open the tab options menu (the chevron on the tab in the footer) and
select **Copy to Clipboard**. The tab's query, chart configuration, and
settings are serialized and written to your system clipboard.
### Pasting a tab
Navigate to the target workbook and press **Cmd+V** (macOS) or
**Ctrl+V** (Windows/Linux) while the workbook is focused. A new tab will
be created with the query and configuration from the clipboard.
Paste only works via the keyboard shortcut. Make sure the focus is not
inside a text input—otherwise the browser's default paste behavior will
take precedence.
You can also copy an exploration from [Explore](/docs/explore-analyze/explore)
and paste it as a new tab in any workbook using the same workflow.
## Duplicating workbooks
You can duplicate a workbook by selecting **Duplicate** from the row
actions menu on the workspace page, from the workbook menu while the
workbook is open, or from a published dashboard's options menu. This
creates a full copy of the workbook, including all its tabs, reports, and
any published dashboard.
A duplicate does not carry the original's sharing over by default. The
copy is visible only to you, plus anyone with access to the folder it is
created in.
If the original is shared — with people, with groups, with your whole
organization, or through signed embedding — the duplicate dialog offers
**Copy sharing and embedding settings**. Selecting it gives the copy the
same audience. The option appears only if you have **Full access** to the
original, since copying its sharing means granting that access again. See
[Duplicates][ref-sharing-duplicates] for what does and doesn't carry over.
## Workbook versions
When you [publish a dashboard][ref-dashboards] from a workbook, Cube
automatically creates a version snapshot. Each time you publish, a new
version is recorded so you can track changes to your dashboard over time.
### Restoring a version
You can restore the state of your workbook from any previously published
dashboard version. Restoring a version rebuilds the workbook with the
charts and configuration from that published snapshot.
Restoring a version replaces **all existing tabs** in the workbook with
the tabs from the selected version. Any current work in the workbook that
has not been published will be lost.
[ref-dashboards]: /docs/explore-analyze/dashboards
[ref-sharing-duplicates]: /docs/organize-content/sharing#duplicates
# Python analysis
Source: https://docs.cube.dev/docs/explore-analyze/workbooks/python-analysis
Attach a Python script to a workbook report or exploration to run forecasting, regression, cohort, and other analysis that SQL can't express.
Python analysis is currently in preview, and the user experience and the script
contract may still change. Reach out to the [Cube support
team](/admin/account-billing/support) to activate this feature for your account.
A workbook report or exploration can carry an attached **Python script** that
transforms its SQL result. Its chart then renders the script's **output** instead
of the raw SQL rows. This turns an analysis that would otherwise scroll away in a
chat transcript into a saved, re-runnable, shareable analysis.
Use it for work SQL can't express — forecasting, regression, cohort analysis,
statistical tests, clustering, and anomaly detection.
## Adding Python to a workbook report or exploration
### From Analytics Chat
Ask for the analysis in natural language — "forecast next quarter's revenue",
"find anomalies in signups" — and the agent runs Python for you, rendering the
result inline in the [chat thread](/docs/explore-analyze/analytics-chat). This
result is ephemeral by default.
Ask to **save it** — to a workbook, or as a standalone
[exploration](/docs/explore-analyze/explore#saving-explorations) — and Cube persists
both the code and the run result. Opening the saved copy renders that output without
re-running; saving re-executes the analysis.
The agent reaches for Python **only** when the answer genuinely needs a statistics
or machine-learning library. Ordinary aggregations, top-N, ratios, running totals,
period-over-period comparisons, and time series all stay in SQL, because a Python
run costs a re-query plus a sandbox start. If you expected Python and got a plain
SQL query, that is usually correct behavior.
### From the toolbar
Workbooks and [Explore](/docs/explore-analyze/explore) share the same flow.
Click **Python** in the toolbar to open the Python panel, then **Add script** to
attach one. Cube seeds a starter script and opens it on the **Script** tab.
Opening the panel does not attach anything by itself — only **Add script** does.
**Remove**, in the panel header, detaches the script, after which the analysis is
SQL-backed again.
Attaching Python clears any existing SQL result: a Python-backed analysis renders
its last Python run, and a freshly attached script has none until you press **Run**.
Attaching Python in Explore is only available on a **saved** exploration. On an
unsaved one the **Python** button is disabled with the tooltip *"Save the
exploration to add Python"* — **Run** executes server-persisted code, so the
analysis needs a saved exploration to live on.
## Writing the script
The script runs in a sandbox against a fixed contract:
* Input data arrives as `data.csv` in the working directory.
* Write results to `output.json` as a **flat JSON array of row objects**, for
example `[{"month": "2026-01", "value": 1.5}, ...]`.
A top-level dict or object is rejected — flatten any nested structure into one
array of uniform rows.
## Python environment
Every run gets a fresh, isolated sandbox running **Python 3.11**. It is created for
the run and destroyed when the run finishes — nothing carries over between runs.
These packages are pre-installed, along with their dependencies:
| Package | Use |
| --------------------------------- | --------------------------------------------------------- |
| `pandas`, `numpy` | Dataframes and numerical computing |
| `scipy` | Statistical tests, optimization, interpolation |
| `scikit-learn` | Regression, classification, clustering, anomaly detection |
| `statsmodels` | ARIMA, exponential smoothing, econometric models |
| `prophet` | Time series forecasting with seasonality and holidays |
| `matplotlib`, `seaborn`, `plotly` | Plotting |
Figures are not a supported output. The chart is built from `output.json`, and
anything a script writes to disk is discarded with the sandbox — so the plotting
libraries are importable, but a saved figure has nowhere to go.
Installing your own packages is not supported yet. Because the sandbox is recreated
for every run, anything a script installs is discarded when the run ends. Support for
adding packages to the environment is coming.
## Running and refreshing
The panel's **Script** tab is editable in place, with line numbers. **Reset**
restores the starter template.
* **Edits do not run anything.** They save with the analysis, and the rendered result
keeps showing the previous run.
* When the code or its input SQL has changed since the last run, the result is
marked **Outdated**, with the tooltip *"The Python code or its input SQL changed
after the last run. Run to refresh the saved result."*
* **Run** executes the stored script in the sandbox and persists the refreshed
result.
* The **Output** tab shows what the last run printed — the script's stdout and
stderr, so `print()` is how you inspect intermediate values. Both are captured up
to the cap in [Limits](#limits), so a chatty script gets truncated.
* The **input SQL panel is read-only** on a Python-backed analysis: that SQL is the
sandbox's input, not what gets charted. It still offers the **Semantic SQL** and
**Generated SQL** tabs, both derived from that input query.
* **A failed run keeps the previous chart.** The error surfaces alongside the last
successful result, which stays rendered.
**Run is the only way the saved result changes.** Opening the workbook report or
exploration, reloading the page, or viewing a dashboard never re-runs anything on
its own.
## On dashboards
Python-backed workbook reports render their **saved output** on dashboards.
Nothing re-runs on dashboard load, so a dashboard full of Python reports costs
no compute to open — each widget shows whatever the last **Run** produced.
A python widget can be opened in [Explore](/docs/explore-analyze/explore) from a
dashboard and run from there.
## Who the analysis runs as
Python runs with the **security context of the person who pressed Run** — or of the
chat user who saved the analysis. The result is then persisted with the workbook
report or exploration, and **anyone who can view that item can see the result**.
Row-level security is applied at **run time**, not at view time. A user with broad
access can Run, and the stored output is then readable by people whose own access
is narrower.
Take this into account when deciding who can run Python analyses or publish Python
reports, the same way you would for any other shared saved result.
## Limits
| Limit | Value |
| ----------------------------- | ------ |
| SQL query timeout | 120s |
| Python execution timeout | 120s |
| Maximum output rows persisted | 10,000 |
| Maximum output size | 2 MB |
| stdout/stderr captured | 16 KB |
Exceeding the output caps means the analysis still returns in chat but **cannot be
saved to a workbook report or exploration**. Aggregate or summarize inside the
script so the output stays within the caps — analysis results such as forecasts,
cohorts, and test statistics are small by nature.
## Learn more
* [Analytics Chat](/docs/explore-analyze/analytics-chat) — the standalone
conversational analytics experience
* [Workbook Agent](/docs/explore-analyze/workbooks/workbook-agent) — the authoring
assistant inside a workbook
* [Source SQL tabs](/docs/explore-analyze/workbooks/source-sql-tabs) — query
connected data sources directly
# Querying data
Source: https://docs.cube.dev/docs/explore-analyze/workbooks/querying-data
In the Semantic Query tab, you can query the semantic layer. On the left, there is a selector for the semantic view that you want to query.
On the left side, you can select a semantic view you would like to query. Inside the semantic view, you can see the measures and dimensions grouped either by cubes or folders. You can search within your semantic view to find the relevant member.
## Running queries
As you build a query — adding members, changing filters, adjusting sorting —
the workbook runs it automatically and refreshes the results. You can turn
this off with the auto-run toggle in the query controls: with auto-run
disabled, the results are marked as outdated while you edit the query, and
the query only runs when you click **Run query**. The toggle applies to the
current tab only.
A data model author can change the default for a semantic view with the
[`meta.auto_run`][ref-auto-run] parameter — selecting a view with
`auto_run: false` starts the tab with auto-run disabled. This is useful for
views that are expensive to query. You can still toggle auto-run back on at
any time.
An admin can also turn auto-run off by default for every tab in the
deployment, from **Model Configuration**. Per-view and per-tab settings still
take precedence over this deployment-wide default. See
[Auto-run][ref-configuration-auto-run] in the Model Configuration reference.
When a cell's value is too long for its column, it is truncated with an
ellipsis. Hover over a truncated cell to see its full value in a tooltip.
## Filtering
Dimensions and measures can be added as filters to focus on specific rows of
data. Pick **Filter** in a member's menu in the left pane, or in a result
column's menu, to add the first one. The filter bar then appears above the
results, with a **Filter** button for adding more — you can filter on any member
of the semantic view, whether or not the query selects it.
Every filter is a chip showing its member, operator and value. Click the chip
to edit it, or use its options menu to remove it. A filter added without a
value yet is highlighted and left out of the query; the query runs as soon as
the value is complete.
Filtering a dimension keeps or drops rows before aggregation, as a SQL `WHERE`
condition. Filtering a measure applies after aggregation, as a `HAVING`
condition — so `revenue greater than 1000` keeps groups whose total exceeds
1000, not individual orders.
### Operators
The operators offered depend on the member's data type, for measures as much as
for dimensions:
| Data type | Operators |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| String | is, is not, contains, not contains, starts with, not starts with, ends with, not ends with, is null, is not null, is empty, is not empty |
| Number | is, is not, greater than, greater than or equal, less than, less than or equal, between, is null, is not null |
| Time | is, is not, after, after or on, before, before or on, between, in the month, not in the month, in the quarter, not in the quarter, in the year, not in the year, is null, is not null |
| Boolean | is, is not, is null, is not null |
On a number member the comparisons appear on the chip as `=`, `≠`, `>`, `≥`,
`<` and `≤`.
### Values
* **is** and **is not** accept several values, matching any of them. On a string
member the value picker lists the member's own values and searches them
server-side, so you select from what the data actually contains.
* The **contains**, **starts with** and **ends with** family matches a
substring, prefix or suffix, ignoring case.
* **between** includes both bounds.
* **is null**, **is not null**, **is empty** and **is not empty** take no value.
Empty means the empty string, which is not the same as null.
* The period operators take any date inside the period — **in the year** with
`2026-03-14` matches all of 2026.
* A date value can be fixed or relative. The date editor's **Relative** tab
takes values like `today`, `7 days ago`, `this month` or `2 weeks from now`,
and resolves them every time the query runs, so the window rolls forward on
its own. The full list is under
[`default_ui_filters`][ref-default-ui-filters].
### Combining filters
Filters in the bar are combined with AND: every one of them must hold. For OR,
and for anything nested, use **Add advanced filter** in the filter picker's
footer. It creates a group, shown as its own chip, whose operator sits as a
toggle between the filters inside it — click it to switch the whole group
between AND and OR. Groups can contain groups, which is how an expression like
`status is completed AND (country is US OR city contains San)` is built.
A group holds either dimension filters or measure filters, never both, because
the two apply at different stages of the query. Its **Filter** button therefore
offers only the kind of member the group already contains.
### Custom SQL filters
**Add custom SQL**, also in the filter picker's footer, takes a SQL condition
and passes it into the query as written. Use it for a condition the operators
above cannot express.
### Default filters
A data model author can define default workbook filters on a semantic view with
the [`meta.default_ui_filters`][ref-default-ui-filters] parameter. The same
defaults are also seeded when a query on that view is started in the Google
Sheets / Excel add-in.
```yaml theme={"dark"}
views:
- name: orders_view
cubes:
- join_path: orders
includes: "*"
meta:
default_ui_filters:
- member: created_at
operator: between
values: ["7 days ago", "today"]
- member: status
operator: is
values: ["completed"]
```
When you select such a view in a new tab, those filters are added to the query
automatically. They behave exactly like filters you add yourself: you can
change their values, switch operators, or remove them.
Date range filters can use relative bounds like `7 days ago` or `this month`.
A relative range resolves each time the query runs, so the window rolls
forward day over day, and the filter's date editor opens on the **Relative**
tab showing its bounds. See
[`default_ui_filters`][ref-default-ui-filters] for the full list of accepted
operators and values.
Resetting the tab or going back to the view list clears default filters
along with the rest of the query; they are seeded again the next time you
select a view.
## Pivoting
You can add dimensions to the pivot by right-clicking on the left pane or in the results header.
Note that chart tables are different from results tables. You can configure pivots in chart tables separately.
## Totals
The totals menu (the **Σ** button) in the results toolbar adds summary rows
and columns to the results table: **Column Totals** (a pinned row at the
bottom with a total for each measure column), **Row Totals** (a total for
each row across the pivoted columns), and **Subtotals** (see below). All
three can be enabled at once; each setting is saved with the report and
carries over when you switch the chart type to [Table][ref-table-chart].
### Subtotals
When the table is [pivoted](#pivoting) by two or more dimensions,
**Subtotals** appends a bold **Total for ‹value›** column after each pivot
group's columns, at every nesting level — for example, pivoting by year and
then status adds a **Total for 2024** column after 2024's status columns.
With fewer than two pivot dimensions the menu item is disabled, and if the
query changes so the pivot no longer qualifies, subtotals switch off
automatically.
Subtotal values are computed by the database — one extra query per pivot
level — rather than by summing the visible cells, so they are correct for
non-additive measures such as counts of distinct values and averages, and
when the [row limit](#limiting) truncates the table. With **Column Totals**
also enabled, the pinned totals row shows a grand total for each pivot group.
* The toggle adds subtotals for all pivot levels at once; there is no
per-group control.
* Subtotals are not available for flat (non-pivoted) tables or in the
[values-as-rows table layout](/docs/explore-analyze/charts/chart-types/table#pivot-presets).
* [Calculations][ref-calculated-fields] based on window functions, such as
**Running total**, are excluded — same as for row and column totals.
## Limiting
Every query in a workbook returns at most a fixed number of rows. By
default, queries are limited to **5,000 rows**, and you can adjust the
limit up to **50,000 rows** from the query controls.
The default and maximum values available in the workbook are governed by
the deployment's [Model Configuration][ref-configuration]:
* **Default row limit** is applied when a query doesn't specify one
explicitly.
* **Maximum row limit** is the upper bound — any value you pick in the
workbook is capped at this number.
Updating those settings at the deployment level changes what every
workbook in the deployment can request.
Large result sets can slow down chart rendering and may hit browser
memory limits, especially for tables. When possible, aggregate in the
semantic layer instead of pulling raw rows into the workbook.
## Sorting
You can sort query results using drop-down menus on column headers in the results table or the
dedicated sorting control.
### Column headers
Hover over a column header in the results table and expand the context menu to sort ascending, sort descending, or clear sorting.
If a column has sorting applied, a chevron icon on the header indicates the current sorting direction.
### Sorting control
The sorting control, available via the **Sort** button, lists all query members and their current state:
unsorted, sorted ascending, or sorted descending. Use it to apply or
change the sort order for any member. Drag and drop members within the sorting control to change their priority in the sort order.
When a [pivot](#pivoting) is applied, an additional section with pivot dimensions appears.
## Filling in missing rows
A time-series query only returns rows for date buckets that actually contain
data. If no orders were created in a given month, that month is simply absent
from the result — which reads as a misleading gap in tables and charts. Fill
in missing rows densifies the result so every date bucket in the range is
present.
To apply it, hover over the header of a time dimension column in the results
table, expand the context menu, and select **Fill in missing rows**. To turn
it off, select **Stop filling missing rows** from the same menu. While the
fill is active, an icon on the column header indicates it.
The fill works on time dimensions with a year, quarter, month, week, or day
granularity. The date range comes from the date filters on that dimension, or
from the earliest and latest dates in the data. In added rows, additive
measures such as counts and sums are set to `0`; other measures are left
empty.
The fill doesn't change the query or the generated SQL — the missing rows are
added when the report is rendered. It is saved with the report and applies
everywhere it appears, including [dashboards][ref-dashboards].
The fill is skipped when it would produce more than 10,000 rows. When that
happens, the column header indicator warns you.
## Period-over-period comparison
You can compare any measure against a prior period — for example, month over
month or year over year. Hover over the header of a measure column in the
results table, expand the context menu, and pick a comparison under
**Compare periods**. The same menu on a time dimension column header applies
the comparison to all measures at once.
The menu is available when the query includes a time dimension with a year,
quarter, month, week, or day granularity. The presets depend on that
granularity — a month grain offers **Previous period**, **Same period
previous quarter**, and **Same period previous year** — so the comparison
period always aligns exactly with the current buckets. A measure can have
several comparisons active at once, such as month over month and year over
year side by side.
Each comparison adds derived columns next to the measure. Choose which ones
to show under **Comparison columns** in the same menu:
* **Previous value** (on by default) — the measure's value in the comparison
period
* **Difference** — the absolute change from the comparison period
* **% change** — the relative change from the comparison period
To remove a single derived column, use **Remove Column** in its header menu;
removing the last one turns the comparison off. You can also untoggle the
preset under **Compare periods**.
Comparisons are saved with the report and apply everywhere it appears —
charts, totals, and [dashboards][ref-dashboards]. Cube fetches the
comparison values with a companion query whose date filters are shifted back
by the comparison offset; you can [inspect it](#inspecting-queries) to check
the math.
## Inspecting queries
Every query in the Semantic Query tab can be inspected as code. Open the SQL panel via the **SQL** button in the toolbar, next to the Results and Chart tabs, to see the query that the workbook generates, switch between representations, and copy it for use elsewhere.
The panel exposes two top-level tabs:
* **Semantic SQL** — the query expressed against the Semantic Model (semantic views, measures, dimensions, and `MEASURE()` calls). This is the source of truth for the query and what the workbook stores.
* **Generated SQL** — the SQL that Cube compiles from the Semantic SQL and sends to your data warehouse. Use it to understand exactly what runs against your database, debug performance, or share with a database administrator.
The **Semantic SQL** tab also has a dropdown for viewing the same query in the formats accepted by Cube's [Core Data APIs][ref-core-data-apis]:
* **Semantic JSON** — the [REST (JSON) API][ref-rest-api] query format
* **Semantic GraphQL** — the [GraphQL API][ref-graphql-api] query format
These views are convenient when you want to take a query you've built interactively in a workbook and reuse it directly through one of the Core Data APIs from your own application.
## Result freshness and provenance
Two indicators next to each result describe the data behind it.
* **Freshness** — a leaf icon shows how recently the underlying data was refreshed. Hover over it to see the last refresh time (for example, "Refreshed 5 minutes ago"); its color shifts as the data ages, so you can tell fresh from stale results at a glance. If the refresh time can't be determined, the leaf turns gray with a "Data age unknown" label. When you have access to [Query History](/admin/monitoring/query-history), clicking the leaf opens the underlying request for that result.
* **Pre-aggregation** — for those same users, a lightning-bolt icon appears next to the leaf with a "Served from a pre-aggregation" label whenever the result was served from a [pre-aggregation](/docs/pre-aggregations). If no icon is shown, the query ran directly against your data source.
## Editing Semantic SQL by hand
You can edit the Semantic SQL directly in the editor to refine a query — for example, to add a `WHERE` clause, change `GROUP BY` order, or apply a different sort. After making changes, click **Save and Run** to apply them to the query and refresh the results, or **Discard** to revert.
When you save a hand-edited query, the workbook keeps the Semantic SQL as the source of truth, and the visual controls (members panel, sort, pivot) stay in sync with what's expressible in the UI. Anything that goes beyond the interactive controls — see [Advanced semantic queries](#advanced-semantic-queries) below — disables the affected UI controls but keeps the rest of the workbook fully functional.
## Advanced semantic queries
Advanced semantic queries typically can contain CTEs (Common Table Expressions) and unions, and can contain multiple semantic views. In many cases for advanced semantic queries, the interactive components of the UI would be disabled—not the whole UI—as these complex queries require manual SQL editing.
Currently, advanced semantic queries can be placed on dashboards, but dashboard filters cannot be applied to advanced semantic queries.
## Applying a security context
When you run a query—in a workbook, on a dashboard, or in [Explore][ref-explore]—Cube
generates a token to execute it against the semantic layer. If you have developer or
admin access, you can attach a [security context][ref-security-context] to that token
directly from the page you're working on. The values you provide are used when the
token is generated, overriding the claims that would otherwise come from your session.
The security context you set is also respected by [AI agents][ref-analytics-chat]:
any query an agent runs on your behalf—whether in the workbook or in the IDE—executes
under the same context. This lets you confirm an agent sees exactly the data a given
end user would.
### What you can override
* **Groups** — choose **Inherit** to keep the Cube Cloud groups resolved from your
session, **Custom** to query as a specific set of groups, or **None** to query with
no groups at all.
* **User attributes** — keep the attributes resolved from your session, or override
them with your own values.
Setting a security context requires developer or admin access (deployment edit
access). Overriding the Cube Cloud **groups** and **user attributes** specifically
requires the admin role. Users without edit access always query with the security
context derived from their own session.
You can't use a security context to impersonate another user. The `email` claim is
never overridable—it's always taken from your own session—so you can't run queries on
behalf of someone else.
This is primarily a testing tool. It lets you reproduce what a specific end user would
see and verify [access control][ref-access-control] and multitenancy behavior without
leaving Cube Cloud. It's especially useful when a query fails or returns no data
because the data model expects security context values that aren't present in your
session—you can supply them here and confirm the model (and any agent that queries it)
behaves as intended.
[ref-core-data-apis]: /reference/core-data-apis
[ref-rest-api]: /reference/core-data-apis/rest-api
[ref-graphql-api]: /reference/core-data-apis/graphql-api
[ref-configuration]: /docs/data-modeling/configuration
[ref-security-context]: /docs/data-modeling/access-control/context
[ref-access-control]: /docs/data-modeling/access-control
[ref-dashboards]: /docs/explore-analyze/dashboards
[ref-explore]: /docs/explore-analyze/explore
[ref-analytics-chat]: /docs/explore-analyze/analytics-chat
[ref-default-ui-filters]: /reference/data-modeling/view#default_ui_filters
[ref-table-chart]: /docs/explore-analyze/charts/chart-types/table
[ref-calculated-fields]: /docs/explore-analyze/workbooks/calculated-fields
[ref-auto-run]: /reference/data-modeling/view#auto_run
[ref-configuration-auto-run]: /docs/data-modeling/configuration#auto-run
# Source SQL Tabs
Source: https://docs.cube.dev/docs/explore-analyze/workbooks/source-sql-tabs
Run warehouse-native SQL inside Workbooks—by hand or with AI help—to explore raw tables before promoting logic into the semantic model.
Source SQL Query tabs allow you to query your connected data sources directly, bypassing the semantic layer. Source SQL Query tabs can be used for prototyping your calculations before they would go into the semantic layer.
## Data sources
The data sources sidebar lists every source declared on the deployment —
both the default source and any named sources added under
[**Settings → Data Sources**](/admin/connect-to-data/multiple-data-sources).
Newly added sources show up immediately; you don't have to wire them into
a cube first.
## Querying methods
You can query source SQL in two ways:
* **Write SQL manually** – Write standard SQL queries directly against your data source for maximum flexibility
* **Ask AI** – Use the AI agent to help you research your data source, build SQL queries, and refine results in real-time
# Workbook Agent
Source: https://docs.cube.dev/docs/explore-analyze/workbooks/workbook-agent
Use the AI agent inside a workbook to build reports, create and edit dashboards, and run analysis with natural language.
Inside a [workbook][ref-workbooks], the AI agent is a full **authoring**
assistant. Where the read-only [Dashboard Agent][ref-dashboard-agent] only
answers questions about a published dashboard, the workbook agent can act on
your behalf — it builds reports, and assembles and edits dashboards, all from
plain-language requests.
**Workbook Agent vs. [Analytics Chat][ref-analytics-chat].** Analytics Chat is
a standalone, chat-first analytics experience for asking questions and saving
results into a workbook. The Workbook Agent lives **inside a workbook** and is
oriented around authoring the workbook's own reports and dashboards. Both run
read-only queries against your semantic model; the Workbook Agent additionally
creates and edits workbook content.
## What the Workbook Agent can do
In addition to everything the read-only [Dashboard Agent][ref-dashboard-agent]
can do (answer questions, run read-only semantic-layer queries, explain metric
and dimension definitions), the Workbook Agent can:
* **Create and save reports** in the workbook from a natural-language request
* **Search existing reports** to find and reuse work already in the workbook
* **Build and edit dashboards** — add and configure chart, KPI, filter,
time-grain, and text widgets
* **Publish dashboards**
## How charts get added to a dashboard
When the agent adds a chart to a dashboard, the chart widget references a
**report in the workbook**, not a one-off query. So to put a new chart on a
dashboard, the agent first saves the query as a report in the workbook, then
adds a widget that points at that report. This keeps every dashboard chart
backed by a reusable, governed report.
## Getting help with Cube
While working in a workbook, you can also ask the agent questions about Cube
itself — it can answer using the Cube documentation, which is handy for
"how do I…" questions without leaving your analysis.
## Limitations
* **Report search covers published reports only.** When the agent searches for
existing reports to reuse, it sees published reports.
## Learn more
* [Dashboard Agent][ref-dashboard-agent] — the read-only Q\&A assistant on
published and embedded dashboards
* [Querying data in workbooks][ref-querying-data] — point-and-click, AI agent,
and Semantic SQL
* [Analytics Chat][ref-analytics-chat] — the standalone conversational
analytics experience
[ref-workbooks]: /docs/explore-analyze/workbooks
[ref-dashboard-agent]: /docs/explore-analyze/dashboards/dashboard-agent
[ref-analytics-chat]: /docs/explore-analyze/analytics-chat
[ref-querying-data]: /docs/explore-analyze/workbooks/querying-data
# Connect your data
Source: https://docs.cube.dev/docs/getting-started/connect-your-data
If you don't have a data source to connect to, you can create a demo deployment to test all Cube features with demo data.
If you don't have a data source to connect to, you can create a demo deployment to test all Cube features with demo data.
Cube supports multiple warehouses and databases.
Once you select your data source, follow the instructions on the connection page.
## AI builds initial semantic model
Select what tables you'd like to use in your initial semantic model and let AI build it.
# Create workbooks and dashboards
Source: https://docs.cube.dev/docs/getting-started/create-workbooks-and-dashboards
Workbooks are places to curate and organize your analysis. You can create multiple reports using tabs, each focusing on different aspects of your data.
## Create a new workbook
You can create a new workbook in two ways:
* Click the **New Workbook** button on the home page or on the workbooks page
* Open **Explore** from Analytics Chat, then convert that exploration into a workbook
## Organize analysis with tabs
Use tabs within a workbook to create different reports. Each tab can contain its own analysis, allowing you to explore various aspects of your data and organize multiple insights in one place. The AI agent can help you build analysis in workbook tabs, creating queries and visualizations based on your questions.
## Share with dashboards
Once you're ready to share your work, you can turn workbooks into [dashboards][ref-dashboards] to organize reports into shareable artifacts. Dashboards let you select and organize reports from your workbooks into polished views for your team and stakeholders. The AI agent can help you create dashboards, manage layout, add filters, and more.
Dashboards can be published to make them accessible to your team. Published dashboards provide stakeholders with direct access to the insights that matter most.
[ref-dashboards]: /docs/explore-analyze/dashboards
# Develop in IDE
Source: https://docs.cube.dev/docs/getting-started/develop-in-ide
Use Cube Cloud IDE branch environments, commits, merges, and AI assistance to iterate on your data model without destabilizing production.
The IDE is a place to develop your data model. Cube supports multiple environments based on Git branches, helping you follow software engineering best practices while developing your data model.
## Branch-based development
Cube uses Git branches to create separate environments for development and production. This allows you to safely make changes to your data model without affecting your production environment. When you need to make a change, you can enter development mode or work on a separate branch to test and refine your changes. This isolation ensures that your production data model remains stable while you experiment and iterate.
## Commit and push to production
When your changes are ready, you can commit and push them to your production branch. The IDE helps manage this entire process without requiring you to know how Git and version control work in detail. You can save changes, commit them, and merge into production—all from within the IDE interface.
## AI-assisted development
The IDE includes an AI agent that helps you create data models from scratch and make changes to existing data models. Simply describe what you want to build or modify, and the AI agent will assist you in implementing the changes.
# Embed analytics
Source: https://docs.cube.dev/docs/getting-started/embed-analytics
Ship agentic embedded analytics in your product — choose among iframe embeds, the conversational Chat API, or headless core APIs for full control.
Cube offers rich options for embedded analytics. You can embed [dashboards][ref-dashboards] and [analytics chat][ref-analytics-chat] as iframes, use the [Chat API][ref-chat-api] directly to create your own conversational analytics experience, or use [Core Data APIs][ref-core-apis] directly to build custom visualization, reporting, and dashboarding experiences.
## Embed with iframes
The simplest way to embed Cube content is using iframes. You can embed both dashboards and analytics chat directly into your applications. Cube supports two authentication methods for iframe embedding: [private embedding][ref-private-embedding] for internal use cases and [signed embedding][ref-signed-embedding] for customer-facing applications.
## Use the Chat API
For more control over the conversational analytics experience, you can use the [Analytics Chat API][ref-chat-api] directly. This API-first approach lets you programmatically integrate AI-powered conversations with your data into your applications, giving you full control over the user experience.
## Build custom experiences
If you want complete control over visualizations and user interfaces, you can use [Cube's core APIs][ref-core-apis]—including REST (JSON), GraphQL, and SQL APIs—directly. This headless approach enables you to build fully custom visualization, reporting, and dashboarding experiences tailored to your specific needs.
[ref-dashboards]: /docs/explore-analyze/dashboards
[ref-analytics-chat]: /docs/explore-analyze/analytics-chat
[ref-chat-api]: /reference/embed-apis/chat-api
[ref-core-apis]: /reference/core-data-apis
[ref-private-embedding]: /embedding/iframe/auth/private
[ref-signed-embedding]: /embedding/iframe/auth/signed
# Import a Bitbucket repository
Source: https://docs.cube.dev/docs/getting-started/migrate-from-core/import-bitbucket-repository-via-ssh
Onboard an existing Cube project by linking a Bitbucket repository over SSH in Cube Cloud and completing database connectivity.
This guide walks you through setting up Cube Cloud, importing a
[Bitbucket][bitbucket] repository with an existing Cube project via SSH, and
connecting to your database.
## Step 1: Create an account
Navigate to [cubecloud.dev](https://cubecloud.dev/), and create a new Cube Cloud
account.
## Step 2: Create a new Deployment
Click **Create Deployment**. This is the first step in the deployment
creation. Give it a name and select the cloud provider and region of your
choice.