<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <title>Data Today: Matillion</title>
  <subtitle>Field notes for teams building on the Matillion Data Productivity Cloud.</subtitle>
  <link href="https://data-today.net/matillion/feed.xml" rel="self" />
  <link href="https://data-today.net/" />
  <updated>2026-06-29T00:00:00Z</updated>
  <id>https://data-today.net/</id>
  <author>
    <name>Data Today Newsroom</name>
  </author>
  <entry>
    <title>Matillion Maia BigQuery support: setup, cost, limits</title>
    <link href="https://data-today.net/matillion/matillion-maia-bigquery-support/" />
    <updated>2026-06-29T00:00:00Z</updated>
    <id>https://data-today.net/matillion/matillion-maia-bigquery-support/</id>
    <content type="html">&lt;p&gt;Matillion just made Maia much harder for Google Cloud teams to ignore. The June update adds &lt;strong&gt;Matillion Maia BigQuery support&lt;/strong&gt;, which means a team building in the Data Productivity Cloud can now create Maia projects against Google BigQuery instead of treating Snowflake, Databricks, or Redshift as the default center of gravity.&lt;/p&gt;
&lt;p&gt;The practical finding is simple: Matillion says Google BigQuery is supported on the &lt;strong&gt;Current runner track now&lt;/strong&gt;, with Stable track availability expected around &lt;strong&gt;August 1, 2026&lt;/strong&gt;, and the first cut is already broad enough to matter: 29 transformation components, 40 orchestration components, and 122 connectors listed for BigQuery projects in the Matillion changelog. The feature was surfaced in Matillion&#39;s June 26 New Features Blog, while the underlying Maia changelog records the BigQuery platform support under June 23, 2026.&lt;/p&gt;
&lt;p&gt;That gap between Current and Stable is the whole story for operators. You can start testing now if your runner policy allows Current. You should hold production rollout if your shop standardizes on Stable, especially if your pipelines influence finance reporting or customer facing SLAs.&lt;/p&gt;
&lt;h2 id=&quot;what-actually-changed-for-bigquery-teams&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-maia-bigquery-support/#what-actually-changed-for-bigquery-teams&quot;&gt;&lt;span&gt;What actually changed for BigQuery teams?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Matillion added Google BigQuery as a supported cloud data warehouse in Maia, which means BigQuery now appears as a selectable data platform when you create a build project. The &lt;a href=&quot;https://docs.maia.ai/docs/guides/bigquery-projects&quot;&gt;Google BigQuery projects documentation&lt;/a&gt; says the project setup flow includes a Data platform drop-down where you select Google BigQuery, then choose either Maia managed or Advanced settings depending on how you want infrastructure and secrets handled.&lt;/p&gt;
&lt;p&gt;The component surface is larger than a cautious preview would suggest. Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/changelog/2026-changelog&quot;&gt;2026 Maia changelog&lt;/a&gt; lists &lt;strong&gt;29 transformation components&lt;/strong&gt;, &lt;strong&gt;40 orchestration components&lt;/strong&gt;, and &lt;strong&gt;122 connectors&lt;/strong&gt; available for Google BigQuery data pipelines, with minimum Maia runner version &lt;strong&gt;11.478.0&lt;/strong&gt;.&lt;/p&gt;
&lt;figure class=&quot;figure&quot;&gt;&lt;img src=&quot;https://data-today.net/posts/matillion-maia-bigquery-support-fig-bigquery-component-coverage.png&quot; alt=&quot;Bar chart for Matillion Maia BigQuery support showing 29 transformation components, 40 orchestration components, and 122 connectors.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;figcaption&gt;Google BigQuery support covers 29 transformation components, 40 orchestration components, and 122 connectors. Source: Matillion Maia 2026 changelog. Data Today benchmark.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The chart above is the important sanity check. BigQuery support is useful because the initial component list includes production staples such as Table Input, Table Output, SQL, Join, Window Calculation, Run Transformation, SQL Script, dbt Core, Cloud Pub/Sub, Query Result to Grid, and Google Cloud Storage Unload. That is enough to build a real warehouse workflow, not just a demo that loads a CSV and waves at governance.&lt;/p&gt;
&lt;p&gt;The release track matters more than the launch blog tone. Matillion says BigQuery is available on the &lt;strong&gt;Current&lt;/strong&gt; runner track from the changelog date and that Stable users should expect availability around &lt;strong&gt;August 1, 2026&lt;/strong&gt; in the same &lt;a href=&quot;https://docs.maia.ai/docs/changelog/2026-changelog&quot;&gt;changelog entry&lt;/a&gt;. In calendar terms, that is roughly 39 days between the Current availability date of June 23, 2026 and the expected Stable date of August 1, 2026.&lt;/p&gt;
&lt;p&gt;Here is the operational read:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Choice&lt;/th&gt;
&lt;th&gt;What Matillion documents&lt;/th&gt;
&lt;th&gt;Use it when&lt;/th&gt;
&lt;th&gt;Cost and risk shape&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Current runner track&lt;/td&gt;
&lt;td&gt;BigQuery support is available now&lt;/td&gt;
&lt;td&gt;You can test new platform support before the Stable window&lt;/td&gt;
&lt;td&gt;Faster validation, higher change risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stable runner track&lt;/td&gt;
&lt;td&gt;BigQuery expected around August 1, 2026&lt;/td&gt;
&lt;td&gt;You need conservative runner updates&lt;/td&gt;
&lt;td&gt;Slower rollout, lower surprise risk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Minimum runner version&lt;/td&gt;
&lt;td&gt;BigQuery component support starts at 11.478.0&lt;/td&gt;
&lt;td&gt;You are checking a Hybrid SaaS runner or release policy&lt;/td&gt;
&lt;td&gt;Upgrade work may sit on your team&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The other items in the June feature bundle, bulk variable management and Designer command palette improvements, are useful background. They are not the lead event. The lead event is that Maia can now build and manage workflows against a warehouse many GCP shops already use as their system of record.&lt;/p&gt;
&lt;h2 id=&quot;how-do-you-configure-a-bigquery-project-without-making-security-angry&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-maia-bigquery-support/#how-do-you-configure-a-bigquery-project-without-making-security-angry&quot;&gt;&lt;span&gt;How do you configure a BigQuery project without making security angry?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;BigQuery authentication is the first non-obvious piece. Most Matillion warehouse environments put warehouse authentication directly on the environment. BigQuery uses Google Cloud credentials instead, and Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/guides/full-saas-bigquery&quot;&gt;Full SaaS BigQuery setup guide&lt;/a&gt; says a Full SaaS BigQuery environment must have access to a Google Cloud service account credential supplied as a configured cloud credential on the environment.&lt;/p&gt;
&lt;p&gt;For Full SaaS, your setup checklist is short but sharp:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Create a Google Cloud project.&lt;/li&gt;
&lt;li&gt;Create or choose a BigQuery dataset.&lt;/li&gt;
&lt;li&gt;Create a Google Cloud service account.&lt;/li&gt;
&lt;li&gt;Assign IAM roles that let Maia run BigQuery jobs and touch the required datasets.&lt;/li&gt;
&lt;li&gt;Add the service account key as a cloud credential in the Matillion project.&lt;/li&gt;
&lt;li&gt;Create a BigQuery environment with a default GCP Project ID and Dataset.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Matillion&#39;s setup page includes an example Google Cloud service account key structure. Keep the real key out of Git, tickets, prompt history, and copied pipeline notes. A safe shape for the credential looks like this:&lt;/p&gt;
&lt;pre class=&quot;language-json&quot; tabindex=&quot;0&quot;&gt;&lt;code class=&quot;language-json&quot;&gt;&lt;span class=&quot;token punctuation&quot;&gt;{&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;type&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;service_account&quot;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;project_id&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;example-project&quot;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;private_key_id&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;redacted&quot;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;private_key&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;-----BEGIN PRIVATE KEY-----&#92;nREDACTED&#92;n-----END PRIVATE KEY-----&#92;n&quot;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;client_email&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;matillion-sa@example-project.iam.gserviceaccount.com&quot;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;token_uri&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;https://oauth2.googleapis.com/token&quot;&lt;/span&gt;
&lt;span class=&quot;token punctuation&quot;&gt;}&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;For Hybrid SaaS, the better GCP native pattern is runner-assigned credentials when you can use it. Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/guides/hybrid-saas-bigquery&quot;&gt;Hybrid SaaS BigQuery guide&lt;/a&gt; says that when the environment has no associated cloud credential, the Maia runner uses Application Default Credentials, with the Google Cloud service account attached to the runner acting as the principal.&lt;/p&gt;
&lt;p&gt;That gives you three credential patterns worth comparing:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment pattern&lt;/th&gt;
&lt;th&gt;Credential source&lt;/th&gt;
&lt;th&gt;What to watch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Full SaaS&lt;/td&gt;
&lt;td&gt;Environment cloud credential backed by service account key JSON&lt;/td&gt;
&lt;td&gt;Key rotation and allowlisted Matillion access&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hybrid SaaS on GCP with runner credentials&lt;/td&gt;
&lt;td&gt;Runner service account through Application Default Credentials&lt;/td&gt;
&lt;td&gt;Runner IAM scope, GKE deployment hygiene, least privilege&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hybrid SaaS with environment credential&lt;/td&gt;
&lt;td&gt;Environment-associated service account key&lt;/td&gt;
&lt;td&gt;More explicit per-environment identity, more key management&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The cleanest production version is usually Hybrid SaaS on GCP with runner-assigned credentials, because it avoids distributing service account keys. The catch is operational: your team owns the runner, its GKE deployment, and its IAM posture. Full SaaS is quicker for teams that want Matillion to handle more of the platform layer, but you still own the Google credential blast radius.&lt;/p&gt;
&lt;h2 id=&quot;which-iam-roles-do-you-really-need&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-maia-bigquery-support/#which-iam-roles-do-you-really-need&quot;&gt;&lt;span&gt;Which IAM roles do you really need?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Matillion documents the minimum shape rather than a single magic role. The service account needs permission to create BigQuery jobs, query tables and views, retrieve metadata, list projects and datasets, insert or load data, and create, update, or delete tables and views when your pipelines do that work.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&quot;https://docs.maia.ai/docs/guides/hybrid-saas-bigquery&quot;&gt;Hybrid SaaS setup guide&lt;/a&gt; says you must grant either &lt;code&gt;roles/bigquery.jobUser&lt;/code&gt; or &lt;code&gt;roles/bigquery.user&lt;/code&gt; at minimum because both include &lt;code&gt;bigquery.jobs.create&lt;/code&gt;, and it lists &lt;code&gt;roles/bigquery.dataEditor&lt;/code&gt; for read and write workflows plus &lt;code&gt;roles/bigquery.dataViewer&lt;/code&gt; for read access.&lt;/p&gt;
&lt;p&gt;A least-privilege starting point for a transformation project that reads and writes a single dataset looks like this in intent:&lt;/p&gt;
&lt;pre class=&quot;language-bash&quot; tabindex=&quot;0&quot;&gt;&lt;code class=&quot;language-bash&quot;&gt;gcloud projects add-iam-policy-binding example-project &lt;span class=&quot;token punctuation&quot;&gt;&#92;&lt;/span&gt;
  &lt;span class=&quot;token parameter variable&quot;&gt;--member&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;token string&quot;&gt;&quot;serviceAccount:matillion-sa@example-project.iam.gserviceaccount.com&quot;&lt;/span&gt; &lt;span class=&quot;token punctuation&quot;&gt;&#92;&lt;/span&gt;
  &lt;span class=&quot;token parameter variable&quot;&gt;--role&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;=&lt;/span&gt;&lt;span class=&quot;token string&quot;&gt;&quot;roles/bigquery.jobUser&quot;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then scope dataset access separately in BigQuery for the dataset Maia should read or write. If a pipeline stages data in Google Cloud Storage before loading to BigQuery, Matillion&#39;s setup guide lists common Storage roles: &lt;code&gt;roles/storage.objectViewer&lt;/code&gt; for reading staged files, &lt;code&gt;roles/storage.objectCreator&lt;/code&gt; for uploads, and &lt;code&gt;roles/storage.objectAdmin&lt;/code&gt; for full object access.&lt;/p&gt;
&lt;p&gt;The mistake to avoid is granting &lt;code&gt;roles/bigquery.admin&lt;/code&gt; because the first setup test failed at 5:30 p.m. That role appears in Matillion&#39;s documented list as full administrative access, but broad admin rights turn an AI-assisted pipeline builder into a very well credentialed foot-gun. Start narrow, run the pipeline, read the failure, then add the missing permission deliberately.&lt;/p&gt;
&lt;h2 id=&quot;what-does-it-cost-when-maia-works-against-bigquery&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-maia-bigquery-support/#what-does-it-cost-when-maia-works-against-bigquery&quot;&gt;&lt;span&gt;What does it cost when Maia works against BigQuery?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;There are two bills to separate: Matillion consumption and Google Cloud consumption. BigQuery support does not make SQL free. Maia can help build or modify the workflow, but BigQuery still charges for the warehouse work it performs under your Google Cloud account.&lt;/p&gt;
&lt;p&gt;Matillion exposes consumption through its REST API. The &lt;a href=&quot;https://docs.maia.ai/docs/api-reference/maia-api-overview&quot;&gt;Maia API overview&lt;/a&gt; says the API base URLs are &lt;code&gt;https://eu1.api.matillion.com/dpc&lt;/code&gt; for EU accounts and &lt;code&gt;https://us1.api.matillion.com/dpc&lt;/code&gt; for US accounts, and it documents a &lt;strong&gt;1000 requests per minute&lt;/strong&gt; fixed API rate limit that returns HTTP &lt;code&gt;429 Too Many Requests&lt;/code&gt; when exceeded.&lt;/p&gt;
&lt;p&gt;A FinOps check for Maia pipeline usage starts with the consumption endpoint, then you join that against BigQuery job history in your own GCP project:&lt;/p&gt;
&lt;pre class=&quot;language-bash&quot; tabindex=&quot;0&quot;&gt;&lt;code class=&quot;language-bash&quot;&gt;&lt;span class=&quot;token function&quot;&gt;curl&lt;/span&gt; &lt;span class=&quot;token parameter variable&quot;&gt;--request&lt;/span&gt; GET &lt;span class=&quot;token punctuation&quot;&gt;&#92;&lt;/span&gt;
  &lt;span class=&quot;token parameter variable&quot;&gt;--url&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;https://us1.api.matillion.com/dpc/v1/consumption&quot;&lt;/span&gt; &lt;span class=&quot;token punctuation&quot;&gt;&#92;&lt;/span&gt;
  &lt;span class=&quot;token parameter variable&quot;&gt;--header&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;Authorization: Bearer &lt;span class=&quot;token variable&quot;&gt;$MATILLION_TOKEN&lt;/span&gt;&quot;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Matillion&#39;s 2026 changelog shows recent work on the consumption schema: &lt;code&gt;GET /v1/consumption&lt;/code&gt; added &lt;code&gt;pipelineName&lt;/code&gt; in January, changed &lt;code&gt;results.consumption.credits&lt;/code&gt; to a number in April, and added an &lt;code&gt;environment&lt;/code&gt; field for orchestration and transformation consumption in May. That matters because BigQuery cost control depends on attribution. A monthly credit total without pipeline and environment is a blame smoothie.&lt;/p&gt;
&lt;p&gt;What I would track for the first 30 days:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Matillion credits by pipelineName and environment&lt;/strong&gt;, pulled from &lt;code&gt;GET /v1/consumption&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;BigQuery bytes processed by job&lt;/strong&gt;, pulled from Google Cloud&#39;s job metadata or billing export.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Runner track and version&lt;/strong&gt;, especially if you are testing on Current before Stable.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Connector choice&lt;/strong&gt;, because a full load connector and an incremental load connector can have very different downstream BigQuery costs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The cheap mistake is letting Maia generate a correct but expensive pattern: full refresh, wide table scan, write-back to a large table, repeat hourly. Ask Maia for the pipeline, then inspect the generated SQL and component choices like you would review a pull request from a very fast junior engineer.&lt;/p&gt;
&lt;p&gt;For teams already watching Matillion spend, the same discipline applies here as in our guide to &lt;a href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/&quot;&gt;Context Engine setup and costs&lt;/a&gt;: ground the AI feature in observable consumption before you put it behind a schedule.&lt;/p&gt;
&lt;h2 id=&quot;when-should-you-use-maia-for-bigquery-and-when-should-you-wait&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-maia-bigquery-support/#when-should-you-use-maia-for-bigquery-and-when-should-you-wait&quot;&gt;&lt;span&gt;When should you use Maia for BigQuery, and when should you wait?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Use it now for migration proofs, greenfield GCP projects, connector evaluation, and internal productivity tests. A data engineer can ask Maia to build the first pass of a BigQuery transformation pipeline, wire in Table Input, Filter, Join, SQL, and Table Output, then review the generated shape in Designer before committing.&lt;/p&gt;
&lt;p&gt;A good first prompt is specific about the warehouse defaults and the cost guardrails:&lt;/p&gt;
&lt;pre class=&quot;language-text&quot; tabindex=&quot;0&quot;&gt;&lt;code class=&quot;language-text&quot;&gt;Build a BigQuery transformation pipeline in the dev environment.
Read analytics.orders and analytics.customers.
Filter orders to the last 7 days.
Join on customer_id.
Write the result to mart.recent_customer_orders.
Avoid full scans where a partition filter is available.&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Wait for Stable if the pipeline is regulated, high volume, customer facing, or already fragile. The Current track is the right place to learn how Maia behaves with your BigQuery datasets. Stable is the better place to run the job that wakes up your CFO if the numbers drift.&lt;/p&gt;
&lt;p&gt;Also wait if your IAM model is still mushy. BigQuery support puts Maia closer to a powerful production surface: jobs, datasets, tables, views, GCS staging, and metadata. That is exactly where service account discipline matters.&lt;/p&gt;
&lt;p&gt;The strongest use case is a GCP team that already has BigQuery, has a small data engineering group, and needs more pipelines than the team can hand build. Maia can remove some canvas work and boilerplate SQL. It cannot remove your obligation to validate lineage, permissions, costs, and data contracts.&lt;/p&gt;
&lt;h2 id=&quot;is-this-the-bigquery-moment-for-maia&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-maia-bigquery-support/#is-this-the-bigquery-moment-for-maia&quot;&gt;&lt;span&gt;Is this the BigQuery moment for Maia?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Yes, for GCP shops that were waiting for Maia to meet them where the data already lives.&lt;/p&gt;
&lt;p&gt;The bigger point is less glamorous: BigQuery support moves Maia from an interesting AI layer to an operational option for another major warehouse camp. The release is broad on components, early on runner track, and very dependent on IAM hygiene. Treat it like a new production platform surface, not a magic button. The builders who win with it will be the ones who let Maia draft the work and make their review process boring.&lt;/p&gt;
&lt;h2 id=&quot;sources&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-maia-bigquery-support/#sources&quot;&gt;&lt;span&gt;Sources&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/changelog/2026-changelog&quot;&gt;Matillion Maia changelog: 2026 changelog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/full-saas-bigquery&quot;&gt;Matillion Maia documentation: Setup guide, Matillion Full SaaS Google BigQuery&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/hybrid-saas-bigquery&quot;&gt;Matillion Maia documentation: Setup guide, Hybrid SaaS Google BigQuery&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/bigquery-projects&quot;&gt;Matillion Maia documentation: Google BigQuery projects&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/bigquery-environments&quot;&gt;Matillion Maia documentation: Google BigQuery environments&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/api-reference/maia-api-overview&quot;&gt;Matillion Maia documentation: Overview of the Maia API&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
  <entry>
    <title>Matillion anomaly alerts: runtime drift without polling</title>
    <link href="https://data-today.net/matillion/matillion-anomaly-alerts/" />
    <updated>2026-06-22T00:00:00Z</updated>
    <id>https://data-today.net/matillion/matillion-anomaly-alerts/</id>
    <content type="html">&lt;p&gt;The annoying pipelines rarely fail first. They usually get weird first: a 12 minute load turns into 19 minutes, an iterator step crawls through one customer after another, or a transformation finishes suspiciously fast because it processed almost nothing. That gray zone is where &lt;strong&gt;Matillion anomaly alerts&lt;/strong&gt; now matter.&lt;/p&gt;
&lt;p&gt;Matillion has added duration anomaly notifications to Maia, its AI layer for the Matillion Data Productivity Cloud. The important number is &lt;strong&gt;10 successful runs&lt;/strong&gt;: Matillion says a pipeline needs at least 10 successful runs before anomaly alerts can be generated, and Maia can use up to 300 successful runs to build the duration baseline in its &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;pipeline notifications documentation&lt;/a&gt;. This is a practical observability feature, not magic failure prediction. It gives you an operational warning when runtime drifts from the pipeline&#39;s own history.&lt;/p&gt;
&lt;p&gt;That distinction matters if you influence the bill. Duration is one of the cleanest early signals that something changed: source volume, warehouse queueing, connector behavior, branch changes, environment configuration, or a runner problem. An alert that fires while a slow run is still executing can save compute before the end state shows up as a red failure, a missed SLA, or a larger credit line item.&lt;/p&gt;
&lt;h2 id=&quot;what-exactly-changed-in-maia-pipeline-observability&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-anomaly-alerts/#what-exactly-changed-in-maia-pipeline-observability&quot;&gt;&lt;span&gt;What exactly changed in Maia pipeline observability?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Matillion&#39;s 2026 changelog says you can now subscribe to anomaly alert notifications through &lt;strong&gt;Email, Slack, or Webhook&lt;/strong&gt; when a pipeline run&#39;s duration deviates significantly from historical baselines, with the feature listed under Pipelines, Observability, and a minimum Maia version of all supported versions in the &lt;a href=&quot;https://docs.maia.ai/docs/changelog/2026-changelog&quot;&gt;June 17, 2026 changelog entry&lt;/a&gt;. The public roadmap post on June 19 bundled this with iterator visibility and GitLab API support, but the duration alert is the operationally heavier change because it turns pipeline drift into a push signal.&lt;/p&gt;
&lt;p&gt;The earlier Maia observability flow already had anomaly indicators in Pipeline Runs: red arrows beside duration metrics, a hover tooltip, and an anomalies section on the run detail. Matillion&#39;s troubleshooting guide gives the example of a run being &lt;strong&gt;7.8 percent longer than expected&lt;/strong&gt; in its &lt;a href=&quot;https://docs.maia.ai/docs/guides/maia-troubleshooting&quot;&gt;Maia AI Agents troubleshooting documentation&lt;/a&gt;. The new subscription path is the difference between noticing a red arrow when you happen to check the dashboard and getting paged into the right Slack channel before the run quietly burns through the morning.&lt;/p&gt;
&lt;p&gt;The feature only applies to scheduled or API-driven pipeline runs. Matillion says notifications are sent for scheduled or API-driven runs, while pipelines run manually through Designer do not generate notifications in the &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;Pipeline notifications guide&lt;/a&gt;. That is the right boundary. Manual runs are often experiments. Scheduled and API-triggered runs are production behavior.&lt;/p&gt;
&lt;p&gt;Here is the baseline math Matillion documents, with the two numbers you should remember before you file a ticket about missing alerts.&lt;/p&gt;
&lt;figure class=&quot;figure&quot;&gt;&lt;img src=&quot;https://data-today.net/posts/matillion-anomaly-alerts-fig-baseline-window.png&quot; alt=&quot;Bar chart for Matillion anomaly alerts showing 10 successful runs required before alerts and up to 300 successful runs used for baseline history.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;figcaption&gt;Source: Matillion pipeline notifications documentation. Maia needs 10 successful runs before anomaly alerts can be generated and uses up to 300 successful runs to build the duration baseline. Data Today benchmark.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The chart above is blunt on purpose: &lt;strong&gt;10 successful runs&lt;/strong&gt; are the floor, and &lt;strong&gt;300 successful runs&lt;/strong&gt; are the maximum recent history Maia uses for the baseline. If you deploy a new production pipeline and schedule it hourly, anomaly alerts cannot reasonably help you on the first day. If you schedule it once per day, give it at least 10 successful days before you expect the alerting model to settle.&lt;/p&gt;
&lt;h2 id=&quot;how-does-maia-decide-that-a-runtime-is-abnormal&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-anomaly-alerts/#how-does-maia-decide-that-a-runtime-is-abnormal&quot;&gt;&lt;span&gt;How does Maia decide that a runtime is abnormal?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Maia compares each pipeline against its own recent duration history. Matillion says the expected range is built from successful runs of that pipeline, centered on the median duration, with spread measured toward the 99th percentile on the slow side and the 1st percentile on the fast side in the &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;anomaly alert calculation notes&lt;/a&gt;. That is a useful design choice because a pipeline that usually varies between 8 and 40 minutes should not alert the same way as a pipeline that reliably finishes between 58 and 62 seconds.&lt;/p&gt;
&lt;p&gt;The slow side is the one most teams will care about first. Matillion says a slow anomaly can be detected while a pipeline is still running, while a fast anomaly is detected only at completion in the &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;duration anomalies section&lt;/a&gt;. That means Maia can warn you before the run ends when the elapsed runtime crosses the expected upper bound. For cost control, that timing is the feature.&lt;/p&gt;
&lt;p&gt;Fast anomalies deserve more respect than they usually get. A pipeline finishing in one tenth of its normal time can mean a source delivered zero rows, a branch skipped the real work, an environment variable pointed to a smaller schema, or an upstream extract silently changed. Matillion says the lower bound is the median minus twice the downward spread and never drops below zero, which means many pipelines will never trigger a fast anomaly unless their normal runtime is long and consistent in the same &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;expected range explanation&lt;/a&gt;. In plain builder terms: do not count on fast alerts for every pipeline. Count on them for stable, long-running jobs where “too quick” is actually suspicious.&lt;/p&gt;
&lt;p&gt;The model is intentionally per-pipeline. That avoids the lazy alerting pattern where every job inherits a single global SLA and the useful signals drown in noise. A five minute Salesforce Load and a two hour warehouse transformation should have different ideas of normal.&lt;/p&gt;
&lt;h2 id=&quot;how-do-you-configure-the-alert-without-wiring-every-pipeline-by-hand&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-anomaly-alerts/#how-do-you-configure-the-alert-without-wiring-every-pipeline-by-hand&quot;&gt;&lt;span&gt;How do you configure the alert without wiring every pipeline by hand?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Configuration lives in Notifications, not on every component canvas. Matillion says registered Maia users can configure pipeline notifications at the project and environment level for completion alerts and anomaly alerts in the &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;notifications overview&lt;/a&gt;. That is the right place for it, because alert routing is an operational concern. You want production to hit a high-priority channel and development to stay in a sandbox channel.&lt;/p&gt;
&lt;p&gt;For email anomaly alerts, the setup is short: open the Profile and Account menu, choose Notifications, add an anomaly alert, select Enable anomaly alerts, choose Email, and add the notification. Matillion states that email alerts go to the email address associated with your Maia account and need no extra setup in its &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;email setup instructions&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Slack needs a webhook URL and a name. Matillion lists &lt;strong&gt;Slack Webhook URL&lt;/strong&gt; and &lt;strong&gt;Slack Webhook Name&lt;/strong&gt; as required fields when configuring Slack anomaly alerts in the &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;Slack anomaly setup section&lt;/a&gt;. That is simple, but it also means channel design matters. Do not point every environment at the same room unless you enjoy training people to ignore red lights.&lt;/p&gt;
&lt;p&gt;Webhook delivery is the most useful option for teams that already route incidents through ServiceNow, PagerDuty, Teams bridges, or internal automation. Matillion says webhook anomaly alerts require a &lt;strong&gt;Webhook URL&lt;/strong&gt;, &lt;strong&gt;Webhook Name&lt;/strong&gt;, and &lt;strong&gt;Payload Template&lt;/strong&gt; in the &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;webhook setup section&lt;/a&gt;. The catch is security: Matillion documents these webhooks as outbound only and says their payloads are not signed with HMAC in the same &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;webhook behavior note&lt;/a&gt;. Put an API gateway, shared secret, allow list, or other control in front of the receiving endpoint. A naked endpoint that takes operational actions from unsigned JSON is how incident automation becomes incident creation.&lt;/p&gt;
&lt;p&gt;A compact webhook payload for a slow anomaly could look like this:&lt;/p&gt;
&lt;pre class=&quot;language-json&quot; tabindex=&quot;0&quot;&gt;&lt;code class=&quot;language-json&quot;&gt;&lt;span class=&quot;token punctuation&quot;&gt;{&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;text&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;${pipelineName} duration anomaly: ${elapsedDurationSeconds}s elapsed, upper bound ${expectedUpperBoundSeconds}s&quot;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;executionId&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;${pipelineExecutionId}&quot;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;environment&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;${environmentDisplayName}&quot;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;,&lt;/span&gt;
  &lt;span class=&quot;token property&quot;&gt;&quot;startedAt&quot;&lt;/span&gt;&lt;span class=&quot;token operator&quot;&gt;:&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;${startedAt}&quot;&lt;/span&gt;
&lt;span class=&quot;token punctuation&quot;&gt;}&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The variable names above come from Matillion&#39;s anomaly webhook template support, which includes &lt;code&gt;${elapsedDurationSeconds}&lt;/code&gt;, &lt;code&gt;${expectedUpperBoundSeconds}&lt;/code&gt;, &lt;code&gt;${expectedLowerBoundSeconds}&lt;/code&gt;, &lt;code&gt;${pipelineExecutionId}&lt;/code&gt;, and &lt;code&gt;${triggerType}&lt;/code&gt; in the &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;available variables list&lt;/a&gt;. Keep the payload small. The first alert should tell the on-call engineer what is drifting and where to click, not dump a novella into Slack.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Alert path&lt;/th&gt;
&lt;th&gt;Trigger behavior&lt;/th&gt;
&lt;th&gt;Best use&lt;/th&gt;
&lt;th&gt;Hard limit to remember&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Email anomaly alert&lt;/td&gt;
&lt;td&gt;Sends to the Maia account email when an anomaly is detected&lt;/td&gt;
&lt;td&gt;Small teams and low-noise production projects&lt;/td&gt;
&lt;td&gt;Requires at least 10 successful pipeline runs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Slack anomaly alert&lt;/td&gt;
&lt;td&gt;Sends to a Slack channel through an incoming webhook URL&lt;/td&gt;
&lt;td&gt;Team triage by project or environment&lt;/td&gt;
&lt;td&gt;Requires a Slack Webhook URL and Slack Webhook Name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Webhook anomaly alert&lt;/td&gt;
&lt;td&gt;Sends an HTTP POST with your custom JSON template&lt;/td&gt;
&lt;td&gt;Incident tools, internal routing, and FinOps automation&lt;/td&gt;
&lt;td&gt;Outbound only and not HMAC signed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&quot;what-does-this-cost-and-where-can-it-save-money&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-anomaly-alerts/#what-does-this-cost-and-where-can-it-save-money&quot;&gt;&lt;span&gt;What does this cost, and where can it save money?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Matillion does not document a separate credit price for anomaly alerts in the notification guide. The cost angle is indirect and more useful: runtime anomalies tell you when a pipeline is consuming time outside its normal envelope, and time is where warehouse, runner, and orchestration costs usually hide.&lt;/p&gt;
&lt;p&gt;The cleanest workflow is to treat a slow anomaly as a cost investigation starter. Matillion&#39;s REST API consumption data is not real time, since the June 10 changelog says &lt;code&gt;GET /v1/consumption&lt;/code&gt; data refreshes every &lt;strong&gt;three hours&lt;/strong&gt; to reflect recent credit usage in the &lt;a href=&quot;https://docs.maia.ai/docs/changelog/2026-changelog&quot;&gt;API update notes&lt;/a&gt;. That lag means anomaly alerts are the faster operational signal, while consumption reporting is the later accounting signal.&lt;/p&gt;
&lt;p&gt;Use both. If Maia alerts that &lt;code&gt;nightly_customer_mart&lt;/code&gt; crossed its expected upper bound at 02:14, you inspect Pipeline Runs immediately. Later, you check consumption by pipeline, project, or user once the usage data refreshes. Matillion&#39;s MCP server documentation explicitly includes cost optimization use cases such as analyzing credit consumption, seeing breakdowns by pipeline and execution frequency, and identifying expensive pipelines in the &lt;a href=&quot;https://docs.maia.ai/docs/api-reference/mcp-server&quot;&gt;MCP server use case table&lt;/a&gt;. The alert tells you where to look. Consumption data tells you what it cost.&lt;/p&gt;
&lt;p&gt;This is also where alert tuning becomes business work, not only engineering hygiene. A pipeline that runs 30 percent long once after a large customer import may be acceptable. A pipeline that drifts 8 percent longer every weekday for two weeks is a margin leak with a nicer UI. If you already use Maia context and governance patterns, connect this operational signal to your broader standards. Our guide to &lt;a href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/&quot;&gt;Matillion Context Engine cost controls&lt;/a&gt; is a useful companion because the same discipline applies: give Maia enough context to help, but keep humans accountable for cost policy.&lt;/p&gt;
&lt;h2 id=&quot;when-should-you-trust-the-alert-and-when-should-you-be-skeptical&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-anomaly-alerts/#when-should-you-trust-the-alert-and-when-should-you-be-skeptical&quot;&gt;&lt;span&gt;When should you trust the alert, and when should you be skeptical?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Trust the alert when the pipeline has stable history, runs on a schedule, and does meaningful work every time. A nightly dimension rebuild, a predictable Salesforce incremental load, or a transformation pipeline behind an executive dashboard are good candidates. If the pipeline has at least &lt;strong&gt;10 successful runs&lt;/strong&gt; and normally lands in a tight runtime band, an alert is a real signal.&lt;/p&gt;
&lt;p&gt;Be skeptical when the workload is naturally lumpy. Month-end finance jobs, backfills, campaign data imports, and pipelines with parameter-driven row counts will produce wider expected ranges because their past durations are wider. Matillion says naturally variable pipelines get a wider range so ordinary fluctuation does not trigger false alerts in the &lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;expected range explanation&lt;/a&gt;. That behavior is good, but it also means a genuinely bad run might need to be very bad before it crosses the threshold.&lt;/p&gt;
&lt;p&gt;Also watch branch and environment changes. Matillion&#39;s pipeline run history shows fields such as status, environment, artifact version, trigger, start time, end time, and duration in the &lt;a href=&quot;https://docs.matillion.com/data-productivity-cloud/hub/docs/pipeline-observability/&quot;&gt;pipeline run history documentation&lt;/a&gt;. If an anomaly follows a new artifact version, start with the diff. If it follows an environment change, start with variables, warehouse sizing, secrets, and connection overrides. If it follows neither, start with source volume and platform health.&lt;/p&gt;
&lt;p&gt;One more practical guardrail: do not auto-kill long-running jobs from the first webhook you receive. Use the webhook to open an incident, enrich a ticket, or notify the owning channel. Termination should require a second rule, such as the run exceeding a business SLA, a known bad artifact version, or a warehouse spend threshold.&lt;/p&gt;
&lt;h2 id=&quot;what-would-i-do-before-turning-this-on-in-production&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-anomaly-alerts/#what-would-i-do-before-turning-this-on-in-production&quot;&gt;&lt;span&gt;What would I do before turning this on in production?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Start with your top five scheduled or API-driven pipelines by business impact, not your loudest pipelines. Pick jobs with clear owners, stable schedules, and visible cost or SLA consequences. Then configure anomaly alerts at the production environment level and route them to a channel where someone is actually expected to respond.&lt;/p&gt;
&lt;p&gt;For each pipeline, capture four pieces of runbook context:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The normal schedule and expected business deadline.&lt;/li&gt;
&lt;li&gt;The primary source systems and largest transformation tables.&lt;/li&gt;
&lt;li&gt;The first Matillion page to inspect, usually Pipeline Runs or the run detail view.&lt;/li&gt;
&lt;li&gt;The cost check to run after the three hour consumption refresh window.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Do not overbuild on day one. The first useful version is one notification, one channel, and one runbook link. After a week, add webhook routing for the pipelines that produce actionable signals. If alerts are noisy, check whether the pipeline is parameterized or lumpy by design before blaming Maia.&lt;/p&gt;
&lt;p&gt;The bigger point is cultural. Runtime drift is often treated as a nuisance until it becomes a failure. Maia&#39;s duration anomaly alerts make drift visible early enough to do something cheaper than a postmortem. That is the feature&#39;s value: it moves the conversation from “why did the dashboard miss?” to “why is this job behaving differently right now?”&lt;/p&gt;
&lt;h2 id=&quot;sources&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-anomaly-alerts/#sources&quot;&gt;&lt;span&gt;Sources&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/changelog/2026-changelog&quot;&gt;Matillion changelog: 2026 changelog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/pipeline-notifications&quot;&gt;Matillion documentation: Pipeline notifications&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/maia-troubleshooting&quot;&gt;Matillion documentation: Troubleshoot pipelines with Maia AI Agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.matillion.com/data-productivity-cloud/hub/docs/pipeline-observability/&quot;&gt;Matillion documentation: Pipeline run history&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/api-reference/mcp-server&quot;&gt;Matillion documentation: Matillion&#39;s MCP server&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
  <entry>
    <title>Matillion Context Engine: setup, limits, and costs</title>
    <link href="https://data-today.net/matillion/matillion-context-engine-costs/" />
    <updated>2026-06-08T00:00:00Z</updated>
    <id>https://data-today.net/matillion/matillion-context-engine-costs/</id>
    <content type="html">&lt;p&gt;If you have ever watched an AI assistant build a technically valid pipeline that misunderstands the business meaning of the table, you already know why Matillion shipped this. Maia can inspect objects and work through pipeline tasks, but a warehouse schema rarely explains which customer table is canonical, which finance metric is blessed, or which lineage path is stale.&lt;/p&gt;
&lt;p&gt;Matillion Context Engine is a public preview knowledge graph feature for Maia AI Agents, designed to give Maia a living map of your business data. The key fact for a data engineer who also watches the bill: &lt;strong&gt;Context Engine is configured through knowledge graphs and crawlers, not through a magic prompt&lt;/strong&gt;, and those crawlers can be scoped, scheduled, paused, and access controlled.&lt;/p&gt;
&lt;p&gt;Matillion announced Context Engine alongside Mission Control on June 1, 2026. The pairing matters. Context Engine gives Maia the domain map. Mission Control gives you a kanban style place to launch, supervise, and review Maia tasks that can use that map. If you need the broader platform context, start with our guide to &lt;a href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/&quot;&gt;Matillion&#39;s Data Productivity Cloud for builders&lt;/a&gt;, then come back here for the Context Engine operating model.&lt;/p&gt;
&lt;h2 id=&quot;what-did-matillion-actually-ship-in-context-engine&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/#what-did-matillion-actually-ship-in-context-engine&quot;&gt;&lt;span&gt;What did Matillion actually ship in Context Engine?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Context Engine adds a new surface in Maia for creating and managing knowledge graphs. In Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/guides/context-engine&quot;&gt;Context Engine documentation&lt;/a&gt;, a knowledge graph captures the structure, relationships, and meaning of your data so Maia AI Agents can use a specific graph while working on a task. That is the important shift. You are no longer relying only on a chat transcript, a project file, or a pasted naming convention.&lt;/p&gt;
&lt;p&gt;The feature is in &lt;strong&gt;public preview&lt;/strong&gt;, so treat it as something to pilot with real work, not something to hand unrestricted production authority on day one. Matillion&#39;s release FAQ says public preview features are available to all users but are not recommended for production workloads. That matters for governance. If a graph teaches Maia the wrong semantic relationship, you want that mistake discovered in a sandbox branch, not after a merge into your main data product.&lt;/p&gt;
&lt;p&gt;The nouns are simple:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;knowledge graph&lt;/strong&gt; is the domain container. Think Finance, Sales, Marketing, Customer 360, or Product Telemetry.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;crawler&lt;/strong&gt; populates the graph from warehouse metadata or pipeline execution history.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Projects&lt;/strong&gt; define which Maia projects can use a restricted graph.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Access&lt;/strong&gt; defines which users can manage the graph.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That is more operational than glamorous. Good. The AI feature you can govern usually beats the dazzling one you cannot explain during a cost review.&lt;/p&gt;
&lt;h2 id=&quot;how-do-knowledge-graphs-and-crawlers-actually-work&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/#how-do-knowledge-graphs-and-crawlers-actually-work&quot;&gt;&lt;span&gt;How do knowledge graphs and crawlers actually work?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;You create a graph from the Context Engine dashboard, then choose whether it is Public or Restricted. A Public graph is available to all projects. A Restricted graph is available only to selected projects. Matillion warns that public graphs should not ingest sensitive data or data that should not be available across projects, which is the sentence your security lead will underline twice.&lt;/p&gt;
&lt;p&gt;To add a graph, the user needs the &lt;strong&gt;Admin or Super Admin&lt;/strong&gt; account role. After the graph exists, it has three tabs: Crawlers, Projects, and Access. The Crawlers tab adds and monitors data ingestion into the graph. Projects controls the allowlist for restricted graphs. Access controls which users can manage the graph.&lt;/p&gt;
&lt;p&gt;The crawler setup has a few concrete choices. In the Add crawler flow, you enter a crawler name, choose the type, select a project, select an environment, then schedule it. Context Engine supports &lt;strong&gt;2 crawler types&lt;/strong&gt; today:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Crawler type&lt;/th&gt;
&lt;th&gt;What it harvests&lt;/th&gt;
&lt;th&gt;Best first use&lt;/th&gt;
&lt;th&gt;Cost watchpoint&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Warehouse data&lt;/td&gt;
&lt;td&gt;Warehouse and structured source metadata&lt;/td&gt;
&lt;td&gt;Domain model discovery, table and column meaning&lt;/td&gt;
&lt;td&gt;Scope databases, catalogs, and schemas tightly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pipeline execution&lt;/td&gt;
&lt;td&gt;Pipeline runs for a selected project and environment&lt;/td&gt;
&lt;td&gt;Freshness, lineage, operational flow&lt;/td&gt;
&lt;td&gt;Avoid crawling noisy dev environments by default&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;For Warehouse data crawlers, the selection step changes by warehouse. Snowflake uses databases and schemas. Databricks uses catalog and schemas. Amazon Redshift uses schemas. That gives you &lt;strong&gt;3 warehouse selection patterns&lt;/strong&gt; to plan for if your estate spans clouds or warehouses.&lt;/p&gt;
&lt;p&gt;The chart below shows the operating surface you need to account for before enabling this broadly: 2 crawler types, 3 warehouse selection patterns, 4 Mission Control task columns, 6 crawler statuses, and a 10 task cap in Mission Control. The numbers are small enough to manage, but large enough to need ownership.&lt;/p&gt;
&lt;figure class=&quot;figure&quot;&gt;&lt;img src=&quot;https://data-today.net/posts/matillion-context-engine-costs-fig-context-engine-surface.png&quot; alt=&quot;Bar chart for Matillion Context Engine showing 2 crawler types, 3 warehouse selection patterns, 4 Mission Control columns, 6 crawler statuses, and a 10 task in progress cap.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;figcaption&gt;Matillion documentation lists 2 Context Engine crawler types, 3 warehouse selection patterns, 6 crawler statuses, 4 Mission Control columns, and a 10 task in progress cap.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Matillion lists &lt;strong&gt;6 crawler statuses&lt;/strong&gt;: Successful, Initializing, Extracting, Pending, Paused, and Failed. That status model is enough for a runbook. If a crawler is Pending, it has not run its first crawl. If it is Paused, it will not run until resumed. If it is Failed, you have a concrete object to investigate instead of a vague complaint that Maia has bad context.&lt;/p&gt;
&lt;h2 id=&quot;how-should-you-configure-the-first-graph-without-creating-semantic-sludge&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/#how-should-you-configure-the-first-graph-without-creating-semantic-sludge&quot;&gt;&lt;span&gt;How should you configure the first graph without creating semantic sludge?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Start with one graph per domain, not one graph for the company. Matillion&#39;s launch post says Context Engine is intended for domain focused knowledge graphs such as Finance, Sales, and Marketing. That is the right default because semantics are local. A field named &lt;code&gt;status&lt;/code&gt; in a billing table and a field named &lt;code&gt;status&lt;/code&gt; in a product event stream may both be valid and still mean completely different things.&lt;/p&gt;
&lt;p&gt;A practical first graph might look like this:&lt;/p&gt;
&lt;pre class=&quot;language-yaml&quot; tabindex=&quot;0&quot;&gt;&lt;code class=&quot;language-yaml&quot;&gt;&lt;span class=&quot;token key atrule&quot;&gt;knowledge_graph&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt;
  &lt;span class=&quot;token key atrule&quot;&gt;name&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; finance_reporting
  &lt;span class=&quot;token key atrule&quot;&gt;availability&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; restricted
  &lt;span class=&quot;token key atrule&quot;&gt;projects&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;token punctuation&quot;&gt;-&lt;/span&gt; finance_dbt_migration
    &lt;span class=&quot;token punctuation&quot;&gt;-&lt;/span&gt; revenue_reporting
  &lt;span class=&quot;token key atrule&quot;&gt;crawlers&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt;
    &lt;span class=&quot;token punctuation&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;token key atrule&quot;&gt;type&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; warehouse_data
      &lt;span class=&quot;token key atrule&quot;&gt;warehouse&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; snowflake
      &lt;span class=&quot;token key atrule&quot;&gt;scope&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; FINANCE_MART.PUBLIC
      &lt;span class=&quot;token key atrule&quot;&gt;schedule&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; standard_daily
    &lt;span class=&quot;token punctuation&quot;&gt;-&lt;/span&gt; &lt;span class=&quot;token key atrule&quot;&gt;type&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; pipeline_execution
      &lt;span class=&quot;token key atrule&quot;&gt;project&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; revenue_reporting
      &lt;span class=&quot;token key atrule&quot;&gt;environment&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;:&lt;/span&gt; dev&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That is not an import format. It is the configuration shape you should decide before clicking through the UI. The point is discipline: &lt;strong&gt;one domain, restricted access, one narrow warehouse scope, one environment for operational history&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Before crawling, improve the metadata Maia will see. Matillion recommends adding metadata such as column descriptions and tags to databases, schemas, or datasets because it helps Maia understand the landscape. In Snowflake, a tiny amount of comment hygiene can save a lot of prompt babysitting:&lt;/p&gt;
&lt;pre class=&quot;language-sql&quot; tabindex=&quot;0&quot;&gt;&lt;code class=&quot;language-sql&quot;&gt;&lt;span class=&quot;token keyword&quot;&gt;comment&lt;/span&gt; &lt;span class=&quot;token keyword&quot;&gt;on&lt;/span&gt; &lt;span class=&quot;token keyword&quot;&gt;table&lt;/span&gt; FINANCE_MART&lt;span class=&quot;token punctuation&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;token keyword&quot;&gt;PUBLIC&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;.&lt;/span&gt;INVOICE_FACT
  &lt;span class=&quot;token operator&quot;&gt;is&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&#39;Canonical invoice fact table for recognized revenue reporting&#39;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;;&lt;/span&gt;

&lt;span class=&quot;token keyword&quot;&gt;comment&lt;/span&gt; &lt;span class=&quot;token keyword&quot;&gt;on&lt;/span&gt; &lt;span class=&quot;token keyword&quot;&gt;column&lt;/span&gt; FINANCE_MART&lt;span class=&quot;token punctuation&quot;&gt;.&lt;/span&gt;&lt;span class=&quot;token keyword&quot;&gt;PUBLIC&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;.&lt;/span&gt;INVOICE_FACT&lt;span class=&quot;token punctuation&quot;&gt;.&lt;/span&gt;NET_REVENUE_USD
  &lt;span class=&quot;token operator&quot;&gt;is&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&#39;Net recognized revenue in USD after credits and discounts&#39;&lt;/span&gt;&lt;span class=&quot;token punctuation&quot;&gt;;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Do this for the five to ten tables people always argue about first. Do not boil the metadata ocean. Context Engine can crawl a wide surface, but your first success metric is whether Maia stops confusing the canonical thing with the convenient thing.&lt;/p&gt;
&lt;h2 id=&quot;what-does-this-change-when-you-use-mission-control&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/#what-does-this-change-when-you-use-mission-control&quot;&gt;&lt;span&gt;What does this change when you use Mission Control?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Mission Control is the other half of the release, and Context Engine is one of the task inputs. When creating a Mission Control task, Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/guides/mission-control&quot;&gt;Mission Control guide&lt;/a&gt; says you choose a project, source branch, environment, and Knowledge Graph, then write the prompt. That Knowledge Graph drop-down is where Context Engine becomes operational.&lt;/p&gt;
&lt;p&gt;Mission Control gives each task its own chat interface and board position. The board has &lt;strong&gt;4 columns&lt;/strong&gt;: Backlog, In progress, Needs attention, and Completed. It also has a hard workflow constraint: a maximum of &lt;strong&gt;10 tasks&lt;/strong&gt; can be in progress at any time. Once 10 are in progress, you can still create a task, but it lands in Backlog. You cannot start another backlog task, and responses to Needs attention tasks are not processed until fewer than 10 tasks are in progress.&lt;/p&gt;
&lt;p&gt;That cap is useful. It prevents the worst version of AI delegation: twenty half reviewed branches, all burning attention, none mergeable.&lt;/p&gt;
&lt;p&gt;The feature also creates a new branch for each task in Mission Control if you select that option. Maia can work on the branch, and you can open it in Designer to review the work. Completing a task in Mission Control does not merge the branch. You still need to commit, push, and merge changes from the task branch to the source branch, and Matillion says Maia cannot merge changes. That separation is exactly what you want for an AI agent touching pipeline logic.&lt;/p&gt;
&lt;p&gt;The workflow I would use is simple:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Create a restricted Context Engine graph for a domain.&lt;/li&gt;
&lt;li&gt;Crawl a dev or analytics environment first.&lt;/li&gt;
&lt;li&gt;Launch a Mission Control task from a fresh branch.&lt;/li&gt;
&lt;li&gt;Select the graph explicitly.&lt;/li&gt;
&lt;li&gt;Keep Ask permission on unless the task is in a disposable sandbox.&lt;/li&gt;
&lt;li&gt;Review the branch in Designer before merge.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Bypass permissions mode exists, but it should be a narrow tool. Matillion says Ask permission is the default and Bypass permissions runs tool calls immediately with no approval prompts. Use it for trusted, hands off sandbox work. Do not use it on a broad graph with production credentials and a PDF from someone you do not know. That is how a demo becomes a cleanup ticket.&lt;/p&gt;
&lt;h2 id=&quot;what-does-context-engine-cost-in-practice&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/#what-does-context-engine-cost-in-practice&quot;&gt;&lt;span&gt;What does Context Engine cost in practice?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Matillion does not publish a separate Context Engine price meter on the public feature documentation. So the honest answer is operational: your cost exposure comes from crawler scope, crawler schedule, Maia task activity, warehouse metadata queries, and the human review time needed to trust the result.&lt;/p&gt;
&lt;p&gt;The bill lever you control first is crawler design. A Warehouse data crawler over a narrow Snowflake schema is a different operational bet than a crawler pointed at every schema reachable by a broad role. A Pipeline execution crawler on a dev project with five representative pipelines is a different bet than one aimed at a noisy shared environment. The docs say crawlers can use Standard schedule settings or Advanced cron expressions, and they can be paused or resumed. Use that.&lt;/p&gt;
&lt;p&gt;For a pilot, avoid the three expensive habits:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Crawling every domain because the button is available.&lt;/li&gt;
&lt;li&gt;Making the first graph Public because it saves two clicks.&lt;/li&gt;
&lt;li&gt;Running Bypass permissions while Maia has access to credentials that are wider than the task.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;There is also a hidden cost in bad context. If Context Engine learns that an old staging table is the customer source of truth, Maia may generate plausible work that takes longer to review than a human built pipeline. The answer is not fear. The answer is graph ownership. Put a named data engineer or analytics engineer on each domain graph, just as you would put an owner on a dbt package or shared pipeline.&lt;/p&gt;
&lt;h2 id=&quot;when-is-context-engine-the-wrong-choice&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/#when-is-context-engine-the-wrong-choice&quot;&gt;&lt;span&gt;When is Context Engine the wrong choice?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Skip Context Engine for one off work where the domain model is obvious and the pipeline is small. A two component load from a clean SaaS source does not need a knowledge graph. A context file in &lt;code&gt;.matillion/maia/rules/&lt;/code&gt; may be enough if all you need is naming rules, project conventions, or a short business rule. Matillion&#39;s context file docs still matter because context files are always read by Maia when stored in the reserved rules folder, and they have a strict &lt;strong&gt;12,000 character&lt;/strong&gt; limit across Markdown files in that folder.&lt;/p&gt;
&lt;p&gt;Use Context Engine when the issue is semantic and operational, not just stylistic. Good fits include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A finance mart with competing definitions of revenue.&lt;/li&gt;
&lt;li&gt;A migration where pipeline execution history tells Maia how data actually flows.&lt;/li&gt;
&lt;li&gt;A governed domain where only selected projects should use the graph.&lt;/li&gt;
&lt;li&gt;A team that wants Mission Control tasks to start from the same domain map.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Bad fits include scratch work, sensitive domains with no access model yet, and production only warehouses with thin metadata. Context Engine can make Maia more useful, but it can also make wrong assumptions more reusable. Reuse is wonderful when the thing is right. It is brutal when the thing is wrong.&lt;/p&gt;
&lt;h2 id=&quot;what-should-you-do-next-if-you-influence-the-matillion-bill&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/#what-should-you-do-next-if-you-influence-the-matillion-bill&quot;&gt;&lt;span&gt;What should you do next if you influence the Matillion bill?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Run a narrow pilot. Pick one domain with real ambiguity and enough metadata to teach Maia something. Create a Restricted graph, add one Warehouse data crawler, add one Pipeline execution crawler only if execution history matters, and schedule both conservatively. Then create two Mission Control tasks against the graph: one pipeline build and one pipeline explanation. Compare the review burden with your normal Maia workflow.&lt;/p&gt;
&lt;p&gt;Your success criteria should be concrete:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Did Maia choose the right source tables without repeated correction?&lt;/li&gt;
&lt;li&gt;Did it use the domain language your team uses in reviews?&lt;/li&gt;
&lt;li&gt;Did the generated branch require fewer manual edits?&lt;/li&gt;
&lt;li&gt;Did crawler failures or stale metadata show up clearly enough to operate?&lt;/li&gt;
&lt;li&gt;Did the schedule add background work you can justify?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If the answer is yes, expand by domain. If the answer is no, fix metadata before widening the crawl. The fastest way to waste money here is to scale uncertainty.&lt;/p&gt;
&lt;p&gt;Context Engine is promising because it moves Maia from prompt memory toward governed domain memory. That is the kind of AI feature builders should want: less sparkle, more accountable surfaces. Give it a map. Just make sure you own the cartography.&lt;/p&gt;
&lt;h2 id=&quot;sources&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine-costs/#sources&quot;&gt;&lt;span&gt;Sources&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/context-engine&quot;&gt;Maia Documentation: Context Engine&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/mission-control&quot;&gt;Maia Documentation: Mission Control&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.maia.ai/resources/changelogs&quot;&gt;Maia Changelog: This week in Maia: Mission Control and Context Engine&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/maia-context-files&quot;&gt;Maia Documentation: Context files&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.matillion.com/data-productivity-cloud/release-faq/&quot;&gt;Maia Documentation: Maia product release FAQ&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
  <entry>
    <title>Matillion&#39;s Data Productivity Cloud, explained for builders</title>
    <link href="https://data-today.net/matillion/matillion-data-productivity-cloud/" />
    <updated>2026-06-07T00:00:00Z</updated>
    <id>https://data-today.net/matillion/matillion-data-productivity-cloud/</id>
    <content type="html">&lt;p&gt;If you have used Matillion before, you probably picture the old ETL tool that ran on a VM you had to size, patch, and babysit. The &lt;strong&gt;Data Productivity Cloud (DPC)&lt;/strong&gt; is the rebuilt, cloud-native successor, and it changes enough about how pipelines run and bill that it is worth a proper walkthrough. This guide is the starting point for our studio: what the platform actually is, how a pipeline executes, and where the money goes.&lt;/p&gt;
&lt;p&gt;Matillion&#39;s pitch is that one platform should let three different kinds of people build the same pipeline: a low-code analyst dragging components in the Designer, an engineer writing dbt, SQL, or Python, and increasingly an AI agent under Matillion&#39;s &lt;strong&gt;Maia&lt;/strong&gt; brand. The interesting question for a builder is not the marketing claim, it is the plumbing underneath: what runs where, what you pay for, and where the platform helps versus gets in the way.&lt;/p&gt;
&lt;h2 id=&quot;what-is-the-data-productivity-cloud-concretely&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/#what-is-the-data-productivity-cloud-concretely&quot;&gt;&lt;span&gt;What is the Data Productivity Cloud, concretely?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The DPC is a fully managed, browser-based environment for data integration. You do not run a Matillion instance anymore. Instead, Matillion hosts the control plane and you connect it to your cloud data warehouse, typically &lt;a href=&quot;https://data-today.net/snowflake/&quot;&gt;Snowflake&lt;/a&gt;, Databricks, Amazon Redshift, Google BigQuery, or Microsoft Fabric.&lt;/p&gt;
&lt;p&gt;Two kinds of work happen in a DPC project:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Ingestion&lt;/strong&gt; pulls data from sources (databases, SaaS APIs, files) into your warehouse using prebuilt connectors and change data capture.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transformation&lt;/strong&gt; reshapes that data once it lands, and this is the part that matters most for cost.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The key architectural fact is &lt;strong&gt;pushdown&lt;/strong&gt;. When you build a transformation pipeline in the Designer, Matillion does not move rows through its own engine. It compiles your pipeline into SQL and pushes that SQL down to your warehouse, which does the heavy lifting. That single design choice explains most of the platform&#39;s cost behaviour, which we will come back to.&lt;/p&gt;
&lt;h2 id=&quot;how-does-a-pipeline-actually-run&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/#how-does-a-pipeline-actually-run&quot;&gt;&lt;span&gt;How does a pipeline actually run?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;A DPC project separates two pipeline types on purpose, and mixing them up is the most common beginner mistake.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Transformation pipelines&lt;/strong&gt; run SQL against your warehouse. They have no orchestration logic; they just transform.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Orchestration pipelines&lt;/strong&gt; are the conductor. They run ingestion jobs, call transformation pipelines, branch on success or failure, loop with iterators, and handle scheduling.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;A healthy project keeps a thin orchestration layer that calls many focused transformation pipelines, rather than one giant pipeline that tries to do everything. The same discipline you would apply to functions in code applies here: small, named, reusable units.&lt;/p&gt;
&lt;p&gt;The chart below shows where the run minutes of a typical project go. The bulk of the time, roughly &lt;strong&gt;60 percent&lt;/strong&gt;, is transformation SQL executing inside your warehouse, around &lt;strong&gt;30 percent&lt;/strong&gt; is ingestion, and the remaining &lt;strong&gt;10 percent&lt;/strong&gt; is orchestration overhead. The exact mix varies, but the shape holds: your warehouse, not Matillion, is doing most of the work.&lt;/p&gt;
&lt;figure class=&quot;figure&quot;&gt;&lt;img src=&quot;https://data-today.net/posts/matillion-data-productivity-cloud-fig.png&quot; alt=&quot;Horizontal bars showing transformation pushdown taking about 60 percent of run minutes, ingestion 30 percent, and orchestration overhead 10 percent.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;figcaption&gt;Illustrative: a typical split of pipeline run minutes across transformation pushdown, ingestion, and orchestration overhead. Data Today.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;where-does-the-cost-actually-go&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/#where-does-the-cost-actually-go&quot;&gt;&lt;span&gt;Where does the cost actually go?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;This is the question that decides whether a DPC project stays affordable, and the answer has two halves that you pay separately.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost layer&lt;/th&gt;
&lt;th&gt;What you pay for&lt;/th&gt;
&lt;th&gt;Who bills you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Matillion credits&lt;/td&gt;
&lt;td&gt;Pipeline runs and platform usage, metered by Matillion&lt;/td&gt;
&lt;td&gt;Matillion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Warehouse compute&lt;/td&gt;
&lt;td&gt;The SQL pushdown that transformations execute&lt;/td&gt;
&lt;td&gt;Your cloud warehouse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ingestion&lt;/td&gt;
&lt;td&gt;Rows or connectors moved, depending on plan&lt;/td&gt;
&lt;td&gt;Matillion&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Because transformation runs as pushdown SQL, a badly written pipeline does not just burn Matillion credits, it runs an expensive query on your warehouse and shows up on a second bill. &lt;strong&gt;The most common cost surprise is a transformation that scans far more data than it needs&lt;/strong&gt;, often because someone left a full reload where an incremental load belonged. Optimizing the SQL your pipeline generates is therefore a warehouse-cost exercise as much as a Matillion one, which is exactly why we treat cost as its own section.&lt;/p&gt;
&lt;h2 id=&quot;what-about-the-runners&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/#what-about-the-runners&quot;&gt;&lt;span&gt;What about the runners?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The DPC runs your pipelines on compute called &lt;strong&gt;runners&lt;/strong&gt;. Matillion offers fully hosted runners so you do not manage infrastructure, and self-hosted or cloud-hosted runner options for teams that need pipelines to execute inside their own network, for example to reach a private database without exposing it to the internet.&lt;/p&gt;
&lt;p&gt;The trade-off is the usual managed-versus-self-hosted one:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hosted runners&lt;/strong&gt; are zero-maintenance and the fastest way to start, but they run in Matillion&#39;s environment.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Self-hosted runners&lt;/strong&gt; keep execution and credentials inside your perimeter, at the cost of you owning the compute and its upkeep.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If you handle regulated data or sit behind strict network controls, the runner choice is the first architectural decision to get right, well before you build a single pipeline.&lt;/p&gt;
&lt;h2 id=&quot;where-does-maia-the-ai-layer-fit&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/#where-does-maia-the-ai-layer-fit&quot;&gt;&lt;span&gt;Where does Maia, the AI layer, fit?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Maia is Matillion&#39;s name for the AI features layered across the DPC: copilots that help build and explain pipelines, agents that can take actions through API endpoints, and assistance for tasks like root cause analysis when a pipeline fails. The honest read is that this is the fastest-moving and least settled part of the platform, which is precisely why it deserves close, sceptical coverage rather than hype.&lt;/p&gt;
&lt;p&gt;For a builder, the practical stance is to let AI accelerate the boring parts, generating boilerplate transformations, suggesting fixes, explaining an unfamiliar pipeline, while keeping a human reviewing anything that touches production data. We will track each Maia capability as it ships and judge whether it is genuinely production-ready or still a demo.&lt;/p&gt;
&lt;h2 id=&quot;what-should-you-do-with-this&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/#what-should-you-do-with-this&quot;&gt;&lt;span&gt;What should you do with this?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;If you are evaluating or adopting the Data Productivity Cloud, a few principles travel well:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Treat your warehouse as the engine.&lt;/strong&gt; Most of your cost and performance lives in the pushdown SQL, not in Matillion. Profile the queries your pipelines generate.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Keep orchestration thin and transformations small.&lt;/strong&gt; Reusable, well-named pipelines age far better than monoliths.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Decide runners early.&lt;/strong&gt; Network and compliance constraints shape the whole project.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Adopt Maia deliberately.&lt;/strong&gt; Use it where review is cheap; gate it where mistakes are expensive.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This guide is the foundation. From here, the guides go deeper on each piece: ingestion connectors and change data capture, transformation patterns in the Designer, orchestration controls like iterators and scheduling, the FinOps habits that keep credits in check, and the Maia AI features as they land. The platform is moving quickly, and the goal here is the same as everywhere on Data Today: tell you what actually changed and what it means for the thing you are building.&lt;/p&gt;
&lt;h2 id=&quot;sources&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/#sources&quot;&gt;&lt;span&gt;Sources&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.matillion.com/&quot;&gt;Matillion Data Productivity Cloud documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://roadmap.matillion.com/&quot;&gt;Matillion changelog and new features blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.matillion.com/&quot;&gt;Matillion product overview&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
  <entry>
    <title>Matillion Context Engine grounds Maia agent work</title>
    <link href="https://data-today.net/matillion/matillion-context-engine/" />
    <updated>2026-06-07T00:00:00Z</updated>
    <id>https://data-today.net/matillion/matillion-context-engine/</id>
    <content type="html">&lt;p&gt;A data agent that can build pipelines is useful. A data agent that knows which tables matter, which pipelines already touch them, and when to stop and ask you before firing a tool call is the version you can let near production.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Matillion Context Engine is the new public preview layer that gives Maia AI Agents knowledge graphs, crawlers, and task context, with Mission Control adding a 10 task in-progress cap around agent work.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Matillion shipped Context Engine and Mission Control together because the old problem with AI assistants in data engineering is not syntax. It is context. Maia can already build orchestration and transformation pipelines, query warehouse data, sample pipeline data, manage files, and commit or push changes inside the Data Productivity Cloud, according to Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/guides/maia-ai-agents-overview&quot;&gt;Maia AI Agents overview&lt;/a&gt;. Context Engine gives those agents a living map of your warehouse metadata, pipeline execution history, business language, and project scope. Mission Control gives you a kanban board where that work becomes task shaped instead of chat shaped.&lt;/p&gt;
&lt;p&gt;If you are still getting oriented around the platform, start with our guide to &lt;a href=&quot;https://data-today.net/matillion/matillion-data-productivity-cloud/&quot;&gt;Matillion&#39;s Data Productivity Cloud for builders&lt;/a&gt;. This piece assumes you already build in Maia or Designer and now need to decide where Context Engine belongs in your operating model.&lt;/p&gt;
&lt;h2 id=&quot;what-did-matillion-actually-ship-in-context-engine&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine/#what-did-matillion-actually-ship-in-context-engine&quot;&gt;&lt;span&gt;What did Matillion actually ship in Context Engine?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Context Engine is in public preview, and the important object is the &lt;strong&gt;knowledge graph&lt;/strong&gt;. Matillion describes it as a way to capture the structure, relationships, and meaning of your data, then let Maia use that graph when it works on a task in Mission Control or chat. In plain builder terms: it is metadata grounding for Maia, scoped to projects and fed by crawlers, not another Markdown rules file with a nicer name.&lt;/p&gt;
&lt;p&gt;The Context Engine dashboard sits under the AI Agents icon in the left navigation. It lists knowledge graphs you can access, lets you filter by project, and supports search by name or description. From there, an Admin or Super Admin can add a knowledge graph, choose whether it is Public or Restricted, and give it a name and description, as Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/guides/context-engine&quot;&gt;Context Engine documentation&lt;/a&gt; spells out.&lt;/p&gt;
&lt;p&gt;That Public or Restricted choice is not cosmetic. A public knowledge graph is available in all projects. A restricted graph is available only to selected projects. Matillion explicitly warns that public graphs should not ingest sensitive data that should not be available to all projects. That is the kind of sentence you should read twice before turning a sales operations graph into a company wide default.&lt;/p&gt;
&lt;p&gt;Once a graph exists, it has three important tabs:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Crawlers&lt;/strong&gt;, where you add and monitor crawlers that populate the graph.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Projects&lt;/strong&gt;, where you manage which projects can use a restricted graph.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Access&lt;/strong&gt;, where you add users who can manage the graph.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The crawler model is refreshingly concrete. Matillion supports &lt;strong&gt;2 crawler types&lt;/strong&gt;: Warehouse data crawlers and Pipeline execution crawlers. Warehouse data crawlers harvest warehouses and structured sources supported through connectors. Pipeline execution crawlers harvest executions for a chosen project and environment, giving Maia operational context about how work actually flows.&lt;/p&gt;
&lt;figure class=&quot;figure&quot;&gt;&lt;img src=&quot;https://data-today.net/posts/matillion-context-engine-fig-context-engine-controls.png&quot; alt=&quot;Bar chart of Matillion Context Engine and Mission Control controls: 2 crawler types, 2 graph availability modes, 4 Mission Control board columns, and a 10 task in-progress limit.&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;&lt;figcaption&gt;Matillion documentation lists 2 Context Engine crawler types, 2 knowledge graph availability modes, 4 Mission Control board columns, and a 10 task in-progress limit.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;As the chart shows, the launch is not just a new button. Context Engine and Mission Control introduce &lt;strong&gt;2 crawler types&lt;/strong&gt;, &lt;strong&gt;2 graph availability modes&lt;/strong&gt;, &lt;strong&gt;4 task board columns&lt;/strong&gt;, and a &lt;strong&gt;10 task in-progress limit&lt;/strong&gt;. Those numbers matter because they turn agent context into something you can scope, schedule, and govern.&lt;/p&gt;
&lt;p&gt;Crawler setup has a few warehouse specific details. For Warehouse data crawlers, the data selection step depends on the target: Snowflake uses databases and schemas, Databricks uses catalog and schemas, and Amazon Redshift uses schemas. You can schedule crawler runs with Standard settings or Advanced schedule settings, including a cron expression. You can also run a crawler on demand with Run now.&lt;/p&gt;
&lt;p&gt;The crawler status model gives you the minimum you need to operate it: Successful, Initializing, Extracting, Pending, Paused, and Failed. You can inspect the latest crawl or crawl history, including status, start time, end time, and duration. That is not observability nirvana, but it is enough to answer the first operational question: did the graph refresh before Maia used it?&lt;/p&gt;
&lt;h2 id=&quot;how-does-maia-use-the-graph-when-it-starts-a-task&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine/#how-does-maia-use-the-graph-when-it-starts-a-task&quot;&gt;&lt;span&gt;How does Maia use the graph when it starts a task?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The graph becomes useful when you attach it to Maia&#39;s work. In Mission Control, the New task dialog includes Project, Branch, Environment, Knowledge Graph, and Prompt fields. The Knowledge Graph drop-down selects the graph Maia will use to inform that task, according to Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/guides/mission-control&quot;&gt;Mission Control guide&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The same idea appears in the Agent Tasks API as &lt;code&gt;graphId&lt;/code&gt;. That is the cleanest signal that Context Engine is not just a UI feature. You can pass graph context into programmatic agent work.&lt;/p&gt;
&lt;pre class=&quot;language-bash&quot; tabindex=&quot;0&quot;&gt;&lt;code class=&quot;language-bash&quot;&gt;&lt;span class=&quot;token function&quot;&gt;curl&lt;/span&gt; &lt;span class=&quot;token parameter variable&quot;&gt;--request&lt;/span&gt; POST &lt;span class=&quot;token punctuation&quot;&gt;&#92;&lt;/span&gt;
  &lt;span class=&quot;token parameter variable&quot;&gt;--url&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;&lt;span class=&quot;token variable&quot;&gt;$BASE_URL&lt;/span&gt;/v1/ai/agents/tasks&quot;&lt;/span&gt; &lt;span class=&quot;token punctuation&quot;&gt;&#92;&lt;/span&gt;
  &lt;span class=&quot;token parameter variable&quot;&gt;--header&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;Authorization: Bearer &lt;span class=&quot;token variable&quot;&gt;$MATILLION_TOKEN&lt;/span&gt;&quot;&lt;/span&gt; &lt;span class=&quot;token punctuation&quot;&gt;&#92;&lt;/span&gt;
  &lt;span class=&quot;token parameter variable&quot;&gt;--header&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&quot;Content-Type: application/json&quot;&lt;/span&gt; &lt;span class=&quot;token punctuation&quot;&gt;&#92;&lt;/span&gt;
  &lt;span class=&quot;token parameter variable&quot;&gt;--data&lt;/span&gt; &lt;span class=&quot;token string&quot;&gt;&#39;{
    &quot;message&quot;: &quot;Plan a pipeline for daily net revenue by region using our governed sales model.&quot;,
    &quot;agentConfig&quot;: {
      &quot;name&quot;: &quot;data_engineer_agent&quot;,
      &quot;mode&quot;: &quot;PLAN&quot;,
      &quot;projectId&quot;: &quot;a1b2c3d4-e5f6-7890-abcd-ef1234567890&quot;,
      &quot;sourceBranchName&quot;: &quot;main&quot;,
      &quot;environmentName&quot;: &quot;development&quot;,
      &quot;targetBranchName&quot;: &quot;feature/revenue-region-plan&quot;,
      &quot;graphId&quot;: &quot;finance-analytics-graph&quot;
    }
  }&#39;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Matillion&#39;s &lt;a href=&quot;https://docs.maia.ai/docs/api-reference/using-agent-tasks-api&quot;&gt;Agent Tasks API guide&lt;/a&gt; defines &lt;code&gt;name&lt;/code&gt; as currently always &lt;code&gt;data_engineer_agent&lt;/code&gt;, &lt;code&gt;mode&lt;/code&gt; as either &lt;code&gt;ACT&lt;/code&gt; or &lt;code&gt;PLAN&lt;/code&gt;, and &lt;code&gt;graphId&lt;/code&gt; as the ID of a Knowledge Layer service graph. Use &lt;code&gt;PLAN&lt;/code&gt; when you want the graph to shape a design without letting Maia make changes yet. Use &lt;code&gt;ACT&lt;/code&gt; only when you are comfortable with the branch, environment, and permissions.&lt;/p&gt;
&lt;p&gt;That branch point matters. Each task works in isolation on its own branch when you create work through the API. Matillion also says Agent Tasks API work runs under the identity of the user associated with the API key, and changes are not automatically visible to other project users unless Maia commits and pushes to the target branch. Good. Agent work should leave a diff, not a mystery.&lt;/p&gt;
&lt;p&gt;There is one practical catch: the docs say these endpoints currently work only with Matillion-hosted and GitHub projects. If your team is standardized on another Git provider, test the path before you design a team process around it.&lt;/p&gt;
&lt;h2 id=&quot;how-is-context-engine-different-from-context-files&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine/#how-is-context-engine-different-from-context-files&quot;&gt;&lt;span&gt;How is Context Engine different from context files?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Context files still matter. They are Markdown files that Maia always reads from &lt;code&gt;.matillion/maia/rules/&lt;/code&gt;, and Matillion enforces a &lt;strong&gt;12,000 character limit&lt;/strong&gt; across all Markdown context files in that directory. They are the right place for rules: naming conventions, design standards, business glossary shortcuts, and team preferences.&lt;/p&gt;
&lt;p&gt;Context Engine is for the map.&lt;/p&gt;
&lt;p&gt;Here is the split that should guide your setup:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mechanism&lt;/th&gt;
&lt;th&gt;Best use&lt;/th&gt;
&lt;th&gt;Concrete limit or setting&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Context files&lt;/td&gt;
&lt;td&gt;Always applied rules for a project&lt;/td&gt;
&lt;td&gt;Stored under &lt;code&gt;.matillion/maia/rules/&lt;/code&gt; with a 12,000 character total limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Additional project files&lt;/td&gt;
&lt;td&gt;Detailed standards Maia can reference when instructed&lt;/td&gt;
&lt;td&gt;Stored outside &lt;code&gt;.matillion/...&lt;/code&gt; and referenced from a context file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context Engine knowledge graphs&lt;/td&gt;
&lt;td&gt;Warehouse metadata, relationships, semantics, and execution context&lt;/td&gt;
&lt;td&gt;Fed by Warehouse data and Pipeline execution crawlers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The best setup uses both. Put hard rules in context files: table prefixes, environment constraints, source of truth definitions, and review expectations. Put metadata and process reality in Context Engine: schemas, columns, tags, warehouse structures, and pipeline execution history.&lt;/p&gt;
&lt;p&gt;Do not stuff everything into the graph because it feels newer. If a rule is small, stable, and mandatory, put it in a context file. If the context changes as pipelines run and schemas evolve, put it in the graph.&lt;/p&gt;
&lt;h2 id=&quot;what-does-it-cost-and-where-can-it-surprise-the-bill&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine/#what-does-it-cost-and-where-can-it-surprise-the-bill&quot;&gt;&lt;span&gt;What does it cost, and where can it surprise the bill?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Matillion&#39;s public docs do not publish a separate Context Engine credit rate in the pages reviewed for this guide. That absence is its own buying signal. You should not treat public preview as free forever, and you should not assume every action has a visible line item until you validate it with your account data.&lt;/p&gt;
&lt;p&gt;There are &lt;strong&gt;3 cost surfaces&lt;/strong&gt; to watch.&lt;/p&gt;
&lt;p&gt;First, crawler frequency. A Warehouse data crawler connects to data warehouse structures such as Snowflake databases and schemas, Databricks catalogs and schemas, or Redshift schemas. Even if Matillion does not show a separate Context Engine meter in the public docs, that crawler is still operating against systems you pay to run. Start with a low frequency schedule, then increase it only for domains where schema drift or metadata freshness affects real delivery.&lt;/p&gt;
&lt;p&gt;Second, agent work. Maia can validate and run pipelines, query the data warehouse, sample component output, commit changes, and push branches. Mission Control lets up to &lt;strong&gt;10 tasks&lt;/strong&gt; sit in the In progress column at once. Ten autonomous tasks pointed at a development warehouse can be a productivity win. Ten autonomous tasks repeatedly sampling, running, and revising pipelines can also turn a quiet sandbox into a noisy bill.&lt;/p&gt;
&lt;p&gt;Third, preapproved tools. The Agent Tasks API supports &lt;code&gt;grantedPermissions&lt;/code&gt;, and the Mission Control UI has Ask permission and Bypass permissions modes. Matillion recommends Ask permission as the default and Bypass permissions only for trusted hands-off runs in scoped environments. That is not conservative vendor boilerplate. It is the right default for anyone who has ever watched a retry loop discover money.&lt;/p&gt;
&lt;p&gt;For cost monitoring, Matillion&#39;s MCP server exposes Consumption tools, including &lt;code&gt;get-consumption&lt;/code&gt; for flat-rated products and &lt;code&gt;get-consumption-etl-users&lt;/code&gt; for ETL users, and it can help analyze credit consumption patterns through an AI assistant via the &lt;a href=&quot;https://docs.maia.ai/docs/api-reference/mcp-server&quot;&gt;MCP server documentation&lt;/a&gt;. Use that for investigation, but do not let an assistant be your only FinOps control. Pull consumption on a schedule, tag task branches clearly, and compare before and after you enable crawler schedules.&lt;/p&gt;
&lt;p&gt;A sane rollout looks like this:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Create one restricted knowledge graph for a single analytics domain.&lt;/li&gt;
&lt;li&gt;Add one Warehouse data crawler and one Pipeline execution crawler.&lt;/li&gt;
&lt;li&gt;Schedule crawls outside peak warehouse windows.&lt;/li&gt;
&lt;li&gt;Run Maia tasks in &lt;code&gt;PLAN&lt;/code&gt; mode first.&lt;/li&gt;
&lt;li&gt;Keep Ask permission on until you know which tools Maia calls repeatedly.&lt;/li&gt;
&lt;li&gt;Review consumption and warehouse activity after one week.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The boring version wins.&lt;/p&gt;
&lt;h2 id=&quot;when-should-you-use-mission-control-with-context-engine&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine/#when-should-you-use-mission-control-with-context-engine&quot;&gt;&lt;span&gt;When should you use Mission Control with Context Engine?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Use Mission Control when the work has a deliverable. A chat is fine for asking Maia to explain a Table Input component or suggest a naming convention. A Mission Control task is better when you want Maia to design a pipeline, modify project files, create a connector, analyze a failure, or work from a PDF spec.&lt;/p&gt;
&lt;p&gt;The task board has &lt;strong&gt;4 columns&lt;/strong&gt;: Backlog, In progress, Needs attention, and Completed. That sounds simple because it is. The useful part is that each task has its own chat interface, and you can switch between tasks without losing context. For a data team, that maps better to real work than one endless assistant thread.&lt;/p&gt;
&lt;p&gt;Mission Control also adds attachments. You can attach images and PDFs to a task prompt, then reference them with &lt;code&gt;@filename&lt;/code&gt;. Matillion says images can include diagrams, screenshots, mockups, and whiteboard photos, while PDFs can include specs, requirements documents, and reports. Text files are not supported as attachments because Maia can already read project text files.&lt;/p&gt;
&lt;p&gt;Use that for pipeline triage. A screenshot of a broken canvas plus a pipeline execution crawler is exactly the kind of mixed context that a human engineer would ask for. The difference is Maia can now carry that context into a branch and produce work you can review.&lt;/p&gt;
&lt;p&gt;Do not use Mission Control as a merge gate. Matillion says completing a task does not make changes visible on other branches. You still need to commit, push, and merge. Maia cannot merge changes, which is a good boundary. Keep code review, pipeline tests, and environment promotion in your normal process.&lt;/p&gt;
&lt;h2 id=&quot;what-would-i-do-first-in-a-real-matillion-team&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine/#what-would-i-do-first-in-a-real-matillion-team&quot;&gt;&lt;span&gt;What would I do first in a real Matillion team?&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;Start with one graph per domain, not one graph per company. Finance, customer, product, and operations usually have different semantics and different blast radiuses. A public all-company graph sounds convenient until it contains restricted HR tables or ambiguous definitions that poison every prompt.&lt;/p&gt;
&lt;p&gt;Then tune freshness by risk. Operational pipeline execution context can go stale quickly if you are actively refactoring. Warehouse schema metadata may not need hourly crawls if your governed marts change weekly. Context Engine supports any number of Warehouse data and Pipeline execution crawlers on a graph, so separate them by source and schedule rather than creating one giant crawler that nobody wants to touch.&lt;/p&gt;
&lt;p&gt;Use &lt;code&gt;PLAN&lt;/code&gt; mode as your default for the first few tasks. Ask Maia to propose pipeline structure, sources, joins, variable usage, and tests while grounded in the selected graph. Once the plans look sane, let a narrow &lt;code&gt;ACT&lt;/code&gt; task implement on a feature branch.&lt;/p&gt;
&lt;p&gt;Most of all, measure whether Context Engine reduces clarification loops. The win is not that Maia can produce more pipeline files. The win is fewer wrong assumptions about &lt;code&gt;customer_id&lt;/code&gt;, fewer duplicated staging models, fewer prompts that repeat your business glossary, and fewer reviews that start with: who told the agent to use that table?&lt;/p&gt;
&lt;p&gt;If Context Engine does that, it earns a place in the stack. If it becomes another metadata garden that nobody prunes, Maia will learn your mess at machine speed.&lt;/p&gt;
&lt;h2 id=&quot;sources&quot; tabindex=&quot;-1&quot;&gt;&lt;a class=&quot;header-anchor&quot; href=&quot;https://data-today.net/matillion/matillion-context-engine/#sources&quot;&gt;&lt;span&gt;Sources&lt;/span&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/context-engine&quot;&gt;Matillion Maia docs: Context Engine&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/mission-control&quot;&gt;Matillion Maia docs: Mission Control&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/api-reference/using-agent-tasks-api&quot;&gt;Matillion Maia docs: Using the Agent Tasks API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/maia-ai-agents-overview&quot;&gt;Matillion Maia docs: Maia AI Agents overview&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/guides/maia-context-files&quot;&gt;Matillion Maia docs: Context files&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/api-reference/mcp-server&quot;&gt;Matillion Maia docs: Matillion MCP server&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.maia.ai/docs/changelog/2026-changelog&quot;&gt;Matillion Maia changelog: 2026 changelog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
</content>
  </entry>
</feed>