Databricks Data Additions
(As of Feb 2026)
DfE’s Databricks instance is an analytical data store — it doesn’t sit behind any operational service. Instead, data from operational DfE systems (such as COLLECT) and external sources is replicated into Databricks via configured mirrors. The built-in data catalogue shows what’s currently available.
Requesting a New Dataset
If the data you need isn’t in the catalogue but exists in a DfE database, contact the Analytical Data Access (ADA) team — they manage the Databricks service and configure new data mirrors.
Email ADA.SUPPORT@education.gov.uk with the following:
- Database, schema, and table identifiers for the data you want added
- A completed Data Supply Agreement (DSA), submitted via the DSA portal
To complete the DSA, the dataset must be registered on the Information Asset Register, and you’ll need approval from the relevant Information Asset Officer (IAO). ADA facilitate access but are not the data owners.
Once approved, ADA will set up an ADF pipeline that refreshes the data approximately daily.
Manual Uploads
Smaller or static datasets don’t need a formal mirror. These can be uploaded directly to Databricks without involving ADA. Bear in mind that manually uploaded datasets have no automated refresh — whoever uploads them is responsible for keeping them up to date.