Automated Databricks Job Failure Digest
Every morning, WebRun opens Databricks, checks overnight job runs against their usual status and duration, saves the error log for any failure to Drive, and posts a Slack summary of what broke.
How do I get a daily digest of failed Databricks job runs?
WebRun checks your Databricks job runs every morning, saves the error log for any failure straight to Google Drive, and posts a Slack summary naming every job that failed overnight. Each Slack entry links to its saved log, so the data engineering team starts the day knowing exactly what broke and where to look.
- Failed jobs get caught the same morning instead of at the next run
- Every failure log is saved and linked, no manual digging in Databricks
- The team starts the day with one Slack summary instead of checking each job
Built for data engineering teams · data platform admins · machine learning engineers · analytics engineers
What does WebRun do on every run?
The exact actions WebRun takes, in order - in plain language, so you can adjust anything.
-
WebRun signs in and gets to work
Opens
databricks.comin a real browser with your saved login - no setup, no API keys. -
1
Databricks - check job run status
WebRun opens Databricks to check job run status. - Open the Databricks Jobs page and list overnight job runs
- Check each run's status and duration against its usual pattern
- Flag any run that failed or timed out
Done when Every overnight job run has a checked status.
-
2
Google Drive - save failed run logs
WebRun opens Google Drive to save failed run logs. - Save the error log for each failed run to the shared Drive folder
- Name the file with the job and run timestamp
Done when Every failed run's log is saved to Drive.
-
3
Slack - post the failure summary
WebRun posts to Slack to share the failure summary. - Post a summary of failed jobs to the data engineering channel
- Link each entry to its saved log in Drive
Done when The team has this morning's failure digest in Slack.
How is each run configured?
Secure by default
Connect once, stays signed in
WebRun signs in once and keeps each session in a persistent environment, so every run picks up right where it left off.
Every action is checked against this policy before it runs.
Questions, answered
Will it re-run a failed job automatically?
No. WebRun only reports the failure and saves the log. Re-running or debugging the job is left to a human.
How does it decide a run failed?
It checks each job's run status and duration in Databricks directly, so the digest reflects the platform's own pass or fail result, not a guess.
Where do the error logs go?
Each failed run's log is saved straight to your shared Drive folder, named with the job and run timestamp, so it's easy to open from the Slack summary.
Put this on autopilot.
Turn it on in minutes - or have our team set it up for you.