When you need this
A Job can fail, a Pod can remain unscheduled and a long previous run can block the next one. kubectl helps diagnose the cluster but does not replace a scheduled expectation of results. An external monitor can detect missing signals even when the cluster is unavailable.
How to connect
- 1
Create a cron monitor with the same schedule and timeZone. Allow for Pod scheduling and normal job duration.
- 2
Add the ping URL to a Secret named tickwatch, key ping-url, through your normal secret-management process.
- 3
Replace the image and job command in the example. Check success, failure and a missed run on a test CronJob.
Ready-to-use example
The image must contain /bin/sh, curl and your /app/run-job.sh script. Allow outgoing HTTPS to the ping endpoint. Adapt this example to your existing manifest, network policies and Secret permissions.
apiVersion: batch/v1
kind: CronJob
metadata:
name: nightly-report
spec:
schedule: "0 3 * * *"
timeZone: "Etc/UTC"
concurrencyPolicy: Forbid
jobTemplate:
spec:
backoffLimit: 1
template:
spec:
restartPolicy: Never
containers:
- name: job
image: your-registry/your-job:version
env:
- name: PING
valueFrom:
secretKeyRef:
name: tickwatch
key: ping-url
- name: RUN_ID
valueFrom:
fieldRef:
fieldPath: metadata.uid
command: ["/bin/sh", "-c"]
args:
- |
signal() {
curl -fsS --connect-timeout 3 --max-time 5 \
-X POST "$PING/$1?rid=$RUN_ID" >/dev/null 2>&1 || true
}
signal start
/app/run-job.sh
result=$?
if [ "$result" -eq 0 ]; then signal ""; else signal fail; fi
exit "$result"Check before going live
Forbid prevents concurrent runs of the same CronJob and can cause a missed next run. Configure the monitor’s maximum runtime.
OOMKill or forced Pod termination may prevent a failure signal. Monitor missing completions as well.
A Secret controls how the key is supplied through Kubernetes; it does not replace RBAC or secret protection in the cluster.
Common questions
Does Tickwatch read the Kubernetes API?
Not in this setup. The container sends its own signals. Tickwatch does not receive access to the cluster or its Secrets.
What happens when a Job retries?
Each new Pod uses its own UID as a run identifier. An initial failure and successful retry can produce an incident followed by recovery; account for that in alert settings.