
The day my monitoring system stopped monitoring
It was a typical Thursday afternoon when I noticed something alarming: our Uptime Kuma monitoring dashboard was barely loading. Pages took over 30 seconds to render, and monitors weren't updating.
The irony wasn't lost on me— our monitoring system, designed to alert us about problems, had become the problem.
What followed was a deep dive into database optimization that uncovered some shocking truths about SQLite at scale and taught me valuable lessons about Kubernetes deployments. Here's the complete story.
The Symptoms: When Everything Feels Broken
Our Uptime Kuma instance was exhibiting multiple concerning behaviors:
UI loading times: 10–30 seconds (sometimes timing out completely) Monitor status: Stuck in "Pending" state or showing sporadic timeouts Error logs: Flooded with connection pool exhaustion errors
The error that kept appearing in logs was particularly troubling:
KnexTimeoutError: Knex: Timeout acquiring a connection.
The pool is probably full. Are you missing a .transacting(trx) call?
At first, I thought it was a networking issue. Maybe the AWS EKS cluster was having problems? But as I dug deeper, the real culprit revealed itself to be far more insidious.
The Investigation: Following the Breadcrumbs
Step 1: Checking the Basics
First, I verified our infrastructure setup:
# Check pod status
kubectl get pods -n monitoring -l app.kubernetes.io/name=uptime-kuma
# Check resource usage
kubectl top pod -n monitoring -l app.kubernetes.io/name=uptime-kuma
Everything looked normal. The pod was running, CPU was at 80–100%, memory around 800MB — nothing screaming "I'm the problem!"
Step 2: Storage Configuration
I suspected storage might be the issue. Uptime Kuma uses SQLite, and SQLite is notoriously sensitive to storage type.
kubectl get pvc uptime-kuma-pvc -n monitoring -o yaml
Relief washed over me:
Storage Class: gp3 (AWS EBS, not NFS) Access Mode: ReadWriteOnce Status: Bound
So it wasn't a storage misconfiguration. The mystery deepened.
Step 3: The Database Investigation
I exec'd into the pod to check the database:
kubectl exec -it -n monitoring <pod-name> -- /bin/sh
#First, let's check SQLite settings:
sqlite3 /app/data/kuma.db "PRAGMA cache_size; PRAGMA journal_mode; PRAGMA synchronous;"
# Output:
-2000 # Cache: 2MB (seems small)
wal # Write-Ahead Logging (good)
2 # FULL synchronous (very safe, but slow)
The cache was tiny. But surely a 2MB cache couldn't cause such dramatic slowdowns?
Then I checked the database size:
ls -lh /app/data/kuma.db
My jaw dropped:
-rwxrwxr-x. 1 root root 5.0G Jan 15 17:06 /app/data/kuma.db
5 GIGABYTES.
For context, a healthy Uptime Kuma database with 150 monitors should be around 50–200MB. Mine was 25–100x larger than normal.
Step 4: The Smoking Gun
I needed to know what was eating all that space:
sqlite3 /app/data/kuma.db "SELECT COUNT(*) FROM heartbeat;"
# I waited. And waited. The query was taking forever. Finally, after what felt like an eternity:
33546184
Thirty-three and a half MILLION heartbeat records.
For 152 monitors, even with years of history, you'd expect 500,000 to 2 million records max. I had 67 times more data than necessary.
Every monitor check writes a heartbeat. Every status change writes a heartbeat. With 152 monitors checking every 30–60 seconds, that's potentially 250,000 writes per day. But clearly, nothing had ever been cleaned up.
The database had been accumulating data since deployment — probably years' worth — and SQLite was choking on the massive table scans required for even simple queries.
The Root Causes: A Perfect Storm
Now that I understood what was wrong, I needed to figure out why. I discovered multiple compounding issues:
1. No Data Retention Policy
Uptime Kuma has a setting called "Keep All Stats Duration" that controls how long historical data is retained. It was set to unlimited.
Every heartbeat since the beginning of time was still in the database.
2. Aggressive Monitor Settings
Looking at our monitor configurations:
sqlite3 /app/data/kuma.db "SELECT id, name, timeout, interval FROM monitor WHERE active=1 LIMIT 5;"
The results were alarming:
Average timeout: 24.8 seconds (should be 10–15s max) Check intervals: 30–60 seconds (too frequent for 152 monitors) Retry attempts: 2 per failure (doubles the failed check load)
With 152 monitors, this meant:
Up to 304 concurrent checks during failures (152 monitors × 2 retries) ~5,000+ database writes per minute during peak times Each 24-second timeout holds a database connection
SQLite's connection pool was drowning.
3. Insufficient Resources
Our Helm values showed:
resources:
limits:
cpu: 1000m # 1 CPU core
memory: 1Gi # 1GB RAM
For 152 monitors hammering a 5GB database, this was woefully inadequate.
4. SQLite Configuration
The environment variables we'd set weren't being applied correctly:
env | grep UPTIME_KUMA
UPTIME_KUMA_DB_CACHE_SIZE=-32000 # Set to 32MB
But SQLite was still using:
-2000 # Only 2MB
Why? Because SQLite PRAGMA settings are per-connection, and Uptime Kuma wasn't applying them on connection initialization.
The Solution: Surgery in Production
With the diagnosis complete, I had a plan. It was risky — modifying a production database always is — but the system was already unusable.
Phase 1: Optimize Monitor Configuration
First, I tackled the low-hanging fruit: monitor settings.
# Reduce timeouts from 24.8s to 10s
sqlite3 /app/data/kuma.db "UPDATE monitor SET timeout = 10 WHERE timeout > 10;"
# Reduce retry attempts from 2 to 1
sqlite3 /app/data/kuma.db "UPDATE monitor SET maxretries = 1 WHERE maxretries > 1;"
# Increase Dev monitor intervals to reduce load
sqlite3 /app/data/kuma.db "UPDATE monitor SET interval = 120 WHERE name LIKE 'Dev.%' AND interval < 120;"
This immediately reduced the concurrent connection pressure.
Phase 2: The Nuclear Option — Delete All Heartbeats
I made a backup, took a deep breath, and executed:
time sqlite3 /app/data/kuma.db "DELETE FROM heartbeat;"
I watched the cursor blink. And blink. And blink some more.
13 minutes and 28 seconds later, it was completed. All 33.5 million records were gone.
But the database file was still 5GB! Why? Because SQLite doesn't automatically reclaim space — it just marks the space as available for reuse. I needed to VACUUM.
Phase 3: The VACUUM
Here's where it got tricky. VACUUM requires:
No active connections to the database Enough disk space to create a temporary copy Time (lots of it)
I scaled down Uptime Kuma:
kubectl scale deployment uptime-kuma -n monitoring --replicas=0
Then created a maintenance pod with the PVC mounted:
kubectl run vacuum-job -n monitoring --rm -i --tty \
--image=louislam/uptime-kuma:1.23.13-debian \
--overrides='...(volume mounts)...'
Inside the maintenance pod:
echo "=== BEFORE VACUUM ==="
ls -lh /app/data/kuma.db
# -rwxrwxr-x. 1 root root 5.0G
time sqlite3 /app/data/kuma.db "VACUUM;"
# Wait 6 minutes...
echo "=== AFTER VACUUM ==="
ls -lh /app/data/kuma.db
# -rwxrwxr-x. 1 root root 2.7M

From 5.0GB to 2.7MB.
A 1,850x reduction.
I literally gasped when I saw it.
Phase 4: Resource and Configuration Optimization
With the database healthy again, I updated our Helm values:
image:
tag: "1.23.13-debian"
podEnv:
- name: "NODE_OPTIONS"
value: "--max-old-space-size=4096"
- name: "UPTIME_KUMA_DB_CACHE_SIZE"
value: "-32000"
- name: "UPTIME_KUMA_DB_MAX_POOL"
value: "15"
- name: "UPTIME_KUMA_PORT"
value: "3001"
resources:
limits:
cpu: 2000m # Doubled from 1000m
memory: 2Gi # Doubled from 1Gi
requests:
cpu: 1000m
memory: 1Gi
livenessProbe:
initialDelaySeconds: 300
periodSeconds: 60
failureThreshold: 10
timeoutSeconds: 10
readinessProbe:
initialDelaySeconds: 60
periodSeconds: 30
failureThreshold: 10
volume:
enabled: true
size: 20Gi
storageClassName: "gp3"
I applied the changes:
helm upgrade uptime-kuma ./chart -n monitoring -f values.yaml --wait
Phase 5: Configure Data Retention
Finally, in the Uptime Kuma UI:
Settings → General → "Keep All Stats Duration": Set to 60 days

This ensures the problem never happens again.
The UI was snappy again. Monitors are updated in real-time. The team's confidence in our monitoring was restored.
Most importantly: no more false alerts.
Lessons Learned: What This Taught Me
1. SQLite Has Limits — Know Them
SQLite is excellent for many use cases, but it struggles with:
Massive tables (>10M rows) High concurrent writes (>100/sec) Large database files (>1GB)
For 150+ monitors with aggressive check intervals, SQLite was at its limit. PostgreSQL would handle this workload without breaking a sweat.
2. Data Retention Isn't Optional
Every application that stores time-series data needs a retention policy. Period.
For Uptime Kuma:
30 days: Minimal, for small deployments 60 days: Recommended for most 90 days: Maximum for large deployments Unlimited: Never, ever, ever
3. Monitor Your Monitoring
The irony of our monitoring system failing without us noticing isn't lost on me. Now I have:
Monthly CronJob to check database size Alerts on Uptime Kuma pod restarts Automated database cleanup
4. Resource Allocation Matters
We had allocated resources based on the initial deployment, not growth over time. As monitors were added and data accumulated, we never increased resources.
Now I review resource allocation quarterly.
5. VACUUM Is Not Optional for SQLite
SQLite's VACUUM command:
Rebuilds the database file from scratch Reclaims unused space Defragments and optimizes
For any SQLite application in production, schedule regular VACUUM operations.
6. Connection Pool Settings Are Critical
The default connection pool size was insufficient. Increasing it to 15 concurrent connections eliminated the timeout errors.
7. When to Migrate to PostgreSQL
We're now planning a migration to PostgreSQL because:
Auto-VACUUM maintains itself Better concurrent connection handling Built-in connection pooling Scales to 1000+ monitors easily More sophisticated query optimization
Prevention: Never Let This Happen Again
I created a Kubernetes CronJob for automated maintenance:
apiVersion: batch/v1
kind: CronJob
metadata:
name: uptime-kuma-db-cleanup
namespace: monitoring
spec:
schedule: "0 2 1 * *" # 1st of every month at 2 AM
jobTemplate:
spec:
template:
spec:
containers:
- name: cleanup
image: louislam/uptime-kuma:1.23.13-debian
command:
- /bin/sh
- -c
- |
echo "Deleting heartbeats older than 90 days..."
sqlite3 /app/data/kuma.db "DELETE FROM heartbeat WHERE time < strftime('%s', 'now', '-90 days');"
echo "Running VACUUM..."
sqlite3 /app/data/kuma.db "PRAGMA optimize; VACUUM;"
echo "Cleanup complete!"
ls -lh /app/data/kuma.db
volumeMounts:
- name: storage
mountPath: /app/data
volumes:
- name: storage
persistentVolumeClaim:
claimName: uptime-kuma-pvc
This runs monthly and prevents database bloat before it becomes a problem.
The Bigger Picture: Technical Debt Is Real
This incident wasn't caused by a single mistake. It was the result of:
Not configuring data retention at deployment Not monitoring resource usage over time Not understanding SQLite's limitations Not implementing preventive maintenance
In other words: technical debt.
We deployed Uptime Kuma, it worked, and we forgot about it. Over months, that debt accumulated — quite literally, as 33.5 million database records.
The lesson? Even "set it and forget it" systems need periodic review and maintenance.
Conclusion: A Happy Ending
What started as a crisis ended as an optimization success story. Our monitoring system is now:
Fast: UI loads in 1–2 seconds Reliable: No connection errors or false alerts Efficient: Using 60% less CPU and memory Sustainable: Automated maintenance prevents future issues
The 13 minutes watching that DELETE statement run were some of the longest of my career. But the 6 minutes watching the VACUUM reclaim 5GB of space were incredibly satisfying.
If you're running Uptime Kuma (or any SQLite-based application) at scale, I hope this story helps you avoid the same mistakes I made.
And if your database is suspiciously large, maybe it's time to check how many heartbeats you're storing.
Key Takeaways
Set data retention policies on day one, not when there's a problem Monitor your database size — growth over time is a red flag VACUUM regularly for SQLite applications Scale resources with data growth, not just initial requirements Consider PostgreSQL for applications with 100+ concurrent operations Implement automated maintenance — future you will thank present you Review "set and forget" systems at least quarterly
Resources
If you found this helpful, here are some resources:
Uptime Kuma GitHub SQLite VACUUM documentation Kubernetes CronJob documentation