Service Status and Diagnostics
Health endpoints, container diagnostics, background jobs, and common operational failure modes.
Service Status and Diagnostics
This page covers assessing the health of a running VeriWorkly deployment and diagnosing failures.
1. Health endpoints
There are two health routes, and the distinction matters.
GET /api/v1/health — liveness
A lightweight check that touches no external dependencies. It confirms the process is up and responding, nothing more. This is the endpoint the Docker health check and any uptime monitor should use, precisely because it does not wake serverless database compute on every poll.
{
"success": true,
"message": "Server is healthy",
"data": {
"status": "ok",
"timestamp": "2026-07-29T10:00:00.000Z"
}
}GET /api/v1/health/ready — readiness
A comprehensive check that actively probes PostgreSQL (SELECT 1) and Redis (PING). Use it for
deploy-time verification and manual diagnosis.
Healthy (200 OK)
{
"success": true,
"message": "Server is healthy",
"data": {
"status": "ok",
"database": "connected",
"redis": "connected",
"timestamp": "2026-07-29T10:00:00.000Z"
}
}Degraded (503 Service Unavailable)
{
"success": false,
"message": "Service Unavailable"
}`/health/ready` fails whole, not partially
If either PostgreSQL or Redis is unreachable, the endpoint returns a plain 503 error response.
It does not report which dependency failed — check the API logs to identify the culprit.
Recommendation: point uptime monitoring (UptimeRobot, Better Stack, Datadog) at /health, and
run /health/ready as a post-deploy gate.
2. Container diagnostics
Service state
docker compose psCompose defines four services: web (the marketing site), api, redis, and redisinsight (a
Redis GUI on port 5540 — a debugging convenience that should not be publicly exposed).
Resource usage
docker statsMemory growth on the api container is most often a Redis-backed structure growing unbounded, or
concurrent AI/ATS requests holding large extracted-document text in memory while waiting on an
external LLM provider.
Document rendering does not load the API
PDF and DOCX generation happens entirely in the browser. Export volume puts no load on the API
container. If you are chasing memory pressure, look at AI/ATS concurrency, uploaded-file parsing
(pdf-parse / mammoth), and Redis — not document rendering.
3. Logs
Run from the directory containing compose.yaml.
API logs
docker compose logs --tail=100 -f apiWatch for: startup configuration-validator failures, Prisma connectivity errors, Redis connection errors, and cron-job lock messages.
Redis logs
bash docker compose logs --tail=100 -f redis Watch for: memory allocation failures and
AOF persistence errors.
Marketing site logs
docker compose logs --tail=100 -f webWatch for: SSR fetch failures and internal network resolution errors (ECONNREFUSED),
usually meaning BACKEND_INTERNAL_URL is wrong.
4. Background jobs
Five node-cron jobs run inside the API. Each acquires a Redis distributed lock first, so only
one worker executes a given job even when the server is clustered across CPU cores.
| Job | Schedule | Purpose |
|---|---|---|
| GitHub sync | GITHUB_SYNC_CRON, default 0 0,12 * * * (twice daily) | Refreshes repository issue and PR statistics. Skipped when existing data is under 12 hours old. |
| Changelog release sync | CHANGELOG_RELEASE_SYNC_CRON, default 0 6 * * * (daily) | Imports missing GitHub Releases into the public changelog. |
| Views flush | every 10 minutes | Moves buffered portfolio and share-link view counts from Redis into PostgreSQL. |
| Usage metrics flush | USAGE_METRICS_FLUSH_CRON, default 10 0 * * * | Flushes buffered usage counters into UsageMetricDaily, only for fully-completed UTC days. |
| Portfolio access | hourly | Suspends portfolios whose billing grace period has expired, and garbage-collects abandoned image uploads (and their R2 objects) that have been PENDING for more than 24 hours. |
Both flush jobs write idempotency-batch markers, so a crashed or racing flush can neither double-count nor silently drop a batch.
Missed cron windows during downtime are partially covered: the GitHub sync, changelog sync, and usage-metrics jobs each run a catch-up check on startup.
5. Common failure modes
| Symptom | Likely cause | Resolution |
|---|---|---|
| API exits immediately on boot | The production configuration validator rejected the environment. | Read the startup log — it names the failing variable. Common causes: a default AUTH_SECRET, a non-HTTPS AUTH_BASE_URL, API_KEY_HASH_SECRET equal to AUTH_SECRET, or AUTH_EMAIL_PROVIDER still set to console. |
/health/ready returns 503 | PostgreSQL or Redis unreachable. | Verify DATABASE_URL and REDIS_URL, and that the API can reach both across the network. |
Login fails but /health is fine | Redis is down. Better-Auth stores sessions in Redis secondary storage. | Restore Redis. The rate limiter has an in-memory fallback; session storage does not. |
| CORS errors in the browser | ALLOWED_ORIGINS does not include the public frontend origin. | Add the exact scheme + host + port. Credentials are only granted to explicitly listed origins, never wildcard-matched portfolio subdomains. |
| AI generation times out | Upstream Anthropic/OpenAI latency, or a dangling credit reservation. | Check the provider status page. Verify the request's credit reservation was released rather than left reserved — reserved credits are not spendable until committed or released. |
| API keys stop working after a payment event | Expected. Keys are invalidated when the owning subscription lapses to cancelled or inactive. | Restore the subscription, or issue keys from an account with an active entitlement. |
| Portfolio goes offline unexpectedly | The billing grace period expired and the hourly portfolio access job suspended it. | Restore the subscription; publication status moves back out of SUSPENDED. |
| Webhook events appear unprocessed | Check BillingWebhookEvent for rows in FAILED or PROCESSING. | Each row records its error and retry count, keyed by the provider's event ID for idempotent reprocessing. |