feature: Observability: Prometheus Metrics and Grafana
This commit is contained in:
@@ -0,0 +1,315 @@
|
||||
# Docker & Docker Compose Guide (VigilCare)
|
||||
|
||||
This guide is a practical reference for running this project with Docker, plus troubleshooting for common issues seen in this repo.
|
||||
|
||||
It is written for junior developers, so each section explains not just what to do, but why.
|
||||
|
||||
---
|
||||
|
||||
## 1) Quick start
|
||||
|
||||
From repo root:
|
||||
|
||||
```bash
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
Check status:
|
||||
|
||||
```bash
|
||||
docker compose ps
|
||||
```
|
||||
|
||||
Stop everything (keep data):
|
||||
|
||||
```bash
|
||||
docker compose stop
|
||||
```
|
||||
|
||||
Start again:
|
||||
|
||||
```bash
|
||||
docker compose start
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2) Core concepts (simple mental model)
|
||||
|
||||
### Container
|
||||
- A running process with its own filesystem and network namespace.
|
||||
- Example: `vigilcare_prometheus` is one container.
|
||||
|
||||
### Service (in `docker-compose.yml`)
|
||||
- A recipe for how to run a container.
|
||||
- Example: the `prometheus:` section defines image, ports, volumes, networks.
|
||||
|
||||
### Image
|
||||
- Template used to create a container.
|
||||
- Example: `prom/prometheus:v2.52.0`.
|
||||
|
||||
### Volume
|
||||
- Persistent storage managed by Docker.
|
||||
- Survives container restarts/recreates.
|
||||
- Example: `prometheus_data`, `grafana_data`, `pg_data`.
|
||||
|
||||
### Network
|
||||
- Virtual network connecting containers.
|
||||
- Containers can reach each other by service name (DNS).
|
||||
- Example: Grafana reaches Prometheus at `http://prometheus:9090` inside Docker.
|
||||
|
||||
---
|
||||
|
||||
## 3) Host ports vs container ports
|
||||
|
||||
In Compose, this format is used:
|
||||
|
||||
```yaml
|
||||
ports:
|
||||
- "HOST:CONTAINER"
|
||||
```
|
||||
|
||||
Example from this project:
|
||||
- Prometheus: `"9101:9090"`
|
||||
- Open in browser with `http://localhost:9101`
|
||||
- Inside Docker, service still listens on `9090`
|
||||
- Grafana: `"3101:3000"`
|
||||
- Open in browser with `http://localhost:3101`
|
||||
|
||||
If a UI is not loading, first verify host port mappings in `docker-compose.yml`.
|
||||
|
||||
---
|
||||
|
||||
## 4) Project networking (`vigilcare_net`)
|
||||
|
||||
This project uses a user-defined bridge network:
|
||||
|
||||
```yaml
|
||||
networks:
|
||||
vigilcare_net:
|
||||
driver: bridge
|
||||
```
|
||||
|
||||
All services should join it:
|
||||
|
||||
```yaml
|
||||
networks:
|
||||
- vigilcare_net
|
||||
```
|
||||
|
||||
Why this matters:
|
||||
- Service-to-service DNS works (`prometheus`, `grafana`, `postgres`, etc.).
|
||||
- Keeps local environment predictable.
|
||||
|
||||
### Important Linux note: `host.docker.internal`
|
||||
|
||||
Prometheus scrapes the API via host address in this project:
|
||||
- `http://host.docker.internal:5270/metrics`
|
||||
|
||||
On Linux, `host.docker.internal` may not resolve by default.
|
||||
Fix by adding this to the `prometheus` service:
|
||||
|
||||
```yaml
|
||||
extra_hosts:
|
||||
- "host.docker.internal:host-gateway"
|
||||
```
|
||||
|
||||
Then recreate Prometheus:
|
||||
|
||||
```bash
|
||||
docker compose up -d --force-recreate prometheus
|
||||
```
|
||||
|
||||
Symptom when missing:
|
||||
- Prometheus target `vigilcare_api` is `down`
|
||||
- Error: `lookup host.docker.internal ... no such host`
|
||||
|
||||
### Important: API must listen on all interfaces (not only `localhost`)
|
||||
|
||||
After `extra_hosts` is fixed, Prometheus may still show:
|
||||
|
||||
```
|
||||
dial tcp 172.17.0.1:5270: connect: connection refused
|
||||
```
|
||||
|
||||
**Why:** `dotnet run` with `http://localhost:5270` binds only to `127.0.0.1`.
|
||||
`localhost` inside a container means the container itself, not your host machine.
|
||||
Prometheus inside Docker reaches the host via the gateway IP (`172.17.0.1` via `host.docker.internal`), not host loopback.
|
||||
|
||||
**Check binding:**
|
||||
|
||||
```bash
|
||||
ss -tlnp | rg ':5270'
|
||||
```
|
||||
|
||||
If you see `127.0.0.1:5270`, Prometheus cannot scrape from Docker.
|
||||
|
||||
**Fix (local dev):** bind on all interfaces in `VigilCareClinicalAPI/Properties/launchSettings.json`:
|
||||
|
||||
```json
|
||||
"applicationUrl": "http://0.0.0.0:5270"
|
||||
```
|
||||
|
||||
Or start the API with:
|
||||
|
||||
```bash
|
||||
ASPNETCORE_URLS=http://0.0.0.0:5270 dotnet run --project VigilCareClinicalAPI
|
||||
```
|
||||
|
||||
Then restart the API and confirm:
|
||||
|
||||
```bash
|
||||
ss -tlnp | rg ':5270' # should show 0.0.0.0:5270
|
||||
curl -sS http://localhost:5270/metrics | head
|
||||
```
|
||||
|
||||
**Security note:** `0.0.0.0` is fine for local development. In production, bind explicitly and use proper network controls.
|
||||
|
||||
---
|
||||
|
||||
## 5) Volumes and persistence
|
||||
|
||||
This repo uses named volumes for persistent data:
|
||||
- `pg_data`
|
||||
- `seq_data`
|
||||
- `kafka_data`
|
||||
- `es_data`
|
||||
- `minio_data`
|
||||
- `prometheus_data`
|
||||
- `grafana_data`
|
||||
|
||||
### Why your data still exists after restart
|
||||
- `docker compose up -d --force-recreate` recreates containers, but volumes remain.
|
||||
- This is expected and usually desired.
|
||||
|
||||
### Full reset (destructive)
|
||||
If you need a totally clean environment:
|
||||
|
||||
```bash
|
||||
docker compose down -v
|
||||
```
|
||||
|
||||
Warning:
|
||||
- `-v` removes named volumes (database/log/index data lost).
|
||||
|
||||
---
|
||||
|
||||
## 6) Common commands and when to use them
|
||||
|
||||
### Apply config change to one service
|
||||
Use when you changed only one section (e.g., Prometheus `extra_hosts`):
|
||||
|
||||
```bash
|
||||
docker compose up -d --force-recreate prometheus
|
||||
```
|
||||
|
||||
### Restart service without recreate
|
||||
Use when config did not change and you just want a restart:
|
||||
|
||||
```bash
|
||||
docker compose restart prometheus
|
||||
```
|
||||
|
||||
### Rebuild image service
|
||||
Use when Dockerfile/app code in image changed:
|
||||
|
||||
```bash
|
||||
docker compose up -d --build <service>
|
||||
```
|
||||
|
||||
### View service logs
|
||||
|
||||
```bash
|
||||
docker compose logs -f prometheus
|
||||
docker compose logs -f grafana
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7) Troubleshooting playbook
|
||||
|
||||
### A) “Service is up but endpoint won’t open”
|
||||
1. Check container state:
|
||||
```bash
|
||||
docker compose ps
|
||||
```
|
||||
2. Verify port mapping in `docker-compose.yml`.
|
||||
3. Check logs:
|
||||
```bash
|
||||
docker compose logs --tail=100 <service>
|
||||
```
|
||||
|
||||
### B) “Prometheus healthy, but target is DOWN”
|
||||
1. Open Prometheus targets page:
|
||||
- `http://localhost:9101/targets`
|
||||
2. Read the exact `lastError`.
|
||||
3. If error mentions `host.docker.internal` on Linux:
|
||||
- add `extra_hosts` fix (section 4)
|
||||
- recreate Prometheus.
|
||||
4. If error is `connection refused` to `172.17.0.1:5270`:
|
||||
- API is likely bound to `127.0.0.1` only
|
||||
- use `http://0.0.0.0:5270` and restart API (section 4).
|
||||
|
||||
### C) “Docker compose command cannot connect to daemon”
|
||||
Example:
|
||||
- `failed to connect to the docker API at unix:///var/run/docker.sock`
|
||||
|
||||
Fix:
|
||||
- Start Docker Desktop / Docker daemon.
|
||||
- Re-run `docker compose ps`.
|
||||
|
||||
### D) “Permission denied writing files under bind-mounted folder”
|
||||
This can happen when directories/files were created as `root`.
|
||||
|
||||
Symptoms:
|
||||
- Cannot create/edit files in folders like Grafana dashboard path.
|
||||
|
||||
Fix options:
|
||||
1. Correct ownership on host:
|
||||
```bash
|
||||
sudo chown -R $USER:$USER <folder>
|
||||
```
|
||||
2. Recreate problematic directory as your user.
|
||||
|
||||
### E) “Script fails preflight even though services seem running”
|
||||
Check these directly:
|
||||
|
||||
```bash
|
||||
curl -sS -o /dev/null -w "%{http_code}\n" http://localhost:5270/api/v1/alert-thresholds
|
||||
curl -sS -o /dev/null -w "%{http_code}\n" http://localhost:9101/-/healthy
|
||||
curl -sS -o /dev/null -w "%{http_code}\n" http://localhost:5345
|
||||
```
|
||||
|
||||
Expected: `200`, `200`, `200`.
|
||||
|
||||
---
|
||||
|
||||
## 8) Current project-specific paths and notes
|
||||
|
||||
- Compose file: `docker-compose.yml`
|
||||
- Prometheus config: `infra/prometheus/prometheus.yml`
|
||||
- Grafana provisioning:
|
||||
- `infra/grafana/provisioning/datasources/prometheus.yml`
|
||||
- `infra/grafana/provisioning/datasources/dashboards/config.yml`
|
||||
- Grafana dashboards expected path:
|
||||
- `infra/grafana/dashboards/`
|
||||
|
||||
Note: if you accidentally create a typo folder like `dashbpards`, Grafana provisioning will not load dashboards from it.
|
||||
|
||||
---
|
||||
|
||||
## 9) Safe workflow for config changes (recommended)
|
||||
|
||||
1. Edit `docker-compose.yml`.
|
||||
2. Recreate only changed services:
|
||||
```bash
|
||||
docker compose up -d --force-recreate <service>
|
||||
```
|
||||
3. Verify logs and health endpoints.
|
||||
4. Run project verification scripts (example):
|
||||
```bash
|
||||
./scripts/run-phase8-verification.sh
|
||||
```
|
||||
|
||||
This avoids unnecessary full resets and speeds up local development.
|
||||
|
||||
@@ -350,7 +350,7 @@ sirs:{encounterId}:WBC_K_UL → "1" (TTL: 30 minutes)
|
||||
|
||||
### 8. RabbitMQ Notification Workers and Escalation
|
||||
|
||||
**Description:** The notification worker reads `alert.generated` events from Kafka and dispatches paging jobs to RabbitMQ. The RabbitMQ consumer sends the page and waits for acknowledgment. If no acknowledgment arrives within five minutes, the dead-letter queue escalates to the on-call backup.
|
||||
**Description:** The notification worker reads `alert.generated` events from Kafka and dispatches paging jobs to RabbitMQ. The RabbitMQ consumer sends the page and waits for acknowledgment. If no acknowledgment arrives within five minutes, the dead-letter queue escalates to the on-call backup. If the API host is stopping while a page is in flight, the cancellation path requeues the message rather than escalating it.
|
||||
|
||||
**Exchange topology:**
|
||||
```
|
||||
@@ -373,6 +373,10 @@ clinical.notifications.exchange (direct)
|
||||
c. After TTL: message routes back to alerts.escalation.queue
|
||||
d. Escalation worker pages the on-call backup
|
||||
e. clinical_alert.status → 'escalated' in PostgreSQL
|
||||
5. If host shutdown occurs during paging wait:
|
||||
a. Cancellation is treated as graceful stop, not failure
|
||||
b. NACK with requeue=true → message returns to alerts.paging.queue
|
||||
c. No DLQ route, so no false escalation during restart/deploy
|
||||
```
|
||||
|
||||
**Discharge summary job:** When an encounter status changes to `discharged`, the outbox relay publishes to Kafka `encounter.status.changed`. The notification Kafka consumer reads this and publishes to `notifications.discharge.queue`. The worker generates a PDF summary (log the content; no real PDF library required), stores it in MinIO under `/discharge-summaries/{encounterId}/summary.pdf`, and marks the job complete.
|
||||
|
||||
Reference in New Issue
Block a user