Música.
running, healthy, and doing nothing.
Two settings that had been quietly ignored for 37 days, a root-owned volume on the Dell, and the restart test that almost scared me.
How this started
The last post ended with two Navidrome settings that had been quietly doing nothing, and my music still sitting on my Mac instead of the Dell. Thursday night I opened the checklist to block 4 and gave it 45 minutes.
Música is the music library tile in Laboratório Brasil. It is also the first tile I want to put on the internet, so before anyone else touches it I wanted to know it works. Not that the pod is green. That I can press play on a bossa nova album and hear the whole song.
It took a lot longer than 45 minutes.
Look before touching anything
My notes already said the two settings were broken. I did not want to fix something because a note told me to. I wanted to see it.
On the control plane:
kubectl -n brasil get pods,svc,ingress | grep -i navidrome # -> pod/navidrome-88b88cff6-l4mk7 1/1 Running 0 2d14h # -> service/navidrome ClusterIP ... 80/TCP 37d # -> ingress.networking.k8s.io/navidrome traefik musica.brasil.local ... 80 37d
Running, zero restarts, up for two days. As far as Kubernetes was concerned, everything was fine.
Then the logs:
kubectl -n brasil logs deploy/navidrome --tail=30 # -> level=info msg="Periodic scan is DISABLED" # -> ... # -> level=info msg="Executing initial scan" # -> level=warning msg="Playlists will not be imported, as there are no admin users yet..." # -> level=info msg="Scanner: Finished scanning all libraries" duration=2.8ms
get tells you what Kubernetes thinks. logs tells you what the app thinks. Kubernetes thought Navidrome was fine. Navidrome was telling me three things were not.
Periodic scan is DISABLED. I had set a 3am daily scan in the manifest. Navidrome did not know about it.- The scan finished in 2.8 milliseconds. Nothing scans a real music library that fast. The library was empty.
no admin users yet. In Navidrome, whoever logs in first becomes the admin. The ingress was already live on my home network. That is a race I wanted to win.
Two settings that did nothing
Here is what was in clusters/staging/brasil/navidrome.yaml:
- name: ND_SCANSCHEDULE value: "0 3 * * *" - name: ND_ENABLESTARTUPSCAN value: "true"
Navidrome renamed these at some point. The current names are ND_SCANNER_SCHEDULE and ND_SCANNER_SCANONSTARTUP. An app does not throw an error on an environment variable it does not recognize. It just skips it. The manifest looked right. The pod was healthy. The scan schedule did not exist.
The startup scan was sneakier. It did run. But only because it is on by default. My setting was being ignored the same way, and the default happened to match what I wanted. I fixed it anyway. I would rather the config say what it means than get lucky.
The checklist had the wrong port
Before the fix, I wanted to open Navidrome in a browser. My checklist said:
kubectl -n brasil port-forward svc/navidrome 4533:4533
The service line above says 80/TCP. When you port-forward to a service, the number on the right has to be a port the service exposes, and mine only exposes 80. Navidrome itself listens on 4533 inside the pod. So there are three ports in a row:
my Mac service pod localhost:4533 --> port 80 --> port 4533
The fixed command, on my Mac this time since that is where the browser is:
kubectl -n brasil port-forward svc/navidrome 4533:80 # -> Forwarding from 127.0.0.1:4533 -> 4533 # -> Forwarding from [::1]:4533 -> 4533
That -> 4533 confused me for a second. kubectl looks up the service, sees that 80 points to 4533, and connects straight to the pod.
I went to localhost:4533 and made the admin account before anyone else on the network could.
The fix goes through Git
On this cluster Flux owns the manifests. If I changed the deployment with kubectl edit, Flux would put it back. So the fix goes in the repo.
I opened the file in vim and ran one substitute for both names. The first try:
E486: Pattern not found: NDND_SCANSCHEDULE
I typed ND twice. Nothing changed, which is the nice thing about a pattern that does not match. I retyped it and it went through. :set number helped too. I had been counting lines by eye.
Before committing I checked exactly what changed:
git diff # -> - - name: ND_SCANSCHEDULE # -> + - name: ND_SCANNER_SCHEDULE # -> ... # -> - - name: ND_ENABLESTARTUPSCAN # -> + - name: ND_SCANNER_SCANONSTARTUP
Two lines out, two lines in, comments and indentation untouched. Then:
git add clusters/staging/brasil/navidrome.yaml git commit -m "fix(navidrome): correct scanner env var names" git push origin main
And on the control plane, so I did not have to wait for Flux’s next sync:
flux reconcile kustomization flux-system --with-source # -> ✔ applied revision main@sha1:1ef2abf... kubectl -n brasil get pods | grep navidrome # -> navidrome-cb8fb96c7-lszdz 1/1 Running 0 29s
The pod name changed. The middle part, 88b88cff6 before and cb8fb96c7 now, is a hash of the pod template. Changing an env var changes the template, so Kubernetes made a new ReplicaSet and replaced the pod.
The logs, one more time:
kubectl -n brasil logs deploy/navidrome | grep -i scan # -> level=info msg="Scheduling periodic scan" schedule="0 3 * * *" # -> level=info msg="Executing initial scan"
DISABLED became Scheduling periodic scan. Fixed, and I can point at the line that proves it.
Two things came with it.
My port-forward died. It was attached to the old pod, and the old pod was gone. Port-forward picks one pod when it starts and stays with it. Hold onto that, it comes back later.
And 3am is not 3am. The log timestamps end in Z, which means UTC. The container runs on UTC, so 0 3 * * * is 11pm in Miami. A library scan at 11pm does not bother me. A backup job would.
Getting the music onto the Dell
I did not have a folder for this. My Brazilian music was mixed in with everything else. So I made one on the daily driver:
mkdir -p ~/Music/brasil
That folder is the only place I add music. The copy on the Dell is what Navidrome serves, and I never touch it by hand. One album first to test with, 96M. Then I got impatient and added three more.
du -sh ~/Music/brasil # -> 701M /Users/op/Music/brasil
Which volume, and where does it live
Navidrome has two volumes:
kubectl -n brasil get pvc | grep -i navidrome # -> navidrome-data Bound pvc-a4be8a71-... 5Gi RWO local-path 37d # -> navidrome-music Bound pvc-20f6e03d-... 100Gi RWO local-path 37d
navidrome-data is the database. navidrome-music is the files. Music in the data volume would not show up, and I would have wondered why for a while.
With local-path, the volume is a regular folder on whichever node first ran the pod. I thought it was the Dell. I checked anyway:
kubectl get pv pvc-20f6e03d-... -o jsonpath='{.spec.hostPath.path}{.spec.local.path}{"\n"}'
# -> /var/lib/rancher/k3s/storage/pvc-20f6e03d-..._brasil_navidrome-music
kubectl get pv pvc-20f6e03d-... -o jsonpath='{.spec.nodeAffinity.required.nodeSelectorTerms[0].matchExpressions[0].values[0]}{"\n"}'
# -> dell-nodeThe Dell, under /var/lib/rancher. Which is owned by root.
rsync and sudo
My user on the Dell cannot write there. So I tried a dry run that runs rsync with sudo on the far end:
rsync -avhn --rsync-path="sudo rsync" ~/Music/brasil/ dell:/var/lib/rancher/k3s/storage/<pv-folder>/ # -> sudo: a terminal is required to read the password; either use ssh's -t option or configure an askpass helper # -> sudo: a password is required # -> rsync error: unexpected end of file
I have seen this exact error before. It is the same one from getting the kubeconfig off the server. rsync runs over SSH with no terminal, so sudo has nowhere to ask for my password.
One fix is a sudoers rule that lets my user run rsync with no password. That is a real security decision and I did not want to make it at 8:30 at night. So I went with two hops.
# daily driver to my home folder on the Dell, which I own rsync -avh ~/Music/brasil/ dell:~/navidrome-upload/ # -> sent 735M bytes received 1276 bytes 33584k bytes/sec # then on the Dell, where sudo has a terminal ssh dell sudo rsync -avh ~/navidrome-upload/ /var/lib/rancher/k3s/storage/<pv-folder>/ sudo ls /var/lib/rancher/k3s/storage/<pv-folder>/ rm -rf ~/navidrome-upload exit
~/Music/brasil/ matters. With it, rsync copies what is inside the folder. Without it, rsync copies the folder itself, and the library ends up at /music/brasil/Artist/…. That is how you end up with folders inside folders you did not ask for.Testing it
Four checks. Each one proves something the one before it does not.
Can I reach it. I started the port-forward again on my Mac and opened localhost:4533. The terminal immediately filled up with this:
E0924 20:40:10.820561 87651 portforward.go:489] "Unhandled Error" err="error copying from remote stream to local connection: ... write: broken pipe"
It looked bad. It was not. The browser closes connections early all the time, when you click away or it cancels an image it does not need anymore. Port-forward logs every one of them. The page was working fine.
Did it find the music. Astrud Gilberto, The Bossa Nova Queen, 12 songs, cover art and all. All four albums were there. The files were in the right volume and the scanner found them.
Does it play. I played a song all the way through. A library that lists albums and cannot stream them is a catalog. Playback proves the whole path works, from the browser through the forward to the pod to the disk on the Dell and back.
Does it survive a restart. This is the one I care about most. If the music or the database lived inside the container, a restart would wipe them. I ran this from my daily driver, which after last week can actually talk to the cluster:
kubectl -n brasil rollout restart deploy/navidrome # -> deployment.apps/navidrome restarted kubectl -n brasil get pods | grep navidrome # -> navidrome-7c4d684dd9-5jfnv 1/1 Running 0 10s
New pod, 10 seconds old. I refreshed the browser.
This site can't be reached localhost refused to connect. ERR_CONNECTION_REFUSED
For a second I thought the library was gone. Then I remembered what happened an hour earlier. The port-forward was attached to the old pod. The old pod was gone, so nothing was listening on 4533 anymore. I stopped the forward, started it again, and refreshed.
Everything was there. The albums, my login, all of it.
That is the whole point of the test. The music lives on navidrome-music, the database lives on navidrome-data, and the container can die whenever it wants.
I hit enough port-forward errors that night to fill a post of their own, so they got one.
Final checklist: is Música actually working
# 1. The pod is up kubectl -n brasil get pods | grep navidrome # -> 1/1 Running # 2. The scan settings are being read, not ignored kubectl -n brasil logs deploy/navidrome | grep -i "periodic scan" # -> Scheduling periodic scan, not DISABLED # 3. The scan found files kubectl -n brasil logs deploy/navidrome | grep -i "finished scanning" # -> longer than a few milliseconds # 4. You can reach it, from the machine with the browser kubectl -n brasil port-forward svc/navidrome 4533:80 # 5. It survives a restart kubectl -n brasil rollout restart deploy/navidrome # -> restart the port-forward, refresh, library is still there
Música is ready for the next step. Right now it only works on my home network. Next it goes behind a Cloudflare Tunnel with a login in front of it, so it can be the first tile anyone else gets to use.