Download from arte.tv with subtitles and multiple audio tracks

I downloaded a film from arte.tv with yt-dlp and --sub-langs all --write-subs. The .vtt files were there, mpv listed them and I could cycle through them with j, but no subtitle was ever displayed.

The cause

ARTE's files start with a long series of STYLE blocks before the first cue:

STYLE
::cue(.black) { color: black; }

ffmpeg does not skip STYLE blocks. It hits one, finds no --> timing line, and stops parsing. The track exists and is empty. mpv uses the same WebVTT demuxer, so it lists a track that never shows anything.

The fix

The obvious approach is to let yt-dlp convert the subtitles on download and get SRT files that mpv has no trouble with. But --convert-subs srt does not help here, because yt-dlp hands the file to ffmpeg and ffmpeg chokes on the same STYLE blocks.

So drop everything before the first cue and put a clean header back:

{ printf 'WEBVTT\n\n'; sed -n '/-->/,$p' "$f"; } > "$f.tmp" && mv "$f.tmp" "$f"

This is done in place because mpv autoloads subtitles by filename.

Doing it automatically

A wrapper gets the final filename from yt-dlp with --print-to-file after_move:filepath and fixes every .vtt next to the video:

#!/usr/bin/env bash
p=$(mktemp)
yt-dlp -f 'bv[height<=720]+mergeall[vcodec=none][format_id!*=clear_voices]' --audio-multistreams --sub-langs all --write-subs \
       --print-to-file after_move:filepath "$p" "$@"
while read -r v; do
  for s in "${v%.*}".*.vtt; do
    { printf 'WEBVTT\n\n'; sed -n '/-->/,$p' "$s"; } > "$s.tmp" && mv "$s.tmp" "$s"
  done
done < "$p"
rm -f "$p"

format_id!*=clear_voices skips ARTE's clear voices tracks. That is an accessibility option: a separate mix of the soundtrack with the dialogue emphasized and music, ambient noise and sound effects turned down, so speech is easier to understand.

The script not only downloads all subtitles but also all audio tracks. A lot of ARTE films have the original audio and a German or French dub if available. This leaves it to me to decide if I am fine with the (dubbed) German version, or prefer to watch the original audio track and read subtitles.

Two tailnets at the same time on one Linux notebook

I use two tailnets on my notebook: my private one and the one from work. Tailscale supports multiple accounts, but only one is active at a time, so I was switching with tailscale switch regularly. From the work tailnet I only need the subnet routes into our AWS infrastructure -- nobody there needs to reach my notebook. So the work tailnet can run as a second tailscaled instance next to the private one.

A second systemd service

The second instance is a copy of tailscaled.service from the Arch Linux package, with its own state file, socket, TUN device and UDP port. I saved it as /etc/systemd/system/tailscaled-work.service:

[Unit]
Description=Tailscale node agent (work tailnet)
Wants=network-pre.target
# systemd-networkd here; with NetworkManager use NetworkManager.service instead
After=network-pre.target systemd-networkd.service systemd-resolved.service tailscaled.service

[Service]
ExecStart=/usr/sbin/tailscaled --state=/var/lib/tailscale-work/tailscaled.state --socket=/run/tailscale-work/tailscaled.sock --tun=tailscale1 --port=41642

Restart=on-failure

RuntimeDirectory=tailscale-work
RuntimeDirectoryMode=0755
StateDirectory=tailscale-work
StateDirectoryMode=0700
CacheDirectory=tailscale-work
CacheDirectoryMode=0750
Type=notify

[Install]
WantedBy=multi-user.target

The original unit has ExecStopPost=/usr/sbin/tailscaled --cleanup, which I left out. Both instances share the same policy routing rules and routing table 52, and the cleanup of one instance should not remove what the other one is using.

This does not help completely, though. Stopping tailscaled-work still deletes the shared rules -- ip rule afterwards only lists local, main and default:

0:      from all lookup local
32766:  from all lookup main
32767:  from all lookup default

Without the lookup 52 rule the private tailnet is unreachable too. Starting the work instance again adds the rules back, and restarting the private one with sudo systemctl restart tailscaled should do the same. As long as both services just run all the time, this does not matter.

Enable and start the new service:

sudo systemctl daemon-reload
sudo systemctl enable --now tailscaled-work
systemctl status tailscaled-work

Logging in

All tailscale commands for the second instance need the socket, so an alias helps:

alias tailscale-work='tailscale --socket=/run/tailscale-work/tailscaled.sock'

Logging in creates a new node in the work tailnet -- the old work profile in the main instance is not reused:

sudo tailscale --socket=/run/tailscale-work/tailscaled.sock up \
  --accept-routes --accept-dns=false --shields-up --netfilter-mode=off --operator=$USER
--accept-routes

Gets the AWS subnet routes.

--accept-dns=false

Keeps MagicDNS and the DNS settings with the private tailnet.

--shields-up

Blocks all incoming connections from the work tailnet.

--netfilter-mode=off

Leaves the iptables chains to the private instance; both would write into the same ts-input chain, and its rule drops 100.x traffic that does not come in on its own interface.

--operator

Allows my user to run tailscale-work status without sudo.

The old work node still has the hostname, so the new machine gets the old name with -1 appended. I deleted the old node in the admin console and renamed the new one to drop the -1.

Result

Both instances are logged in and each has its own interface:

tailscale status --peers=false
tailscale-work status --peers=false
ip -br addr show | grep tailscale

The subnet routes from the work tailnet land in table 52 on tailscale1, next to the peers of the private tailnet on tailscale0:

$ ip route show table 52 | grep -v '^100\.'
10.x.0.0/16 dev tailscale1
172.x.0.0/16 dev tailscale1

ip route get with an address from one of these subnets shows dev tailscale1 table 52. A host at work that is only reachable via the tailnet loads in the browser, and my private peers still answer tailscale ping. It is only this simple because the two tailnets do not share any 100.x addresses and my private tailnet has no routes into AWS.

Finally, remove the old work profile from the main instance (tailscale switch --list shows the names):

tailscale switch work
tailscale logout
tailscale switch <private-profile>

No more tailscale switch when a dev setup needs AWS access for work.

Matrix status messages from a cron job

The Wikidata cache now updates itself from cron. If an update is happening I want to get a status message about the results. The script posts into a Matrix room when it loaded something, and stays quiet when no new Wikidata dump was found and no update happened.

The bot uses my own Matrix server instance and has its own account.

Creating the room

The bot account creates the room and invites me:

room = httpx.post(
    f"{base}/_matrix/client/v3/createRoom",
    headers={"Authorization": f"Bearer {token}"},
    json={
        "name": "wikidata-cache",
        "topic": "Status of the automatic Wikidata dump loads",
        "preset": "private_chat",
        "invite": ["@me:cress.space"],
    },
).json()

The response is one line:

{"room_id": "!AbCdEfGhIjKlMnOpQr-StUvWxYz0123456789abcdef"}

That id is stored and used by the script to send messages later. An alias like #wikidata-cache:cress.space has to be resolved into the id first, one request more per message.

Minting a token

The password, given by the create-user command in the conduit admin room, is only used once, to get an access token:

login = httpx.post(
    f"{base}/_matrix/client/v3/login",
    json={
        "type": "m.login.password",
        "identifier": {"type": "m.id.user", "user": "@bot:cress.space"},
        "password": password,
        "device_id": "wikidata-cache-cron",
        "initial_device_display_name": "wikidata-cache-cron",
    },
).json()

device_id is fixed, so a second login replaces that device instead of registering another one.

The token, not the password, is what goes on disk. A token can be revoked from any Matrix client without changing the account password. It does not expire on its own either: a login that does not ask for a refresh token gets one that is valid until the device is logged out, the password changes, or someone revokes it. Mine lives in a matrix.toml next to the script:

homeserver = "https://chat.cress.space"
room = "!AbCdEfGhIjKlMnOpQr-StUvWxYz0123456789abcdef"
token = "syt_..."

Sending the message

Sending is a PUT, with a transaction id in the URL that makes a repeated send idempotent:

httpx.put(
    f"{base}/_matrix/client/v3/rooms/{room}/send/m.room.message/{time.time_ns()}",
    headers={"Authorization": f"Bearer {token}"},
    json={"msgtype": "m.notice", "body": text},
).raise_for_status()

m.notice instead of m.text, so clients that mute notices can do so.

The room id goes percent-encoded into the URL: it starts with ! and holds a :, so urllib.parse.quote(room, safe="") makes %21AbCdEf...%3Acress.space of it.

Sending an image

Another cron of mine pushes a photo once a day, until now to ntfy. I moved it to Matrix so ntfy only gets alarms.

Matrix can handle images too, but it needs two calls instead of one. The image is uploaded to the media repository first, which answers with an mxc:// URI:

MXC=$(curl -sf -X POST \
    -H "Authorization: Bearer $MATRIX_TOKEN" \
    -H "Content-Type: image/jpeg" \
    --data-binary "@$IMAGE" \
    "$MATRIX_HOMESERVER/_matrix/media/v3/upload?filename=$NAME" | jq -r .content_uri)

The media endpoint takes a POST with --data-binary.

Then the same PUT as a text message, with m.image and the URI in it:

curl -sf -X PUT \
    -H "Authorization: Bearer $MATRIX_TOKEN" \
    -H "Content-Type: application/json" \
    -d "$(jq -n --arg url "$MXC" --arg name "$NAME" \
        '{msgtype:"m.image", body:$name, url:$url, info:{mimetype:"image/jpeg"}}')" \
    "$MATRIX_HOMESERVER/_matrix/client/v3/rooms/$ROOM/send/m.room.message/$(date +%s%N)"

ntfy's prio:low has no equivalent here; m.image has no notice variant, so quiet is a per-room setting in the client.