Playback & client runtime
This page covers issues in the Android app and player SDK at runtime. For build-time problems, see Android & React Native builds.
Posters and video silently fail to load (blank tiles, no error anywhere)
- Cause: Android API 28+ blocks cleartext HTTP by default. The app serves
media and art from
http://127.0.0.1:<port>(the embedded P2P node), so image loaders and ExoPlayer fail silently. - Fix: the app ships a
network_security_config.xmlthat permits cleartext only to loopback (plus the emulator host aliases for Metro in dev). The manifest references this file. See Client build. - Diagnostic that isolates it: run
adb forward tcp:<x> tcp:<port>, thencurlfrom the host. If that returns 200 while the app shows nothing, this is the cause.
Playback fails with OPLOG_CORRUPT: Oplog file appears corrupt or out of date
- Cause: the app process died mid-write (a crash or force-kill) and corrupted the local Corestore replica cache. Without recovery this is permanent until you wipe the app's data.
- Fix (shipped): the engine detects corruption codes (
OPLOG_CORRUPT,INVALID_CHECKSUM, and others) on open and on read. It purges the whole store and retries once (sdk/recover.js, exercised bynpm run test:corrupt). The store is a disposable replica cache — everything re-replicates from peers, and the in-memory session survives. No re-login is needed. - Manual fallback (always safe): clear the app's data/storage. The same reasoning applies — nothing of value is lost.
Login spins forever on "not connected to panel" / "Cannot reach the service"
Check these causes in order of likelihood:
- The panel is wedged or down. Verify from another machine with a small
hyperswarm read of
catalog/*before you blame the client. A wedged panel can look alive in the process list while it answers nothing — restart it (see Operating the panel & broadcaster). - The first DHT dial after a fresh install legitimately takes 30–90 s. The login screen retries for about a minute, then gives up. Pressing Sign in again restarts the retry loop.
- Stale swarm state sits on the device after the panel restarted, or after the app's data was cleared mid-session. Force-stop and relaunch the app — retrying inside the app is not enough, because the embedded node needs a fresh swarm.
- Device network says connected but isn't validated. Check with
adb shell dumpsys connectivity | grep -i validated.
Player shows black + spinner right after opening a channel
- Normal for a few seconds: the live edge is replicating from peers. The media server holds a request for a not-yet-replicated playlist or segment (bounded at 6 s, under ExoPlayer's 8 s read timeout) and serves it the moment it lands. So the usual cost is the actual replication time, not a 404-to-retry-remount cycle at 2.5 s. Only a path still missing after the 6 s bound 404s, and the player's 2.5 s auto-retry stays as the fallback behind that.
- Persistent spinner with
0 peers: no seeder is found. Either the broadcaster is down, its ffmpeg ingest died while the process kept "seeding" a frozen playlist, or the device holds a stale DHT record. The last case happens because the broadcaster restarted since the last lookup — its feed swarms are ephemeral identities, and hyperswarm re-queries a topic only every ~10 min. The broadcaster side self-heals the same failure via PanelLink. - Fix (shipped) — tune self-heal: while a tune is incomplete (the
playlist is not advancing — merely existing is not enough, see the
wedge below), the engine forces fresh DHT lookups on a 5 s → 60 s backoff.
At 30 s it evicts the cached feed open and re-opens fresh once (the
feed:retunebreadcrumb). At 60 s it destroys the swarm connections serving the feed and dials fresh (feed:reconnect). If that also expires, it surfaces a friendlytune timeouterror instead of spinning forever — zap to the channel again to retry. Worst case is ≤ 90 s to the error at defaults. Pre-fix builds could sit on the spinner indefinitely (S22, 2026-07-16: a zap stuck at "90 %" for 10+ min against a healthy VPS; only an app restart — a fresh swarm — cleared it, because the cached dead open poisoned every retry). Thetest:sdktune section guards the whole cycle. - Persistent spinner with peers connected (
1 peershowing): this is the wedged connection class. A network flap (Wi-Fi degrade, radio cycle) can leave the hyperswarm/UDX connection alive at transport level while replication over it moves zero bytes. Peer counts look healthy on both ends, no error fires, and because hyperswarm keeps one connection per peer across all topics, a retune faithfully reuses the same dead pipe. With prewarm, one wedged connection to the broadcaster starves every channel at once (S22, 2026-07-16: 15+ min stuck at "90 %" with "P2P — 1 peer" while a fresh client played the same feed in 10 s). Fix (shipped): the tune watchdog requires the playlist to advance (a stale pre-flap playlist in the replica no longer counts as tuned), and it tears the wedged connections down on its second expiry (feed:reconnect) so the swarm dials fresh. Thetest:sdkwedged-connection section reproduces the exact signature with a paused socket. - A redirect channel never hits this failure class at all — there is no P2P feed behind it. The host player fetches the operator's URL directly and owns its own errors. (P2P channels have no CDN failover by design — the self-heal ladder above is their recovery story.)
Video freezes while everything looks healthy (clock ticks, peers connected)
- Symptom: the picture stops dead mid-watch. Peer count and worklet heartbeats stay healthy, the UI stays alive, and no error fires. Zapping away and back fixes it.
- Cause: the HLS live window is short (the code default is 8×2 s = 16 s; reference deploys now run 12×2 s = 24 s). A network blip longer than the window slides it past ExoPlayer's position, and react-native-video raises no error event for that — the surface just freezes.
- Fix (shipped):
<AliranVideo>watches the playhead. Once a mount has played and the position sits still for 12 s (stallTimeoutMs) while not paused, it remounts onto a fresh playlist load at the live edge — the same thing the manual zap did — and firesonStallplus anonTunestart. The app's tuning pill restarts and stays until the resync mount's first real playback (onTuneplaying). - If the resync remount itself never plays within another window, the
freeze is not a slid live window but a wedged connection (see the
tune section above). The stall ladder then escalates to
backend.reconnect(), which tears down the engine's connections serving the feed and dials fresh. The engine's re-armed tune watchdog then drives the outcome — playback resumes, or a friendly error replaces the silently frozen frame. - Widen the margin (operators): the standard
HLS_LIST_SIZE=12(24 s) gives clients room to recover from blips; go to16for very flaky viewer networks — the same lever as the rebuffer cushion in sizing the segment window.
The guide can change before the picture does
The viewer apps play ~10 s behind the live edge. The programme guide's "now playing" label follows the schedule clock, so it can change up to ~10 s before the picture does. This is normal, not a fault.
Channel zapping is slow, or flipping back to a channel hangs
- How long a zap should take: switching happens inside a warm (logged-in) session, so it skips panel connect and login. Expect ~1 s to a new channel and ~0.3 s back to a channel you already watched this session. That is far below the cold time-to-play (~10 s+ with login) — if a zap takes that long, you are not actually in a warm session (the player was torn down between switches).
- Each channel is a separate P2P feed/DHT topic, so the first zap to a channel can't be instant like cable — it joins that feed's topic and pulls its first segments. Subsequent visits are near-instant, because the SDK keeps opened feeds warm.
- What a zap costs since the 2026-07-16 latency pass (all shipped,
covered by
test:serve+test:sdk): - Segment bytes stream to the player as blocks replicate (block-progressive bodies — decode starts on the first 64 KB, and every segment opens on a keyframe).
- Requests for a not-yet-replicated playlist or segment are held and served on arrival (bounded), which kills the old 404-to-2.5 s-retry quantization.
- Serving a live playlist now replicates the whole live window in parallel for the active stream (metered networks keep the newest-3 read-ahead), so replication overlaps the player's sequential fetches.
- ExoPlayer starts at ~1 s buffered instead of ~2.5 s (the
<AliranVideo>bufferConfigdefault). The stall-resync/self-heal ladder covers the slightly higher rebuffer risk this creates. - Optional
zapPrefetchkeeps the adjacent channels' newest segment warm (off by default — it costs standing bandwidth). - Fixed: flipping back to a channel used to hang.
resolve()opened a duplicate Hyperdrive on the same store namespace and deadlocked.sdk/player.js serveFeednow reuses the cached feed perfeedKey, so make sure your build includes this fix (thetest:sdkzapnews → movies → newsregression guards it). - First zap also warm (pre-warm): the SDK opens entitled feeds' topics
at login (the
prewarmoption; the app enables it), so even the first play/zap to a channel is a cache hit — verified on-device asfeed:readywith nofeed:open. See P2P feed buffer & tuning.
App dies the moment the player seeks / switches source / closes (worklet SIGABRT)
- Symptom: the whole app exits. Logcat shows
Uncaught StreamError: Writable stream closedand an abort inside the Bare runtime. - Cause: video players routinely abort in-flight HTTP requests. Writing into the closed response was an unhandled stream error, and any uncaught exception in the embedded worklet aborts the entire app process.
- Fix (shipped): the media server tolerates client aborts on every path.
The worklet also installs a last-resort
uncaughtExceptionguard that reports the error over IPC instead of crashing. If you embed the SDK in your own runtime, keep both.
Reading the app's own diagnostics (dev builds)
Every backend→UI IPC message is logged. adb logcat -s ReactNativeJS shows
[backend] {"type": ...} lines, including feed:open / feed:ready
breadcrumbs, peer-count ticks, fallback / source-changed events, and
store:reset (corruption recovery). Read these before guessing from the
screen.