Teaching Without Translating
Overview
A workplace safety instructor teaches occupational health and safety to factory workers. Her material is visual, and her students are Vietnamese speakers. Her videos were in English.
She requires that her videos be subtitled, not just translated. This is not a request for subtitles that play automatically, but rather for subtitles that are displayed on top of the video.
She said:
Translation is necessary to avoid the need for interpreters. Teachers don’t need to speak while the video is playing.

Every session, she was translating aloud over her own video. This was a problem, as she was splitting her attention between delivering the lesson and interpreting it. Her students got a lesson competing with a translation. She got neither job done as well as she could do either one alone.
I had been building a media platform since mid-2024 that would eventually do this. It was not ready. Her course was scheduled.
Engineering Challenge
The hard parts of this project were:
- Displaying subtitles correctly. The subtitles had to be displayed on top of the video, without interfering with the video playback.
- Handling long operations. The video processing and other operations could take a long time, and could fail halfway through.
- Managing authorization. The platform had to handle multiple users, and ensure that only authorized users could access certain features.
Engineering Initiatives
What the video shows, in text
The clip has no audio, so this is its text alternative rather than a caption track.
What the video shows, in text
The clip has no audio, so this is its text alternative rather than a caption track.
- A media file —
q3-planning-call.mp4— open in the library with its duration, resolution and tags. - Transcript. Speaker-labelled utterances with timestamps. Marcus and Dana discuss a migration window; one block highlights as playback reaches it.
- Subtitle generation. A language is typed into the picker — Vietnamese — and generation runs.
- Subtitles. English and Vietnamese both listed as Ready, each offering Download and Export video.
- Playback. The video plays with subtitles burned into the picture rather than layered over it.
- Summary. A generated summary, a chapter list with real timestamps, and generated topics and tags.
- Search. A term is typed and the transcript filters to matching blocks, each with its timestamp.
The Thirteen-Day Stopgap
A standalone CLI pipeline was created instead of rushing to finish the platform work on time — then the pipeline's logic was added to the platform once it was ready.
The Thirteen-Day Stopgap
A standalone CLI pipeline was created instead of rushing to finish the platform work on time — then the pipeline's logic was added to the platform once it was ready.
Source is public under MIT: github.com/dallanj/language-barrier.
Problem
The client needed a video with subtitles for a scheduled course. The platform was almost finished but couldn’t meet the deadline. Two options were bad: rush the platform to meet the deadline and waste time on the architecture, or tell the client the tool wasn’t ready.
The client’s problem was small, easy to understand, and could be solved without the platform’s help.
Solution
A CLI pipeline that works on local files:
- Uploads the file to AssemblyAI to transcribe it automatically
- Waits for the job to finish
- Gets the transcript as a WebVTT file
- Translates the transcript through Google Cloud Translation
- Burns the subtitles into the video with FFmpeg — one output per target language
Technical Highlights
- Using WebVTT as the format for the transcript. This makes it easy to see what’s happening in the pipeline and what the final output will look like.
- Each stage can be run again independently. This means that if something goes wrong, it’s easy to try again without having to start from the beginning.
- Translation is batched in chunks of 128 cues. This helps avoid running into API limits and ensures that the translation is accurate.
- Progress is shown for long operations. This helps the operator know what’s happening and when.
Engineering Considerations
The hard part was not the API integrations.
Getting the subtitles right took trial and error. The subtitles have to be perfect because they’re permanent.
FFmpeg’s stderr pipe had to be read carefully. This helped avoid hanging the process or using up too much CPU.
Escaping paths for the subtitles= filter was tricky. This required a separate layer of escaping that was different from shell escaping.
Outcome
- Delivered a Vietnamese-subtitled video on time for the course.
- Kept the platform’s build from being rushed by a deadline.
- Published the pipeline under MIT and left it standing.
- The pipeline’s logic is now in production inside the platform.
Media Ingestion: Failing Safely at Every Stage
A staged upload pipeline where the routing decision, the type column, and the processor cannot disagree — with race-safe quotas, weighted progress, and cleanup for files a database rollback can't touch.
Media Ingestion: Failing Safely at Every Stage
A staged upload pipeline where the routing decision, the type column, and the processor cannot disagree — with race-safe quotas, weighted progress, and cleanup for files a database rollback can't touch.
Problem
Uploading a file is not one operation. It involves: receiving bytes, deciding what kind of file it is, checking whether the user is allowed to store it, probing it, transcoding it, generating a thumbnail, a preview clip, and an animated preview, persisting several related rows, and then handing it off to transcription.
Any of those steps can fail. Several of them write files to disk before the failure. A database transaction rolls back rows; it does not roll back an 80MB preview clip that FFmpeg already wrote.
Three specific failures made this concrete:
- A QuickTime voice recording is
video/quicktimewith no video track. The pipeline was chosen from the MIME prefix, so it was bound to the video processor before anything inspected the file. Downstream,->videos()->first()returned null and->get()on it producedCall to a member function get() on null— a message describing the symptom and naming nothing. - Photos never uploaded at all. The type column normalized
imagetophoto; the pipeline selector didn’t, sogetPipelines()fell through to its default and threwUnsupported media type: imagefor every image. - Two concurrent uploads could both pass a storage quota check and both commit, because request-time validation reads a total that the other upload hasn’t written yet.
Solution
One resolver, consulted once. MediaTypeResolver is the single source of truth for which pipeline a file belongs to. It probes with php-ffmpeg — the same library the pipelines themselves use — so the routing decision and the processing agree by construction rather than by coincidence. Crucially it treats “tagged as video” and “carries a picture” as different questions: a track can be marked video and hold no image, which is exactly what QuickTime timecode and metadata tracks do.
return $this->hasPicture($storedPath) ? self::VIDEO : self::AUDIO;
MediaService resolves the type once and passes it into the pipeline as data. CreateMediaModel is explicitly forbidden from re-deriving it.
Quota checked under a row lock, inside the transaction. CheckStorageQuota runs first and takes lockForUpdate() on the user row, using it as a mutex for the transaction’s duration. A second concurrent upload for the same user blocks there until this one commits or rolls back — so by the time it re-reads used bytes, it sees this upload’s committed bytes. Request-time validation stays in place as fast feedback; it simply cannot close that window on its own.
Progress driven by named stages, not hand-guessed percentages. Stage weights live in config/processing_stages.php per job type and subtype, and pipeline stages call advanceStage('transcoding') rather than setting a number. Progress is the cumulative weight up to that stage, capped at 99 — only markCompleted() is allowed to say 100.
Technical Highlights
tries = 1on the upload job, deliberately, with the reason in the code. An uploaded file isn’t safely retryable without re-sending the bytes. Because that means a failure is final, the job cleans up its own stored file in every catch branch — otherwise every failure would leave the raw upload inuploads/pendingforever.- Quota rejection is logged at
info, noterror. Exceeding your plan’s limit is an expected, user-facing outcome, not a defect. Logging it as an error trains you to ignore errors. - Failure logs record origin and a trimmed trace, because the message alone wasn’t diagnosable.
"Call to a member function get() on null"names neither the probe that returned null nor the stage it happened in. Failures now capture the exception class, the base-path-relative file and line, and the first six frames — enough to diagnose from the log instead of by reproduction. - The outer try/catch inside
ProcessVideowas removed on purpose. Swallowing the error there produced a log line plus a confusing null-object failure further downstream. Letting it propagate means the transaction rolls back and the job records the real message. - Files written before a failure are tracked explicitly. Pipeline stages call
trackGeneratedFile()after each write. On success it’s informational; on failure it’s the listuploads:cleanupuses to delete every artifact the pipeline produced — precisely the files a DB rollback never touches. - Cleanup runs in three passes. Files from failed jobs (a backstop for the inline cleanup, in case the job was killed before reaching its catch block), orphaned files no
ProcessingJobreferences at all (crashes before the job row existed), and empty directories. Orphans younger than two hours are left alone so the sweep can’t race an upload that is mid-flight.--dry-runreports without deleting. - One event, many consumers.
markCompleted()fires a genericProcessingJobCompleted, and listeners decide what happens next per job type. That is the extension point that dispatches transcription after an upload finishes, rather than the upload job knowing about transcription.
Engineering Considerations
The listener that chains upload into transcription originally checked for video alone. Once audio files were being classified correctly rather than mis-routed, that check became the bug: a plain MP3, or a QuickTime voice recording now properly identified as audio, was stored and then silently never transcribed. No error, no failed job — a file that simply never gained a transcript.
Fixing the classifier is what surfaced it. That is the ordinary shape of this kind of work: correcting one thing exposes the code downstream that was quietly relying on the original mistake.
Outcome
- Routing, the type column, and the chosen processor became a single decision that cannot drift.
- Closed the concurrent-upload quota race that request-time validation structurally cannot catch.
- Made failures diagnosable from logs — origin, stage, and real message rather than a downstream null.
- No orphaned media on disk after a failure, including files written before the failing step.
- Progress reporting that reflects actual stage weights instead of invented percentages.
Transcription & AI Enrichment
Speaker-diarized transcription with LLM-generated chapters and metadata — built so the model never invents a timestamp, and so a webhook that never arrives can't strand a job.
Transcription & AI Enrichment
Speaker-diarized transcription with LLM-generated chapters and metadata — built so the model never invents a timestamp, and so a webhook that never arrives can't strand a job.
Problem
Transcription is a long-running external operation. Three things follow from that:
- The webhook is not guaranteed. It can be missed, dropped, delivered while the app is deploying, or never sent. A pipeline that treats the callback as its only completion path has jobs that hang forever.
- Not every file the user uploads is a file the API accepts. AssemblyAI’s supported container list doesn’t include QuickTime
.qt, so a screen or voice recording in that container is rejected at upload. - A raw transcript is not a product. Wall-to-wall text with no structure is barely more useful than the video. It needs chapters, a summary, searchable metadata, and word-level timings for playback.
Solution
Prepare files for the API without degrading them. AssemblyAI’s own guidance is to submit audio in its native format, because re-encoding costs quality and they downsample to 16kHz mono internally regardless. AudioTranscodeService makes a three-way decision rather than converting everything to MP3:
| Case | Action |
|---|---|
| Extension already supported | Returned untouched. FFmpeg never runs. |
| Unsupported container, codec has a supported one | Remuxed with -c:a copy — no decode, no generational loss, finishes in about the time it takes to read the file |
| Anything else | FLAC. Lossless, and far smaller to upload than WAV |
The original upload is never modified; derived files land in a separate directory so the user’s file stays exactly as they sent it.
Assume the webhook won’t come. ProcessTranscription submits the job and ends — the webhook takes it from there. transcriptions:sweep runs every five minutes with withoutOverlapping(), finds transcript jobs untouched for five minutes, and asks AssemblyAI directly what their status is. Completed ones get dispatched onward, errored ones fail, in-progress ones reset the clock. Anything still open after two hours is given up on permanently, whatever the API says at that point.
Let the model group, not measure. Chapters are generated through AssemblyAI’s LLM Gateway, replacing their deprecated auto_chapters flag. The prompt hands the model numbered paragraphs and asks for paragraph index ranges — never timestamps:
/*
* Deliberately does NOT ask the model to invent start/end millisecond
* timestamps itself — LLMs are unreliable at producing exact numbers
* for data they weren't given precisely. Instead it hands back
* paragraph indices, and resolveTimestamps() below maps those back to
* AssemblyAI's own real timestamps.
*/
The model does what it’s good at (deciding where a topic changes) and is structurally prevented from doing what it’s bad at (producing exact numbers). Chapter timings are therefore exact by construction, not by luck.
Technical Highlights
- Output shape matched to the thing it replaced. The generated chapters return
{gist, headline, summary, start, end}— the same shape AssemblyAI’s deprecatedauto_chaptersreturned — so swapping the implementation required no frontend change at all. - Model output is parsed defensively. The response is stripped of stray markdown fences before decoding, and a chapter referencing an out-of-range paragraph index is skipped with a warning rather than crashing the job. You do not trust a model to honour a format instruction every time.
- Enrichment is layered so cheap inputs feed expensive ones. Metadata generation (title, description, tags, topics, keywords) is fed the chapter summaries — already condensed, cheap to re-send — and only falls back to raw transcript text if chapter generation produced nothing.
- Every enrichment step is independently failable. Transcript blocks, chapters and metadata each sit in their own try/catch, and whatever succeeded is saved. A failure in chapter generation costs you chapters; it does not cost you the transcript.
- Word-level blocks built for playback, not for reading. Sentences are regrouped into blocks of roughly eight words, preferring to break at punctuation once that target is reached and hard-capping at eleven. Each block keeps every word’s individual start/end in milliseconds, which is what makes clickable timestamps and word-by-word highlighting possible.
- The LLM client is a sibling of the transcription client, not a special case.
LLMGatewayClientmirrorsTranscriptsClientin structure but targets a different host, because LLM Gateway is a separate service that happens to share billing and auth with the rest of AssemblyAI. - Speaker diarization with user-assigned names. Utterances carry speaker labels; users map those labels to real names, and exports resolve them at render time rather than rewriting the stored transcript.
Engineering Considerations
The reconciliation sweep is the same pattern I built into a claims platform at an agency, where a locking system depended on a third-party WebSocket service and a payment pipeline depended on a bank’s webhooks. In both cases the answer was identical: do not trust a single event to have fired correctly — verify independently on a schedule.
It is worth being precise about why the sweep isn’t redundant with the webhook. They fail differently. The webhook fails by not arriving; the sweep fails by being slow. Running both means the common case is fast and the failure case is bounded — five minutes to notice, two hours to give up — rather than unbounded.
Outcome
- Files the API would have rejected are accepted, without re-encoding anything that didn’t need it.
- A missed webhook costs five minutes, not a permanently stranded job.
- Chapter timestamps are exact because the model was never allowed near them.
- A partial enrichment failure degrades the transcript rather than losing it.
- Word-level timings power clickable transcript navigation and synchronized highlighting during playback.
Subtitles, Translation & Export
The original script's transcribe-translate-burn logic was rewritten as a job system with two generator strategies and three export formats.
Subtitles, Translation & Export
The original script's transcribe-translate-burn logic was rewritten as a job system with two generator strategies and three export formats.
Problem
The script was designed for one person on one machine. To make it a product, it needs to handle multiple users asking for multiple languages on the same media, potentially at the same time, where each request is a few minutes of API calls and CPU-bound encoding — and where asking twice should not cost twice.
Solution
Two generators behind one interface. GenerateSubtitleJob picks a strategy by comparing the requested language against the transcript’s own:
NativeSubtitleGenerator— fetches the VTT file already produced, stores it. No translation, no cost.TranslatedSubtitleGenerator— takes the native subtitle (reusing it if it exists, fetching it if not), parses the cues, translates in chunks of 128, rebuilds the VTT with original timings and translated text.
The translated generator’s reuse of the native file is the script’s idempotency idea, kept: don’t pay for work already done.
Deduplication at dispatch. SubtitleGenerationPipeline::dispatchFor() looks for an existing queued or processing job for the same media and language before creating one, and returns the existing job if it finds it. Two users asking for Vietnamese on the same video get one job and both watch it.
Timing survives translation, because timing is never translated. VttParser extracts cues; only the text array goes to the translation API; VttBuilder reassembles using the original timestamps with translated text slotted in by index, falling back to the source text where a translation is missing. Cue timing cannot drift because it never leaves the application.
Technical Highlights
- The VTT parser handles input the spec permits but tools produce inconsistently. Optional hours, either
.or,as the decimal separator (SRT-style commas normalized to VTT dots), trailing cue settings likealign:andposition:ignored, bare numeric cue identifiers skipped, and timestamps padded toHH:MM:SS.mmmwhen hours were omitted. - Burn-in ported from the script with the mechanics upgraded. Same libx264/AAC re-encode and same subtitle styling;
proc_open()replaced by Symfony Process, which handles argument escaping and exposes a progress callback natively. The FFmpeg filter-syntax escaping is still there and still commented, because it is still a separate concern from shell escaping. - Encode progress parsed from FFmpeg’s stderr into a real percentage. Duration is probed with ffprobe first; the process callback matches
time=HH:MM:SS.mson each chunk and updates the job’s progress. The user watching a nine-minute encode sees it move. - WebM inputs are re-wrapped to MP4 on export, because burning subtitles into WebM through this filter chain isn’t reliable — the same fix the script needed, carried across.
- Timeouts matched to the work. The burn job runs with
tries = 2,backoff = 30,timeout = 3600, with a comment explaining the difference: video encodes are CPU-bound local work, unlike the transcription jobs which are just waiting on an external API. - PDF export by dompdf, chosen after Browsershot failed. Headless Chrome’s crashpad handler wouldn’t launch in this Docker environment regardless of HOME overrides or disable flags — a well-documented, unresolved issue. dompdf is pure PHP: no Node, no browser binary, no Docker changes, nothing to break on an unrelated Chrome update. The trade-off is a weaker CSS engine, so the template is written defensively with tables and floats instead of flexbox.
- Two PDF layouts, same words. Grouped by time, or grouped by speaker — with an automatic fall back to time-grouping when the transcript has no diarization to group by.
- Language list cached in two layers. A static property memoizes within the request, and
Cache::flexiblekeeps it warm across requests with a six-hour fresh window and a seven-day stale window, so a slow upstream call never blocks a page render.
Outcome
- Requesting a language you already have costs nothing; requesting one twice runs once.
- Subtitle timing is exact after translation because timings never leave the application.
- Long encodes report real progress instead of appearing hung.
- Three export paths — subtitle file, burned-in video, PDF transcript — from one transcript.
Sharing as an Authorization Model
Folder sharing with three roles, consent-based invites and ownership transfer — and a privilege escalation that let a contributor seize a subfolder and lock its owner out.
Sharing as an Authorization Model
Folder sharing with three roles, consent-based invites and ownership transfer — and a privilege escalation that let a contributor seize a subfolder and lock its owner out.
Problem
Sharing looks like a simple operation, but it’s not. Authorization failures don’t give any clear error messages, so you have to use the application to find problems.
There are 12 merged pull requests, three of which are bugfix branches against work from earlier in the same sequence:
#165 feature/164-folder-invites-schema #181 feature/180-permission-sync-and-redirects
#167 feature/166-close-visitor-permission-gap #182 bugfix/177-revoke-ownership-leaks
#169 feature/168-invite-and-revoke-endpoints #184 feature/183-ownership-transfer-and-leaving
#171 feature/170-invitation-emails #187 bugfix/185-user-route-name-collision
#174 feature/173-shared-with-me #188 feature/186-sharing-changelog
#176 feature/175-broadcasting-invitation
#179 bugfix/178-login-log
These bugfix branches are left visible on purpose. The defects were found by using the application, not by a later audit.
The Escalation
A user invited to contribute to a folder could take a private space inside it and remove the folder’s owner from that space entirely.
Reproduction:
child.user_id=B? yes
B pivot role on child: contributor
B can INVITE on child (owner-only ability): YES ← escalation
BEFORE: A can view child: yes
AFTER : A can view child: NO
AFTER : A can delete child: NO
AFTER : A still owns parent: yes ← A owns the parent
AFTER : child still inside parent: yes ← and it is still in it
AFTER : B can view child: yes
Read the last three lines together. A owns the parent. The child is still inside it. A cannot see it.
How it was found: not by auditing the policy. A subfolder appeared under “Shared with me” where it had no business being — the subtree-collapsing query excluded children whose parent was also shared, but not children whose parent was owned. The visible bug looked cosmetic. Pulling that thread produced this.
Two Compounding Causes. FolderService::getContributors() was written assuming the parent’s owner owns a nested folder, while CreateFolder set user_id to the creator — two pieces of code holding different beliefs about one column. And FolderPolicy::isOwnerOrContributor() short-circuited on $folder->user_id === $user->id and returned before the $roles filter was considered, so being recorded as the creator granted every ability the policy could grant, invite included. Either alone is a bug. Together they are an escalation path.
The Fix
The fix was to stop one column carrying two facts:
| Fact | Where it lives | What it grants |
|---|---|---|
| Ownership | folders.user_id |
Administrative authority |
| Authorship | folders.created_by_user_id (new) |
Writer-level only, never administrative |
| Membership | pivot role | Whatever the invite specified |
That distinction is what makes “a contributor may delete what they created, and nothing else” a sentence the code can express. Before the split it wasn’t expressible — no column meant “made it but doesn’t own it.”
child owner is A: yes
B can invite/revoke on child: no
after B attempts revoke -- A can view child: yes
after B attempts revoke -- A can delete child: yes
Test: FolderOwnershipTest::it does not let a contributor manage access to a subfolder they created. Reverting the single line that sets subfolder ownership fails three tests, this one among them.
Two Rules Where There Should Have Been One
A separate defect, same root shape: a folder nobody could delete. The UI offered Delete because the policy allowed it; the API refused because the pipeline didn’t.
A can delete ZSub (policy): NO ← A had removed themselves; no pivot row
B can delete ZSub (policy): yes ← user_id short-circuit
B pivot role on ZSub: contributor ← but CanDestroy required Owner
B attempted delete -> still there
A attempted delete -> still there
FolderPolicy::delete() and the CanDestroy pipeline stage each carried their own copy of the rule, and the copies had drifted. HasPermission had the same shape. A folder satisfying one and failing the other was permanently stranded — not a permissions error a user could act on, a dead object in their library.
The pipeline stages now delegate to the policy. One rule, one place, consulted by everything.
This is usually filed as a tidiness problem. It isn’t. Two rules that agree in every case you tested are one refactor away from disagreeing, and when authorization rules disagree the failure isn’t an exception — it’s an object your user can see and cannot touch.
Design Decisions
- Invites Require Consent, Even from People Who Already Have Accounts. Originally, inviting an existing user attached them to the folder immediately — faster, no email round-trip. Under that behaviour, anyone who knew your email could push a folder into your library with a name and description they controlled. That is a message channel, and message channels get abused. The fix removed a code path rather than adding one.
- Contributors Cannot Share. Sharing is not itself shareable. Someone who needs to bring others in must be given ownership — which is why transfer exists. This closes the class of problem where access spreads through a graph nobody is watching.
- Transfer Reparents and Cascades. A folder lives in its owner’s tree, so handing it over moves it into the recipient’s library and carries every descendant. Leaving descendants behind would recreate the split-ownership bug one level down. Nested folders cannot be transferred at all — that would carve a hole in the current owner’s structure.
- An Owner Cannot Leave a Shared Folder. They transfer it or unshare it. This guarantees every folder has exactly one owner who can reach it — the invariant preventing orphans, and a deliberate refusal to build a button users will occasionally look for.
- The Changelog Logs from Services, Not Broadcast Events. Reusing the events is tempting since they already fire on every access change, but they cannot do audit:
leave()andrevokeContributor()emit the identical event, and inviting an address with no account emits nothing at all. Keeping open tabs fresh is a different job from recording what happened. - Notifications Split Actionable from Informational, Under One Invariant. A badge once counted items nothing could display: invite notifications whose invites had been answered were filtered out of the actionable list, excluded from the informational feed, still counted unread, and skipped by both bulk-clear controls — four behaviours, each correct alone. Replaced by a single rule: actionable + unread means a decision is live; actionable + read means resolved. That also closed a slow leak, since the prune had been sparing all actionable rows and resolved ones would have accumulated forever.
Outcome
- Closed a privilege escalation before any user encountered it, and made the underlying confusion structurally impossible rather than patched.
- Eliminated the “two authorization paths disagree” class of defect by making the policy the single source.
- Made delegated write access safe on nested structures — the case every sharing model gets wrong first.
- 75 of the platform’s 129 tests cover sharing, permissions, invites and notifications.
Architecture
The platform was built using Laravel 12 with Octane, Sanctum for authentication, and Reverb for broadcasting. The frontend was built using Vue 3 with Inertia and Pinia, Tailwind, MySQL, and Sail for local parity. AssemblyAI was used for transcription, diarization, and LLM enrichment, while Google Cloud Translation was used for translation. FFmpeg was used for all media work.
The platform used a polymorphic subject, a type, a stage list, and a weighted progress value to manage long work runs. Every status transition was broadcast on a per-user private channel, so a browser tab could show real progress on work happening in a queue worker that had no connection to it.
Results
The platform was able to:
- Remove the need for live translation in every training session, which was a recurring labor cost.
- Improve comprehension on the floor, and move the instructor into a government training role.
- Deliver under a client deadline without compromising a two-year architecture to meet it.
- Build an ingestion pipeline where a failure at any stage left no orphaned files, no partial rows, and a diagnosable log entry.
- Integrate LLM enrichment in a way that structurally prevented the model from producing unreliable output.
- Close a privilege escalation, a permanent-stranding defect, a concurrent-quota race, and an unbounded table growth path.
Engineering Practice
The platform used several engineering practices, including:
- Regression testing. Tests were verified by reverting the fix, and the test that passed both before and after proved nothing.
- Comments. Comments explained why certain decisions were made, and named the failure they prevented.
- The server is authoritative. The client was told, never trusted.
- Known duplication is documented. Duplicate code was documented in place, rather than being left to be rediscovered.
On AI Pairing
The platform was built in an intensive AI-paired session. The architecture, permission model, and every product decision were mine; the implementation was paired.
I state that plainly because the record supports it. The privilege escalation was found by me, using the running application, after noticing a subfolder in the wrong list.
What I’d Do Differently
- Separate ownership from authorship from the start. They look like one fact until a contributor creates something inside your folder.
- Resolve derived facts once, at the boundary. Media type, deletion authority, storage totals — every one of these bit me by being computed in two places.
- Never let authorization logic exist in two places. The policy and the pipeline each seemed like a reasonable home for the rule. Having both was the bug.
route:cachebelongs in CI on day one, along with anything else that can only fail at deploy time.- Test the roles nobody uses yet. The Visitor permission gap was latent purely because nobody had been a Visitor.
Scope
Backend
- Staged media ingestion pipeline with race-safe storage quotas
- FFmpeg transcoding, thumbnails, animated previews, preview clips
- AssemblyAI transcription with speaker diarization
- Container-aware audio preparation (remux-or-FLAC, never blind re-encode)
- Webhook completion plus scheduled reconciliation sweep
- LLM chapter and metadata generation via AssemblyAI LLM Gateway
- Subtitle generation with native and translated strategies
- Google Cloud Translation with batched cues
- Subtitle burn-in export, PDF transcript export
- Folder sharing, invitations, ownership transfer, audit changelog
- Policy-based authorization with pipeline delegation
- Plan and storage-limit enforcement
- Self-healing cleanup for failed and orphaned uploads
Frontend
- Vue 3 + Inertia application shell, Pinia state
- Media library, folder tree, sharing and permission management
- Transcript viewer with word-level highlighting and clickable timestamps
- Live processing-queue widget driven by broadcast job progress
- Real-time notification surfaces
Infrastructure
- Laravel Octane, MySQL, Laravel Reverb, Laravel Sail
- Queue workers for all long-running media work
- Scheduled reconciliation:
transcriptions:sweep(5 min),uploads:cleanupandnotifications:prune(nightly)
language-barrier is public under MIT at github.com/dallanj/language-barrier. The platform is private; the reproductions, tests and commit history described here can be walked through on request.



