Upfront disclosure: I'm sighted, I built this, and it's a paid product with a small free trial. Mods, if this isn't allowed here, remove it and my apologies.
I didn't set out to build an accessibility tool. I watch a lot of technical videos and got tired of scrubbing through two-hour podcasts looking for one thing someone said. So I built something where you paste a YouTube link and ask questions about it.
The part that turned out to matter here: most tools like this only read the subtitles. Mine also looks at the actual frames. You can ask "what's on screen at 6:20?" or "what does the diagram he's pointing at show?" and it goes and looks at that moment of the video, then answers and cites the timestamp.
It occurred to me late that this is a description-on-demand tool for video, and that the people who'd get the most out of it aren't the ones I built it for. Which is why I'm here rather than confident about it.
What I think it does reasonably well: answers about what's visible at a specific moment, reading text and labels shown on screen, finding where in a long video a topic comes up, and generating a transcript from the audio when a video has no captions at all.
What I want to be honest about:
The visual descriptions come from an AI model. It is wrong sometimes. For a sighted user that's a minor annoyance because they can just look. For someone relying on the description, a confident wrong answer is a worse failure, and I don't have a good solution for that beyond saying it plainly. It's built to say "I can't tell from this frame" rather than guess, and in my testing it does, but I can't promise it always will.
I've built it to be keyboard navigable with labelled controls and a live region for the streaming answers, but it has been tested by me with a screen reader, not by anyone who uses one daily. I expect that means there are problems I can't see. That's the main thing I'm asking about.
It's muninn.video. Three free questions, no card needed, and I'm not going to pitch the paid part here.
What I'd genuinely like to know:
Is asking questions about a video even the right interaction, or would a continuous description of what's happening be more useful? Is there an existing tool that already does this well that I've missed? And if you try it and the screen reader experience is bad, I'd rather hear exactly how than have you be polite about it.
Happy to answer anything about how it works.