When you upload a file to Soontra, you can have it AI-indexed — read, watched, and listened to in the background so it becomes something you can search and ask about, not just a name in a folder. Indexing is what turns a two-hour recording into a set of moments you can find in seconds.
This article covers what indexing extracts, how to track its progress, and how your content is — and isn't — used.
Indexing is opt-in per upload
You choose whether to index each file at upload time. If you skip indexing — say, for raw footage you know you'll never search — the file is still stored and available like any other; it just matches searches by name only. You can index it later at any time, so nothing is locked in by that first decision.
What indexing produces
For every indexed file, Soontra builds a structured understanding:
- Word-level transcripts of videos and audio, so you can find a line by what was said and jump to the second it was spoken.
- Visual scene descriptions of what appears on screen, even when nothing is being said.
- On-screen text (OCR) — signs, slides, captions, and text printed inside images.
- Objects, actions, topics, and people, each tied to timestamped moments.
- Embeddings, the numeric fingerprints that power semantic search, so "the part where I talk about pricing" matches even if that exact phrase never appears.
Because spoken words and visuals are lined up moment by moment, a search can land you on the right second rather than just the right file.
Track progress: the status pill and notifications
Indexing runs quietly in the background — you can close the tab and keep working. The pill in Drive shows your workspace's indexing state at a glance: a solid dot means everything is indexed and searchable, and a spinner means files are still being indexed. Click the pill anytime to see what's in progress. You also get a notification the moment a file is ready.
Which file types can be indexed
- Video — full transcript plus visual understanding, aligned moment by moment.
- Audio — podcasts, voice memos, and recorded calls become clean, word-level transcripts.
- Images — a plain description of what's in the frame, plus a read of any text inside it.
- Documents — PDFs and Word documents have their text extracted and made searchable, down to the page.
- Spreadsheets — indexed so their contents are searchable too.
Is my content used to train AI models?
No. Soontra uses enterprise AI APIs that don't train on customer data. Your footage, scripts, and conversations stay yours — Soontra processes them to give you results, not to improve anyone else's model.