Awesome!

We have received your request and will reach out shortly with more information. We’re super excited to show you what Uplifted can do for you!

Close
Oops! Something went wrong while submitting the form.
Skills library

Natural-Language Footage Finder

Organize
by
Uplifted
94
Saves
Overview
What it does
Finds the exact clip from a plain-language description of what's in it — subject, attributes, action, setting, or spoken line.
When to use it
You need a specific clip and you only know what's in it — "the shot of a long-bearded guy who isn't white," "the red-shirt scene even though the color was never in the transcript," "the b-roll of the product on a marble counter." Instead of scrubbing folders, you describe it and get it.
Pairs with MCP
Uplifted indexes the visual content of every frame plus the transcript, so the model can find footage by what the camera actually shows — not just by filename or what someone happened to type in a description.
Best for
Video editors, content marketers, creative teams.
Install
SKILL.md · natural-language-footage-finder
---
name: natural-language-footage-finder
description: Use this skill when the user wants to find specific clips by describing what's in them — people, objects, settings, actions, on-screen text, demographics, or spoken content — even when the detail isn't in any filename or written description. Searches Uplifted's multimodal visual + transcript index.
---

# Natural-Language Footage Finder

You find footage by what it actually contains, across visual content and transcript.

## Data you need
- A natural-language description of the desired footage from the user (subject, action, setting, attributes, on-screen text, spoken phrase)
- Uplifted's multimodal index via MCP: per-clip visual tags/embeddings, detected objects/people/attributes, on-screen text (OCR), and transcript

If multimodal visual search isn't available for some assets, say so and search transcript + tags only, flagging the gap.

## How to search
1. Parse the request into searchable facets: subject(s), attributes (e.g., beard length, clothing color, age/skin tone where the user specifies it for casting needs), action, setting, on-screen text, spoken phrase.
2. Query the visual index AND the transcript — a "red shirt" should be found from the picture even if no one says "red shirt."
3. Rank results by match confidence across facets; show the strongest matches first.
4. For each result, give the clip ID, a one-line description of why it matches, the timecode of the matching moment, and the source asset it lives in.

## Output format
RESULTS for: "[the query]"
Ranked list — per clip: Clip ID | Why it matches | Timecode | Source asset | Match confidence
NEAR MISSES — clips that match some facets but not all, in case the brief flexes
NOT FOUND — if a facet had no matches, say so plainly so the user knows to shoot it

## Guidelines
- Search the picture, not just the words — explicitly use the visual index, not only the transcript.
- Be honest about confidence; don't pad results with weak matches presented as strong.
- If nothing matches, say so and suggest the closest available alternative or that it needs shooting.
Prompt
Find footage in my library by what's actually in it.

I'm looking for: {{describe the clip — subject, attributes, action, setting, on-screen text, or spoken phrase}}.

Data: Uplifted's multimodal index via MCP — per-clip visual tags, detected objects/people/attributes, on-screen text (OCR), and transcript. {{confirm MCP or paste an index export}}

1. Parse my request into facets (subject, attributes, action, setting, on-screen text, spoken phrase).
2. Search BOTH the visual index and the transcript — find a "red shirt" from the picture even if no one says it.
3. Rank by match confidence across facets.
4. Per result: Clip ID | Why it matches | Timecode | Source asset | Confidence.

Also list near misses and, if a facet has no matches, say so.

Search the picture, not just the words. Be honest about confidence; no weak matches dressed up as strong.

EXPECTED OUTPUT:
- A ranked list of matching clips with IDs, timecodes, and why each matches
- Near misses in case the brief can flex
- A clear "not found — shoot this" note for anything missing