Script icon
Docker icon
Request icon
If icon
Write icon
Schedule icon

Arcmira: YouTube Transcript Search to CSV

Search YouTube transcripts with Arcmira and export timestamped excerpts and source links to a CSV for research and video planning.

Categories
BusinessData

Search indexed YouTube transcripts for one research topic and export the returned passages to a source sheet. Keep the exact query and API response beside the CSV so reviewers can inspect each source and the search's coverage notes.

How it works

  1. build_query validates the topic, UTC date window and result limit, then URL-encodes them into one search request. User text is passed through environment variables.
  2. search sends one authenticated GET request to Arcmira's /v1/search endpoint. HTTP errors fail the flow; there are no retries or fallback sources.
  3. has_matches exports returned passages through source_sheet, or preserves the empty response through no_matches. The flow does not paginate, download media or generate transcripts.

What you get

  • outputs.source_sheet.outputFiles['source-sheet.csv'] contains video_id, published_at, start_seconds, watch_url and text. Leading spreadsheet formula characters in text cells are escaped with an apostrophe.
  • outputs.source_sheet.outputFiles['response.json'] preserves the unchanged API response, including partial-result, index-health and access notes.
  • outputs.source_sheet.outputFiles['request.json'] records the exact query and window.
  • With no matches, outputs.no_matches.uri contains the response instead of a CSV. An empty result does not establish that a topic was never discussed.

Who it's for

  • Researchers collecting source-linked passages about a topic across indexed videos.
  • Video editors finding candidate quotes or clips to review in the original recordings.

Why orchestrate this with Kestra

Kestra keeps the schedule, secret-backed request, export and execution history in one flow. Each run retains its query and response, so a reviewer can trace a CSV row back to the search that produced it. Other tasks can consume the output files for storage or reporting without embedding an API key in a script.

Prerequisites

  • An Arcmira account and an API key with research read access.
  • Kestra with the Python scripting plugin and a Docker task runner able to pull python:3.13.15-slim. The scripts use only the Python standard library.
  • Network access from the HTTP task to https://api.arcmira.com.

Secrets

  • ARCMIRA_API_KEY: supplies the Bearer token for the authenticated search request. Store it in Kestra's configured secrets mechanism and read it with {{ secret('ARCMIRA_API_KEY') }}. Do not paste a key into the flow or its inputs.

Quick start

  1. Set the secret and import this flow into your namespace.
  2. Run manually with query, a string of at least two characters. It defaults to creator economy.
  3. Optionally set published_after and published_before as UTC dates, with the start before the end. The start is inclusive and the end is exclusive. Omitted dates use a rolling window from 60 to 31 days before execution, outside the Free plan's freshness gate. To reproduce the documented example, use 2026-08-01 and 2026-09-01.
  4. limit is an integer, defaults to 3 and accepts 1 to 20. The first five search passages per request are free; results after those use credits. A paid read uses credits from your plan, then any top-up credits, then your on-demand budget. Account limits still apply.
  5. Download the CSV and inspect each watch_url alongside the preserved response. Results are research leads; check context, attribution and reuse rights before quoting a passage or editing a recording. The flow supplies no video rights.
  6. Enable the disabled Monday 09:00 UTC schedule only after a manual run. It uses the rolling window; consecutive windows overlap and outputs are separate per execution.

Coverage depends on the index and the account's plan. The flow does not change plans or bypass access gates. Inspect partial-result and access notes before drawing conclusions.

How to extend

  • Add a storage task after source_sheet to archive the CSV, request and response together under the execution ID.
  • Add a notification task that links reviewers to the output files for each run.
  • When combining scheduled exports, deduplicate by video_id and start_seconds because the rolling date windows overlap.

Links

Share this Blueprint
See How

New to Kestra?

Use blueprints to kickstart your first workflows.