Run icon
Get icon
Assert icon
Log icon
Write icon
FileTransform icon
IonToCsv icon
Query icon
CopyIn icon

Verify and load source-backed company enrichment with Apify

Enrich one company domain with Apify, verify explicit website-scrape provenance, transform the result, and bulk-load it into Postgres.

Categories
BusinessData

Turn one company domain into review-ready website metadata loaded directly into Postgres without silently treating estimates as company facts. This flow runs the fixed Brainiall Company Enrichment Actor, fetches its single result, fails unless the item declares successful website-scrape provenance and the expected Kestra route label, reshapes the candidate record, and bulk-loads it into Postgres.

How it works

  1. enrich_company (io.kestra.plugin.apify.actor.Run) calls Actor tFyauKW1J9LcscDT8 with the selected bare domain and the bounded kestra-c19 route label. maxItems: 1, maxTotalChargeUsd: 0.02, and memory: MB_256 limit the run before it starts.
  2. get_source_backed_result (io.kestra.plugin.apify.dataset.Get) reads the run's {{ outputs.enrich_company.defaultDatasetId }}.
  3. verify_source_backed_result (io.kestra.plugin.core.execution.Assert) requires exactly one item, success: true, provenance method website_metadata_scrape, and integration source kestra-c19.
  4. log_audit_receipt (io.kestra.plugin.core.log.Log) records the non-secret Actor run ID, dataset ID, and platform-reported usage.
  5. write_json (io.kestra.plugin.core.storage.Write) materializes the validated dataset array to Kestra internal storage as a JSON file.
  6. reshape (io.kestra.plugin.graalvm.python.FileTransform) extracts the domain, company name candidate, description candidate, provenance method, source URL, observation timestamp, and integration source into structured rows.
  7. to_csv (io.kestra.plugin.serdes.csv.IonToCsv) serializes the reshaped enrichment record to CSV.
  8. create_table (io.kestra.plugin.jdbc.postgresql.Query) creates the public.source_backed_company_enrichment table in Postgres if missing.
  9. load_data (io.kestra.plugin.jdbc.postgresql.CopyIn) bulk-loads the CSV record into the Postgres table.

What you get

  • One schema-checked website-metadata candidate record loaded to Postgres.
  • Explicit source URL and caveat stored alongside the candidate metadata.
  • Actor run and dataset identifiers for auditability.
  • A hard total-charge cap; actual Apify platform and Actor charges can apply.

Who it's for

  • RevOps teams populating company databases with verified candidate records.
  • Research and automation teams requiring explicit source URLs in SQL storage.
  • Data engineers looking for an auditable enrichment-to-warehouse pipeline.

Prerequisites

  • A Kestra instance with Apify, Core, GraalVM Python, SerDes, and JDBC PostgreSQL plugins.
  • An Apify account permitted to run public Actors.
  • A reachable PostgreSQL database.

Secrets

  • APIFY_API_TOKEN: a scoped Apify token allowed to run the public Actor and read the resulting default dataset. Never place the token in the flow YAML.
  • POSTGRES_HOST: hostname or IP address of the PostgreSQL database server.
  • POSTGRES_USERNAME: username for PostgreSQL database authentication.
  • POSTGRES_PASSWORD: password for PostgreSQL database authentication.

Inputs

  • domain (STRING, default example.com): one bare public company domain. Schemes, paths, ports, credentials, wildcards, and batch input are rejected.

Quick start

  1. Store APIFY_API_TOKEN, POSTGRES_HOST, POSTGRES_USERNAME, and POSTGRES_PASSWORD in your Kestra instance secrets.
  2. Import the flow and set domain to a public company website you are entitled to process.
  3. Execute the flow to run the Actor, verify provenance, transform the result, and bulk-load the record into Postgres.
  4. Query public.source_backed_company_enrichment in Postgres to inspect the ingested candidate record.

Expected outputs

  • {{ outputs.enrich_company.id }}: the Apify Actor run ID.
  • {{ outputs.enrich_company.defaultDatasetId }}: the default dataset ID.
  • {{ outputs.get_source_backed_result.dataset }}: an array containing one validated candidate item with provenance and observedAt.
  • {{ outputs.reshape.uri }}: internal URI of the transformed record.
  • {{ outputs.to_csv.uri }}: internal URI of the CSV record.
  • {{ outputs.load_data.rowCount }}: row count loaded into Postgres.

Common pitfalls

  • A failed or source-less upstream response intentionally writes zero items, causing the assertion to fail before transformation or database loading.
  • The route label is caller-supplied attribution, not buyer identity, payment, payout, settlement, or revenue evidence.
  • Candidate fields in Postgres are website-derived candidates and should be reviewed against authoritative sources before CRM, credit, or compliance decisions.

Links

See How

New to Kestra?

Use blueprints to kickstart your first workflows.