Hugging Face
Business software
Sync Hugging Face model, dataset, and space data into your warehouse. Ingest builds this connector the first time a customer asks for it, then keeps the tables current in a data warehouse you own, on the schedule you choose.
The Hugging Face connector syncs your AI/ML hub data into analytics-ready tables in your data warehouse. Pull models, datasets, spaces, and metrics from the Hugging Face API, with Ingest handling auth, pagination, and retries so the tables stay current.
What you would get
Each of these arrives as its own table, kept current on the schedule you choose:
- Models
- Datasets
- Spaces
- Metrics
How it gets built
- Choose Hugging Face when you set up a pipeline, with the warehouse it goes to and a schedule.
- Ingest reads Hugging Face's documentation, and you choose the tables you want.
- Ingest builds the connector and tests it against Hugging Face itself.
- A person at Ingest reviews and publishes it, and your pipeline starts on its own.
What you will need
Bearer token (User Access Token). Create a token at huggingface.co/settings/tokens with a read, write, or fine-grained scope.
You enter it once, in a form in Ingest, and it is kept in a secret store set aside for your organization.
Questions
- Does Ingest have a Hugging Face connector?
- Not as a finished connector yet. Ingest builds it from Hugging Face's own documentation the first time a customer asks, tests it against the real service, and lists it once it passes.
- Which warehouses can Hugging Face data go to?
- Amazon S3, Google Cloud Storage, Azure Blob Storage, Apache Iceberg, Ingest Managed Lakehouse, MotherDuck, Postgres, Amazon Athena, Databricks, Redshift, BigQuery, ClickHouse, MySQL / generic SQL and Snowflake, in an account you own.
- Do I need to write code?
- No. Ingest builds, runs and maintains the connector. Someone with access to your warehouse connects it once, following a short guide.
Hugging Face API documentation: https://huggingface.co/docs/hub/security-tokens
Where this would land, and what else connects
Setting up the destination is its own short guide, one per warehouse or lake: Amazon S3, Google Cloud Storage, Azure Blob Storage, Apache Iceberg, Ingest Managed Lakehouse, MotherDuck, Postgres, Amazon Athena, Databricks, Redshift, BigQuery, ClickHouse, MySQL / generic SQL, Snowflake.
Other sources Ingest connects for SaaS teams: Asana, Jira, Airtable, Clockify, Facebook Ads, Frankfurter.