Wikipedia
Business software
Sync Wikipedia pageview and traffic analytics data into your warehouse. Ingest builds this connector the first time a customer asks for it, then keeps the tables current in a data warehouse you own, on the schedule you choose.
The Wikipedia connector syncs your pageview analytics data into analytics-ready tables in your data warehouse. Pull pageviews and articles from the Wikimedia Pageviews API, with Ingest handling auth, pagination, and retries so the tables stay current.
What you would get
Each of these arrives as its own table, kept current on the schedule you choose:
- Pageviews
- Articles
How it gets built
- Choose Wikipedia when you set up a pipeline, with the warehouse it goes to and a schedule.
- Ingest reads Wikipedia's documentation, and you choose the tables you want.
- Ingest builds the connector and tests it against Wikipedia itself.
- A person at Ingest reviews and publishes it, and your pipeline starts on its own.
What you will need
None (public, open access). No signup or key required; requests just need a descriptive User-Agent header identifying the client.
You enter it once, in a form in Ingest, and it is kept in a secret store set aside for your organization.
Questions
- Does Ingest have a Wikipedia connector?
- Not as a finished connector yet. Ingest builds it from Wikipedia's own documentation the first time a customer asks, tests it against the real service, and lists it once it passes.
- Which warehouses can Wikipedia data go to?
- Amazon S3, Google Cloud Storage, Azure Blob Storage, Apache Iceberg, Ingest Managed Lakehouse, MotherDuck, Postgres, Amazon Athena, Databricks, Redshift, BigQuery, ClickHouse, MySQL / generic SQL and Snowflake, in an account you own.
- Do I need to write code?
- No. Ingest builds, runs and maintains the connector. Someone with access to your warehouse connects it once, following a short guide.
Wikipedia API documentation: https://doc.wikimedia.org/generated-data-platform/aqs/analytics-api/reference/page-views.html
Where this would land, and what else connects
Setting up the destination is its own short guide, one per warehouse or lake: Amazon S3, Google Cloud Storage, Azure Blob Storage, Apache Iceberg, Ingest Managed Lakehouse, MotherDuck, Postgres, Amazon Athena, Databricks, Redshift, BigQuery, ClickHouse, MySQL / generic SQL, Snowflake.
Other sources Ingest connects for SaaS teams: Asana, Jira, Airtable, Clockify, Facebook Ads, Frankfurter.