Data Integration Cost: ETL vs iPaaS vs Custom

Published by Bles Software, a custom software and AI company based in Yehud-Monoson, Israel, building web apps, AI agents and API integrations for clients in Israel, the US, the UK and the EU.

ETL, iPaaS or custom: the billing unit decides the bill

Data integration pricing is not one market. Three vendors can quote the same monthly figure for moving the same tables and bill you wildly differently a year later, because they do not meter the same thing. One charges per row that changed. One charges per credit, where a credit buys rows from an API and gigabytes from a database. One charges per task. One charges per connected flow regardless of volume. Choose the meter before you choose the vendor, because the meter is what grows.

Every price below was read on 9 October 2026 from the vendor's own page, with the unit it bills in stated next to it.

The managed ELT meter

Now read the definitions, which is where the money is.

Fivetran defines a Monthly Active Row as a distinct row synced from source to destination in a calendar month, counted once per month even if it syncs many times, and it charges inserts, updates and deletes, with deletes counting from 1 January 2026 and history-mode changes counting as paid MAR [7]. That definition is generous on frequency and unforgiving on churn: a table that updates every row daily costs the same as one that updates each row once, but a wide slowly-changing dimension with history on is expensive.

Airbyte's credit converts differently by source type: APIs and custom sources cost 6 credits per million rows, databases and files cost 4 credits per 1 GB [8]. So the same pipeline is priced by rows on one leg and by bytes on the other.

Portable is the outlier worth knowing about: it prices flows, not volume, and states unlimited data volumes [5]. For a small number of very large feeds that is the cheapest shape on this page. For forty small feeds it is the most expensive.

Matillion publishes no dollar price, only a consumption-based credit system across Developer, Teams and Scale tiers [9].

The self-hosted option, and what it actually asks of you

Open source removes the vendor meter and replaces it with your own infrastructure and your own people. Airbyte's own quickstart states the requirement plainly: Docker, the abctl CLI, and 4 or more CPUs with at least 8 GB of memory, or a low-resource mode on 2 CPUs and 8 GB without the Connector Builder [10]. dbt Core is Apache 2.0 and free to use [11]; the managed platform is not, with Developer free at 1 seat and 3,000 successful models built per month and Starter at $100 per user per month with five developer seats, 15,000 models and 5,000 queried metrics a month [12].

Orchestration has its own meter. Dagster+ is $10 per month plus $0.040 per credit on Solo and $100 per month plus $0.035 per credit on Starter, with serverless compute at $0.010 per minute, where a credit is the sum of asset materializations and ops executed [13]. Prefect Cloud is free on Hobby, $100 per month on Starter and $100 per user per month on Team [14]. Meltano's managed tier bills compute hours rather than rows and publishes no dollar figure [15].

The iPaaS middle, and the half of it that has no price

Celigo at least publishes its philosophy rather than its price: pay for endpoints and flows, not per task or transaction [18]. Make defines one operation as one credit for non-AI apps [19]. When four vendors in a category publish nothing, the comparison you can actually make is between the two that do and your own hours.

The cost that eats the budget after the pipeline works

This is the line most data integration estimates leave out entirely, and it is usually larger than the pipeline. Loading data is cheap. Querying it is not.

Note the regional spread on Snowflake: the same Enterprise credit is $3.00 in US East and $3.90 in EU Frankfurt [20], a 30 per cent difference decided by where you put the account. The credit consumption table we read carries an effective date of 13 May 2025 [20], so check it against your own contract before budgeting from it.

Egress, the line nobody quotes you

Cross-cloud pipelines pay to leave. AWS charges $0.09 per GB for the first 10 TB out of N. Virginia, falling to $0.085, $0.07 and $0.05 at higher tiers, with 100 GB a month free [23]. Google Cloud Premium Tier to North America is $0.12 per GiB from 1 GiB to 1,024 GiB, then $0.11, then $0.08, with inter-region within North America at $0.02 per GiB [24]. Azure gives the first 100 GB a month free, then $0.087 per GB for the next 10 TB [25]. Cloudflare R2 publishes egress to the internet as free, with storage at $0.015 per GB-month [26].

Moving a terabyte a day out of one cloud into another's warehouse is a five-figure annual line on its own at those rates [23][24][25]. It is also the easiest one to design away.

The people line, and the one survey figure worth quoting

The US Bureau of Labor Statistics has no separate code for data engineers, so the two closest are database architects (15-1243) at a mean $69.44 per hour and $144,440 a year, and software developers (15-1252) at a mean $71.20 per hour and $148,100 a year, both from the May 2025 occupational survey [27].

On how much of that time maintenance takes, the honest answer is that the available figures come from a vendor. Fivetran's Enterprise Data Infrastructure Benchmark Report 2026, published 26 March 2026, reports 53 per cent of engineering time spent on pipeline maintenance and $2.2 million a year in engineering labour, from 500 senior data and technology leaders at organisations of 5,000 or more employees, fielded in the fourth quarter of 2025, with a 95 per cent confidence level and a stated margin of error of 4.4 per cent [28]. Fivetran fielded it itself and sells the alternative, so weigh it accordingly. We quote it because it is the only figure of its kind with a published methodology, not because it is neutral.

How to decide

Pick by meter and by churn, in this order. If your volume is large and your feed count is small, a flow-priced vendor wins [5]. If your feed count is large and each is small, a row or credit meter wins [1][2]. If your data churns heavily with history, read Fivetran's MAR definition twice before signing [7]. If you have engineers and more than a handful of non-standard sources, self-hosting plus dbt Core is genuinely cheaper [10][11], and the cost moves onto your payroll [27] where it is visible.

Then price the warehouse [20][21][22] and the egress [23][24] before you sign anything, because those two outlive every pipeline decision on this page.

Sources

More Costs and Timelines from Bles Software