MartinAI
August 27, 2026·10 min read

Connecting utility data to your energy tools: one clean dataset, many destinations

Energy teams feed utility bills and interval data into benchmarking, feasibility, BI, and emissions tools. The real bottleneck is getting clean, standardized data out of many utilities first. Here is how the shared foundation works.

If you run energy, facilities, or sustainability for a portfolio, you already know the pattern. The same utility bills and interval reads have to end up in four or five different places: a benchmarking tool for performance scores, a feasibility and measurement-and-verification model for projects, a business intelligence dashboard for the executive team, and an emissions calculation for reporting. Every one of those destinations wants the data in its own format, at its own granularity, with its own assumptions about units and time. And every one of them stalls for the same reason: nobody has a complete, clean, standardized dataset to feed them in the first place.

This article is the map for the rest of the cluster. It explains why the data foundation, not the destination tool, is where teams lose the most time, and it summarizes each place your utility data needs to go. Deeper pieces on each destination are linked as you read.

The bottleneck is upstream of every tool

The tools themselves are mature. The friction sits before them, in acquisition and preparation. A single portfolio can pull from dozens of utility accounts across electricity, gas, water, and steam, each with a different bill layout, a different meter numbering scheme, and a different way of expressing the same quantity. Interval data arrives as one file format from one utility and a completely different one from the next. Scanned PDF bills carry the numbers you need locked inside a page image.

Researchers who study building analytics keep landing on the same conclusion: the shortage is not analysis technique, it is usable measured data. A widely cited review of the barriers to building energy work points to data gaps, missing values, and the lack of standardized formats as the invisible barrier that slows everything downstream. If the foundation is incomplete or inconsistent, no benchmarking score, project model, or dashboard built on top of it can be trusted.

The core idea

Prepare the data once, correctly, and reuse it everywhere. A benchmarking score, a feasibility model, a BI dashboard, and an emissions figure can all draw from the same clean, standardized, complete dataset instead of four separate one-off exports.

One clean dataset, many destinations

Think of the utility data layer as a single source that each tool reads in its own way. The table below shows where the same underlying consumption record has to travel, and what each destination actually needs from it.

DestinationWhat it needsWhy it matters
BenchmarkingMonthly consumption and cost by meter, complete 12-month coverage, correct property and floor areaProduces a comparable performance score across the portfolio
Feasibility and M&VBaseline consumption, weather data, project period readsSizes savings and verifies them after a retrofit
BI dashboardsStandardized time series, tidy dimensions (site, meter, commodity, cost)Lets non-specialists explore trends and outliers
EmissionsElectricity and fuel quantities by location and periodTurns consumption into Scope 1 and Scope 2 figures

None of these destinations is hard once the input is clean. All of them are painful when it is not. That is why the work described in the rest of this cluster starts with acquisition and standardization, and only then moves to the specific tool.

Destination 1: benchmarking in ENERGY STAR Portfolio Manager

Benchmarking answers a simple executive question: how does this building compare to similar ones? ENERGY STAR Portfolio Manager, run by the US Environmental Protection Agency, is the most widely used tool for that. Its cleaned public dataset covers more than 150,000 US properties, and the free tool is also used to benchmark buildings in Canada against Canadian energy use intensity medians.

The value is well documented. In the largest analysis to date, properties that benchmarked consistently saved an average of 2.4 percent per year, roughly 7 percent over three years, with the lowest performers improving the most. But the tool only reflects what you feed it. Missing a month of gas data or misclassifying floor area quietly corrupts the score. That is a data-completeness problem, not a benchmarking problem. The deeper mechanics of pushing clean monthly data in are covered in automating ENERGY STAR Portfolio Manager.

Destination 2: feasibility and measurement-and-verification in RETScreen

When a project is on the table, the question shifts from how are we doing to what will this save. RETScreen, the free clean energy management software from Natural Resources Canada, is a standard tool for feasibility analysis and for measurement and verification after a project goes in. It is available in 36 languages and used by energy and facility professionals worldwide.

RETScreen models depend on a credible baseline, and a baseline is only as good as the historical consumption behind it. Gaps, estimated bills, and unit mismatches all distort the savings estimate. Getting a clean, continuous history into the model is the unglamorous prerequisite, and it is where the shared data foundation pays off. The full workflow is in using utility data in RETScreen.

Destination 3: BI dashboards in Power BI

Once the leadership team wants to explore the data themselves, it usually ends up in a business intelligence tool. Power BI is the common landing spot: it holds roughly 22 percent of the BI platform market and Microsoft has been named a Leader in the Gartner Magic Quadrant for analytics and BI platforms for nineteen consecutive years.

Power BI is only as useful as the model behind it. It expects tidy, standardized time series with consistent dimensions for site, meter, commodity, and cost. Energy teams routinely spend more time reshaping messy exports than building visuals. When the upstream data is already standardized, the dashboard becomes a modeling exercise instead of a cleanup project. See Power BI for utility and energy data for the data model patterns that hold up.

Destination 4: Scope 1 and Scope 2 emissions

The same consumption numbers also drive emissions reporting. Under the GHG Protocol, on-site fuel combustion becomes Scope 1 and purchased electricity becomes Scope 2. The location-based method multiplies electricity use by a regional grid emission factor, and in the US those factors come from the EPA's Emissions and Generation Resource Integrated Database (eGRID).

The calculation is arithmetic. The hard part is having complete, correctly attributed electricity and fuel quantities for every site and period, which is exactly what the shared foundation provides. Walk through the full method in turning utility bills into Scope 1 and 2 emissions.

What clean and standardized actually means

The phrase gets used loosely, so here is what it has to mean in practice before any downstream tool can rely on it.

  • Complete coverage: every meter, every period, with estimated or missing reads flagged rather than silently dropped.
  • Consistent units: energy, demand, and cost expressed the same way across every utility and commodity.
  • Standardized structure: one schema for site, meter, commodity, timestamp, quantity, and cost, regardless of which utility or bill it came from.
  • Reconciled reads: interval data and billed totals lined up so the two views agree.
  • Traceable source: every figure tied back to the original bill or interval file so it can be audited.

Do this once and every destination in this article gets easier. Skip it and you rebuild the same messy export four times, each with its own errors.

How MartinAI helps

MartinAI is a Canadian utility-data platform that focuses on exactly this foundation. It collects utility data from utility connections, Green Button feeds, and scanned or PDF bills, then cleans and standardizes it into one consistent dataset. From that single source, the same numbers can flow into benchmarking, feasibility and M&V models, BI dashboards, and emissions calculations without being re-keyed or reshaped for each one.

The point is not another dashboard. It is ending the manual data wrangling that sits in front of every energy tool your team already uses, so the people doing the analysis spend their time on decisions instead of on cleanup.

Read the deeper pieces next: benchmarking, RETScreen, Power BI, and emissions each get their own article in this cluster, and they all start from the same clean dataset described here.

Frequently asked questions

Why not just export data separately for each tool?

You can, but you end up cleaning the same messy utility data multiple times, once per destination, and each pass introduces its own errors. Preparing one clean, standardized dataset and reusing it keeps the numbers consistent across benchmarking, feasibility, BI, and emissions.

What is the single biggest cause of delay in energy analytics?

Data acquisition and preparation, not analysis. Studies of building energy work repeatedly identify missing data, gaps, and inconsistent formats as the main barrier, well before any tool-specific step.

Do I need interval data or are monthly bills enough?

It depends on the destination. Benchmarking and emissions can work from monthly totals, while diagnostics, load analysis, and detailed M&V benefit strongly from interval data. A good foundation captures both and reconciles them.

Can the same dataset feed both benchmarking and emissions?

Yes. The same consumption and cost records that produce a benchmarking score also supply the electricity and fuel quantities needed for Scope 1 and Scope 2 emissions, as long as the data is complete and correctly attributed to each site and period.