← Ground Truth

Normalizing revenue across US GAAP, JP GAAP, and K-GAAP — the hard parts

OpenFilings··8 min read
XBRLcross-GAAPnormalizationEDINETDART

Everyone agrees revenue should be comparable across markets. Almost no cross-market spreadsheet actually is.

The labels translate on the surface. They don't map cleanly underneath. A developer who parses Toyota's EDINET Yūho and Apple's 10-K and puts both revenue numbers into the same column has made a meaningful comparison only if they have also confirmed: consolidated or parent-only? Financial services included or excluded? Net of intercompany eliminations? Which XBRL tag was chosen by the filer?

Most data vendors don't tell you. Most models can't tell you. The error is silent and structural.

The label problem

Here is what "revenue" looks like across the three systems we'll cover:

MarketFiling systemTypical primary labelStandard
US SECEDGARRevenues or RevenueFromContractWithCustomerExcludingAssessedTaxUS GAAP (ASC 606)
Japan EDINETEDINET売上高 (uriage-daka)J-GAAP or IFRS
Korea DARTOpenDART매출액 (maechul-aek)K-GAAP or K-IFRS

The XBRL tags differ even within a single standard. The SEC accepts at least a dozen valid revenue tags depending on whether the filer uses ASC 606, whether it nets assessments, and whether it uses a custom extension. EDINET's XBRL taxonomy for J-GAAP and the IFRS taxonomy it also accepts are structurally different. DART's Korean XBRL uses Korean-language account codes (ifrs-full:Revenue in the international taxonomy, or domestic equivalents for K-GAAP filers).

Parsing the XML/XBRL correctly gets you the number. Knowing whether that number is the right number for a cross-market comparison is a separate problem. Here are the five places where it breaks.

Problem 1 — Gross vs net revenue

US GAAP ASC 606 requires companies to report revenue on a gross basis when they are the principal in a transaction — they control the goods or services before delivery — and on a net basis when they are an agent. The distinction matters enormously for marketplace businesses. An e-commerce platform that guarantees delivery, holds inventory risk, and sets prices is principal and reports gross. One that merely connects buyers to sellers is agent and reports net commission.

J-GAAP 売上高 is typically gross, but the interpretation of agent/principal is applied under Japanese accounting guidance which, until the adoption of IFRS or JMIS, had different thresholds. J-GAAP has historically been more permissive about gross reporting — some Japanese platforms that would be net reporters under US GAAP or IFRS report gross under J-GAAP.

K-GAAP has its own guidance, though K-IFRS (adopted for listed company consolidated statements) aligns with IFRS 15.

The practical effect: two companies with identical underlying transaction volumes can report revenues that differ by 3× because of this single agent/principal determination. When comparing a US SaaS company to a Japanese platform business, checking this is not optional. [1]

Problem 2 — Consolidated vs parent-only

Toyota Group includes Toyota Motor Corporation (the parent), Toyota Financial Services, Toyota Tsusho, Daihatsu, Hino Motors, and many others. Its consolidated Yūho for FY2025 reports ¥48.0 trillion in revenue. The parent-only financial statements — also included in the Yūho as required by J-GAAP — report a much smaller number that excludes subsidiary revenues.

Samsung Electronics is even more extreme. The consolidated K-IFRS financial statements include Samsung Display, Samsung SDI, Samsung Electro-Mechanics, and Samsung SDS. The parent-only K-GAAP statements include none of these.

EntityStatement typeFY revenue (approx)
Toyota (consolidated, IFRS)FY2025 Yūho¥48.0 trillion
Toyota (parent-only)FY2025 Yūhosubstantially lower
Samsung (consolidated, K-IFRS)FY2024 annual₩300.9 trillion
Samsung (parent-only, K-GAAP)FY2024 annuallower

The same EDINET or DART document contains both consolidated and parent-only statements. A naive parser that grabs the first 売上高 or 매출액 it encounters may pick up either one depending on document structure. For any global comparison, you want consolidated.

How to check: the XBRL context for a consolidated statement will have a ConsolidatedMember or entireGroup dimension. Parent-only statements use different context identifiers. OpenFilings extracts only the consolidated figure. [2]

Problem 3 — Financial services embedded in industrials

Toyota Financial Services (TFS) — car loans, dealer financing, operating leases — is fully consolidated into Toyota's revenue and operating income. This creates a comparability trap when doing automotive peer analysis.

Under Toyota's IFRS presentation, interest income from financing operations flows through "net revenues" at the segment level but shows up in "non-operating income" at the consolidated P&L level — or sometimes in revenues depending on the fiscal year and restructuring of segment reporting. The precise treatment has varied across Toyota's IFRS transition.

For FY2024 (ended March 2024), Toyota's operating_income of ¥5.35 trillion on ¥45.1 trillion revenue implies an 11.9% operating margin. A pure automotive OEM without a captive finance arm (say, Stellantis or Renault) at 11% operating margin is not the same — Toyota's margin includes the TFS contribution, Stellantis's does not.

This problem exists whenever a manufacturing company has a large embedded financial services division. The solution is not to strip TFS out (you usually can't from the top-level filing), but to note the embedded segment and adjust comparisons accordingly, or to use segment-level data where available.

Problem 4 — XBRL extension tags

XBRL was designed to produce machine-readable, comparable financial data. In practice, US filers use extension tags — custom XBRL elements beyond the standard taxonomy — for a significant portion of their disclosures. The SEC has repeatedly warned about overuse of custom extensions, but the practice persists.

A US technology company might report revenue as CompanyNameTotalNetRevenues — a custom extension tag — rather than the standard us-gaap:Revenues. A naive parser that only looks for standard taxonomy tags will miss this revenue entirely and report zero.

In Japan, EDINET's XBRL uses a hierarchical taxonomy where 売上高 can appear under different parent nodes depending on whether the filer uses the 2024 taxonomy revision or an earlier one. The node path changes across taxonomy years.

In Korea, DART's XBRL uses Korean-language account code labels that must be mapped to canonical equivalents. The label 영업수익 (operating revenues, used by financial companies) is not the same line item as 매출액 (sales revenues, used by industrial companies) — though both could colloquially be called "revenue."

The implication: a working revenue extractor needs to handle standard tags, extension tags, segment rollups, and Korean/Japanese taxonomy navigation simultaneously. This is not a one-afternoon project.

Problem 5 — Period and currency misalignment

Apple's FY2024 10-K covers the fiscal year ended September 28, 2024 and reports in USD.

Toyota's FY2025 Yūho covers the fiscal year ended March 31, 2025 and reports in JPY (millions).

Samsung's FY2024 annual report covers the fiscal year ended December 31, 2024 and reports in KRW (millions).

Three things are wrong with a naive side-by-side comparison:

  1. Period overlap: Apple's "FY2024" (Oct 2023 – Sep 2024) and Samsung's "FY2024" (Jan–Dec 2024) share three quarters but not the same year-end. Revenue for Q4 calendar 2024 appears in Samsung's FY2024 and Apple's FY2025.

  2. Currency: ¥1 ≠ ₩1 ≠ $1. JPY/USD ≈ 150. KRW/USD ≈ 1,380. A revenue comparison without explicit FX handling or currency labels is meaningless.

  3. Scale notation: Japanese EDINET reports values in millions of yen (百万円). Korean DART reports in millions of won. US SEC typically reports in thousands. A parser that doesn't account for the denominator unit will produce figures off by 1,000×.

The canonical approach is to store currency, period_end, fiscal_year, and unit_multiplier alongside every extracted number — and to expose these fields to any downstream consumer that wants to do cross-market comparison.

The canonical vocabulary approach

A canonical vocabulary is a single set of field names — revenue, operating_income, net_income, operating_cash_flow, capex, free_cash_flow — that map consistently to whatever the primary filing calls the equivalent line item, given its GAAP, country, and consolidation scope.

Building this mapping requires:

  1. Per-GAAP alias tables: mapping US-GAAP XBRL tags, J-GAAP node paths, and Korean DART account codes to canonical names
  2. Consolidation scope rules: always prefer consolidated, document the cases where parent-only is the only available figure
  3. Period and currency tagging: attach period_end, fiscal_year, currency, and accounting_standard to every extracted value
  4. Extension tag resolution: catch common custom extensions by matching on semantic patterns, not just exact tag names
  5. Regression testing: run against known correct outputs from audited filings (e.g., Apple's 10-K where the correct revenue is publicly reported and verifiable)

The alternative — scraping the numbers from PDF or HTML renderings — is faster to implement and structurally unreliable for anything that requires precision. PDF layouts change across years, table parsers break on unusual formatting, and numeric extraction from unstructured text produces silent errors at scale.

Three verified examples

Here are canonical KPIs extracted from primary filings, with explicit metadata:

Apple Inc. (AAPL) — FY2024
Standard: US GAAP · Registry: SEC · Period: Oct 2023 – Sep 28, 2024 · Currency: USD

FieldValue
revenue$391,035,000,000
operating_income$123,216,000,000
operating_margin31.5%
free_cash_flow$108,807,000,000

Source: 10-K, accession 0000320193-24-000123, EDGAR.


Toyota Motor Corporation (7203) — FY2025
Standard: IFRS · Registry: EDINET · Period: Apr 2024 – Mar 31, 2025 · Currency: JPY

FieldValue
revenue¥48,036,704,000,000
operating_income¥4,795,586,000,000
operating_margin10.0%

Source: Yūho, doc ID S100VWVY, EDINET.


Samsung Electronics (005930) — FY2024
Standard: K-IFRS · Registry: DART · Period: Jan–Dec 31, 2024 · Currency: KRW

FieldValue
revenue₩300,869,000,000,000
operating_income₩32,726,000,000,000
operating_margin~10.9%
operating_cash_flow₩44,617,000,000,000
capex₩57,097,000,000,000
free_cash_flow~₩−12,480,000,000,000

Source: Annual business report, OpenDART.


These numbers are comparable because they carry their metadata. You know what period each covers, what standard was used, what currency it's in, and where to verify the source. Strip the metadata and they become three numbers in a column that looks comparative but isn't.


Sources & notes

  1. The gross/net determination for platform businesses is one of the most consequential and contested accounting policy decisions in tech. Uber, for example, reports net revenue (commissions) under ASC 606 for its ride-hailing segment but gross revenue for some other segments. Rakuten reports 売上高 on a gross basis for its marketplace, which inflates its revenue relative to a US peer reporting on a net basis for the same economic activity.
  2. XBRL context identifiers for consolidated vs. parent-only differ by filing system. In J-GAAP EDINET filings, consolidated data uses context FilingDateInstant_NonConsolidatedMember vs. FilingDateInstant_ConsolidatedMember (or their duration equivalents). In K-IFRS DART filings, the Korean XBRL taxonomy uses separate report sections labelled 연결 (consolidated) vs. 별도 (separate/parent-only). Parsers that do not filter by context dimension will mix the two.