Synthetic Form 1040-NR - U.S. Nonresident Alien Income Tax Return Data

Synthetic training data — no real PII, fully coherent identities

Tax2023

Generate synthetic Form 1040-NR nonresident alien returns with foreign addresses, ITIN-style identifiers, and treaty-position fields. Requesting this form switches the simulated population to nonresident-alien identities, so every generated return is a coherent cross-border filer.

134

Fields per document

2

Pages

Tax

Category

What this document is

Form 1040-NR is the U.S. income tax return filed by nonresident aliens with U.S.-source income — foreign students and scholars, temporary workers, cross-border investors, and non-resident landlords. It mirrors the domestic 1040's income-through-tax-computation spine but adds foreign address blocks, country-of-residence and visa-status fields, treaty-based return positions, and a schedule for income not effectively connected with a U.S. trade or business.

Why generate synthetically

Cross-border tax documents are the hardest category to source for training: they are individually rare, they are concentrated in a handful of specialist providers, and every real example carries an identifiable foreign national's data. Synthetic 1040-NRs let you build a nonresident corpus at whatever volume your model needs, with foreign addresses and treaty fields populated, and with no privacy or export-control exposure at all.

What makes synthetic data useful

The 1040-NR carries a nonresident-alien requirement in its definition, so when you request it the generator constrains the entire simulated population to nonresident-alien identities with foreign addresses rather than filtering a domestic population down to nothing. Every return is therefore a coherent cross-border filer: address, country of residence, identifier format, and income mix all agree, and the income-to-tax computation resolves through the same arithmetic engine that drives the domestic 1040.

Training challenges

Foreign address blocks break every assumption a U.S.-trained extractor holds: there is no two-letter state code, postal codes do not match the five-digit ZIP pattern, and the province and country lines are separate fields that models trained on domestic forms routinely merge or drop. The identifier field accepts ITIN-formatted values that look like SSNs but are not, so a validator keyed on SSN area ranges will reject valid documents. The form carries a large checkbox population — filing status, treaty positions, and residency questions — packed into small targets, and its conditional logic is heavier than the domestic 1040's relative to its size, because whole sections activate or stay blank depending on residency and treaty answers.

Generate synthetic Form 1040-NR - U.S. Nonresident Alien Income Tax Return data

Start with 500 free credits. No credit card required.

Generate Now

Who uses this data

Cross-border and expatriate tax-prep providers (Sprintax-class products), global-mobility and relocation platforms, university international-student tax offices, withholding-agent compliance tooling at banks and brokerages, and KYC teams that have to read non-U.S. addresses and identifiers correctly. It is also the sharpest available out-of-distribution test for any extractor that assumes a domestic address schema.

Document complexity profile

134 fields across 2 pages: 52 currency amounts, 49 text, 28 checkbox targets, and 5 identifier fields. 134 annotation relations. The binding graph is dense for the form's size — 23 conditional bindings, 5 arithmetic bindings, and 91 function calls with a maximum expression depth of 3 — because residency status and treaty answers gate whole regions of the return rather than individual lines.

Key stats from our synthetic corpus

Quantitative characteristics of the Form 1040-NR - U.S. Nonresident Alien Income Tax Return documents our generator produces.

MetricValueDetail
Checkbox targets2828 of the form's fields are checkboxes — filing status, residency questions, and treaty-position answers. Checkbox state on the 1040-NR is load-bearing: it decides which downstream sections carry values at all.
Conditional bindings2323 conditional bindings gate sections of the return on residency and treaty answers, the highest conditional density relative to field count among the 1040-family forms we ship.
Function calls in bindings9191 FORMAT and IF function calls drive identifier formatting, foreign-address assembly, and the income-to-tax chain, so a generated return is internally consistent rather than field-wise random.
Annotation relations134134 label-to-value relations are exported per document, giving relation-extraction models full supervision over a layout whose label vocabulary differs substantially from the domestic 1040.
General-population eligibility0%No identity in a default domestic corpus qualifies for the 1040-NR — the form requires a nonresident alien. Requesting it reconfigures the population to nonresident identities, which is why a 1040-NR job yields a full corpus while a mixed general-population job yields none.

How this document co-occurs with others

Rates at which identities in our corpus that produce a Form 1040-NR - U.S. Nonresident Alien Income Tax Return also produce other documents.

CorrelationRateDetail
Domestic counterpart, same computation spine100%The domestic Form 1040 has no residency requirement, so it co-generates in any 1040-NR job. Pairing the two gives a resident/nonresident contrast set over the same income-to-tax structure — the training pair for residency classification.
Co-generates with a W-9100%The W-9 declares no requirements and co-generates with any 1040-NR identity. In practice a nonresident would file a W-8 series form instead, which makes this pair a useful adversarial case for models that classify tax-identity documents.
Co-generates with a 4868 extension100%Form 4868 has no residency requirement and co-generates in a nonresident job, producing the extension-plus-return document pair that cross-border preparers file constantly.
Co-generates with a CMS-1500 claim100%The CMS-1500 carries no residency requirement either, so a nonresident identity can also produce a healthcare claim — the visiting-scholar and international-student case that healthcare RCM systems handle badly.

Structural figures above (field counts, checkbox and relation counts, binding densities) are derived from the shipped Form 1040-NR definition in the SymageDocs form library. Unlike the domestic forms, the population figure is reported against a default domestic corpus of 1,000 identities, in which no identity is nonresident by construction; a 1040-NR job reconfigures that population. No real taxpayer data was used at any stage.

Frequently asked questions

What data format do synthetic Form 1040-NR documents include?
Each generated identity produces a rendered PDF plus a structured JSON annotation file with bounding boxes and ground-truth values for all 134 fields across both pages, including the foreign address and treaty-position regions. COCO, YOLO, FUNSD, and BIO/NER exports come from the same job.
Do you actually generate nonresident identities, or just relabel domestic ones?
Actual nonresident identities. The form definition declares a nonresident-alien requirement, and the generator resolves form requirements into the population configuration before it simulates anyone — so a 1040-NR job produces identities with foreign addresses and nonresident attributes from the start, not domestic identities with a checkbox flipped.
Will a general-population corpus contain 1040-NRs?
No, and that is by design. A default synthetic population is domestic, so no identity in it qualifies for the 1040-NR. Requesting the form is what switches the population. If you want mixed domestic and nonresident data, generate two jobs and combine them — that also gives you a clean labeled split for residency classification.
How does labeling work for the foreign address fields?
Foreign country, province, and postal-code fields are annotated as distinct labeled regions with their own bounding boxes, not as one address blob. That is what lets a model learn the non-U.S. address schema instead of forcing every document into a street/city/state/ZIP shape.
Can I use this data commercially?
Yes. All synthetic data is generated from statistical models, contains no real PII, and is licensed for commercial use including ML model training and benchmarking.

Related Tax Forms