What happens to the rows that were fine? A bulk-import correctness axis proposal (September 2026)
The only file an importer is ever given during development is one the exporter wrote, so it cannot fail. This proposal scores what is actually in the table when a real file stops at row 7,412, who is told which row it was, and what happens when the corrected file is uploaded again.
Updated on September 30, 2026
On this page
Every AI app builder will wire up a form that creates one record. Ask for a way to load a thousand records from a file and most of them will produce something that works the first time you try it. The interesting question is not whether the import succeeds. It is what is in the database when it does not: how many rows landed, which one stopped it, whether the person who uploaded the file can tell, and what happens when they fix the bad line and upload it again. That is the question this axis proposes to measure. We are not scoring builders today. We are publishing a candidate axis, its rubric, its posture levels and a reproduction protocol, and opening all of it for comment before it enters the composite.
Quick Answer
Bulk-import correctness is a proposed BuilderProof benchmark axis, drafted September 30, 2026, that scores what the code an AI app builder emits does when a person uploads a multi-record file. It is measured from the exported project and from a fixed file battery: does a failure at row 7,412 roll back the 7,411 rows before it, or leave them committed with nothing recording where it stopped; is the failing row identified by number and column, or only as a count; is re-submitting the corrected file safe, or does it duplicate everything that already landed; and is the file decoded on a determined encoding rather than an assumed one. The rubric weights seven signals. The cohort under consideration is the five commercial builders we already track. This page is an axis proposal open for community edits, not a leaderboard. No builder is named, scored or assigned a posture here.
Why this is not already covered
This axis sits between three we have already published, and each boundary is worth stating rather than assuming, because in one case the neighbouring rubric does not merely fail to see this failure. It scores it backwards.
Against input-validation and data-integrity
Our input-validation and data-integrity proposal asks whether an emitted write endpoint rejects a malformed request body at the server boundary. Its rubric carries a signal it calls rejection semantics, and its protocol is explicit about what counts: a 400 or 422 with a structured error is a pass, and a 200 that persists the row, a 500, or a silent coercion is a fail.
That is correct for the thing it measures, and it inverts on a file of ten thousand rows.
You cannot reject a ten-thousand-row upload because row 7,412 has a date in the wrong format. Returning a 4xx and writing nothing is the neighbour's pass condition and it discards 9,999 rows that were fine. Meanwhile a 200 that persists rows, which is the neighbour's fail condition, is close to what a well-built importer does: it applies what it can, refuses what it cannot, and tells you which was which. The two rubrics do not disagree because one of them is wrong. They disagree because the unit changed. A single request has one verdict available to it. A file has as many verdicts as it has rows, and the interesting engineering is entirely in how those are reconciled into one response.
That is the sharpest boundary we have drawn between two of our own axes, and it is worth being precise about the direction: on a bulk-import path, the input-validation rubric scores the correct behaviour as a failure and the destructive behaviour as a pass. Neither axis needs amending. They need scoping, which is what this section is.
The other half of the boundary is the battery. That axis sends a fixed set of single payloads: a missing required field, a wrong-typed field, an oversize string, an unexpected extra field. Every one is one request carrying one record. There is no payload in that battery which can be partly valid, and partial validity is the entire subject here.
Against file upload and media handling
Our file-upload and media-handling proposal draws its own boundary in its own words, and it excludes this axis explicitly. Its scope section states that it covers binary objects and their access path, not rows. Its six signals are object-level authorization, server-side content verification, size-ceiling handling, object-key generation, serving posture and lifecycle coupling. Every one is a property of an opaque object and the service it comes to rest in.
That axis ends where the bytes are safely stored. This one starts by opening them. A project can score at the top of the upload axis, with a private bucket, owner-scoped policies, signature-checked content and a generated object key, and still parse that object in a loop that stops dead halfway through and leaves the table in a state nobody recorded.
Against data export and handoff
Our data-export proposal scores the file the application writes. This one scores the file somebody else wrote. Our print-output proposal scores a third direction across the same boundary, a page rendered for a person with no program at all, and the three together are the only places records leave or enter this application as an artefact somebody carries somewhere else. That sounds like a tidy mirror and it is not, because the two are not symmetric in the way that matters: the exporter controls every byte of its output, and the importer controls none of its input.
The sharper point is that the export axis's best remedy manufactures this axis's blind spot. Its seventh signal rewards a project whose export round-trips through a conformant reader, on the reasoning that otherwise the only evidence the file is well formed is that it opened once, in one program. That is good advice. It also means the one file a project is most likely to have fed through its own importer is the file its own exporter produced: same encoding, same delimiter, same columns in the same order, no missing values, no surprises. A project that scores full marks on that signal has proved its importer works on precisely the input that can never fail. We return to this under the named trap, because it is the trap.
Against outbound-call failure
Our outbound-call failure proposal covers idempotency, and re-uploading a file after a failure is obviously an idempotency question, so the boundary needs stating. That axis asks whether one operation can be retried safely when the caller does not know whether it landed. This one asks what happens when an unknown prefix of N operations definitely did land, the caller knows roughly how many, and the retry is a different file from the original because a human edited one line of it. An idempotency key protects a single call. It says nothing about which of ten thousand rows it is safe to send again.
Named without links, to stay inside the internal-link budget
Three further neighbours compose with this axis rather than bounding it, and saying so is more useful than drawing a line that is not there. Our background-work correctness proposal asks whether a scheduled or deferred job runs at all; a large import is very often such a job, and the two questions stack, because a job that never runs and a job that dies at row 7,412 leave the table in different states and only one of them is this axis's subject. Our concurrent-write safety proposal asks what happens when two writers contend for one row; this asks what happens to one writer applying ten thousand rows nobody else is touching. Our pagination proposal asks whether a large collection can be read completely; this asks whether one can be written completely.
Four documented facts that make this measurable
The reason this axis can be scored rather than merely discussed is that the database underneath most of these builders has already made every one of these decisions explicit, named them, and chosen defaults. We do not have to invent a standard of care. We can read one.
1. The default is to fail the whole command, and the database names the alternative
PostgreSQL's
COPY command is the canonical bulk loader, and its documentation states the position in one sentence: "By default, COPY will fail if it encounters an error during processing. For use cases where a best-effort attempt at loading the entire file is desired, the ON_ERROR clause can be used to specify some other behavior."
Two things are settled there. The first is that all-or-nothing is the default, so a project that wants partial application has to ask for it. The second is that the database calls the alternative a best-effort attempt at loading the entire file, which is a description of an intent rather than of a mechanism. That phrase is doing real work: it concedes that partial loading is a product decision, not a correctness property.
The option itself has exactly two values and the documentation is equally plain: "An error_action value of stop means fail the command, while ignore means discard the input row and continue with the next one. The default is stop."
This matters for the rubric because it means a generated importer sits somewhere on a ladder the database has already drawn, and the position is checkable rather than aesthetic. A project that commits each row as it parses it has not chosen ignore. It has landed in a third place the database does not offer, where a failure stops the loop wherever it happens to be and everything before that point is already committed.
2. Once you opt out, the default tolerance is unlimited
There is a second rung, and it was added recently. PostgreSQL 18's COPY gained a REJECT_LIMIT clause, documented as follows: "Specifies the maximum number of errors tolerated while converting a column's input value to its data type, when ON_ERROR is set to ignore. If the input causes more errors than the specified value, the COPY command fails, even with ON_ERROR set to ignore."
The sentence that follows is the one worth quoting to anybody building an importer: "If not specified, ON_ERROR = ignore allows an unlimited number of errors, meaning COPY will skip all erroneous data."
So the ladder has three rungs and two of them are opt-in. Fail on the first bad row. Or tolerate bad rows, with no ceiling, so a file that is entirely garbage loads zero rows and does not raise. Or tolerate bad rows up to a stated bound, which is the only rung on which a file that is mostly broken is distinguishable from a file that is mostly fine.
A generated importer that catches per-row exceptions and continues has implemented rung two without implementing rung three. That is a real posture, it is common, and it is the one where a user uploads the wrong file entirely, sees no error, and finds an empty table.
3. The row number is a separate opt-in from the row skipping
This is the fact we found most surprising, and it is the anchor for the second-heaviest signal in the rubric.
Skipping a bad row and telling anybody which row it was are two different options. PostgreSQL's documentation separates them: "A NOTICE message containing the ignored row count is emitted at the end of the COPY FROM if at least one row was discarded. When LOG_VERBOSITY option is set to verbose, a NOTICE message containing the line of the input file and the column name whose input conversion has failed is emitted for each discarded row."
By default you are told how many rows were discarded. You are not told which. The line number of the input file and the name of the offending column arrive only if a second, separate option is set to verbose. PostgreSQL 18 adds a third value in the other direction, where setting it to silent means no message is emitted regarding ignored rows at all.
If the database that has thought about this problem for thirty years makes the row number opt-in, a generated importer is unlikely to volunteer it. And the difference is not cosmetic. "23 rows could not be imported" is not actionable on a ten-thousand-row file. "Row 7,412, column joined_on, could not be converted" is a thing a person can fix in ninety seconds.
4. The mechanism has a scope that almost nobody checks
Read the ON_ERROR definition once more, because the first clause is a limit and it is easy to skim past: "Specifies how to behave when encountering an error converting a column's input value into its data type."
ON_ERROR = ignore covers type-conversion errors. It does not cover a unique-constraint violation, a foreign-key violation or a failed check constraint, because none of those is a conversion failure. So the one mechanism the database offers for partial loading does not apply to the single most common real-world import failure, which is a duplicate email address in a customer list.
We are not treating that as a defect in PostgreSQL. The scope is stated plainly and the behaviour is correct. We are treating it as evidence that partial-application semantics are genuinely hard, that the layer underneath solves only part of the problem, and that the rest has to be solved in the emitted application or not at all. An importer that wraps COPY and believes it has handled bad rows has handled one class of them.
And one fact about the file itself
The fourth signal is about decoding, and it needs a different source, because the failure there is not that the import stops. It is that it does not.
The web platform's decoder has two relevant defaults and MDN states both. The encoding label "Defaults to
utf-8". And on error handling: "A boolean value indicating if the TextDecoder.decode() method must throw a TypeError when decoding invalid data. It defaults to false, which means that the decoder will substitute malformed data with a replacement character."
Put those together. A file that a spreadsheet application wrote on a Windows machine in a legacy single-byte encoding, handed to a decoder nobody configured, will not raise. It will decode, every accented character will become U+FFFD, and the import will report complete success. There is no error to catch, no row to skip, and no count to report, because from the parser's point of view nothing went wrong. Every name with an accent in it is now stored incorrectly and the only party who will ever find out is the person whose name it is.
The same decoder handles the other half of the problem by default. On the byte order mark: "It defaults to false, which means that the byte order mark will be skipped over when decoding and will not be included in the decoded text." So the platform silently handles one of the two encoding hazards and silently mishandles the other, which is worth knowing before you decide how much of signal 4 is the builder's responsibility.
The Encoding Standard is blunter about why this is a mess at all: "The problems outlined here go away when exclusively using UTF-8, which is one of the many reasons that is now the mandatory encoding for all things." True, and an importer does not get to choose what its users upload.
The proposed rubric
Seven signals. Weights are a starting point for community revision, not a settled formula.
Scroll to see more
| Signal | Weight | What a failing case looks like |
|---|---|---|
| Partial application is a decision, not an accident | 22 | Rows are committed one at a time as they are parsed, so a failure at row 7,412 leaves 7,411 rows in the table, nothing records where it stopped, and neither the user nor the code can say what the file did |
| The failing row is identified, not merely counted | 20 | The response is a bare failure or a count of rejects, with no line number, no column name and no offending value, so a person with a ten-thousand-row file has nowhere to start |
| Re-submitting the corrected file is safe | 16 | Rows are inserted with a generated key and nothing in the file identifies a record, so uploading the fixed file after a partial failure duplicates everything that already landed |
| The encoding is determined rather than assumed | 14 | The file is decoded as UTF-8 with the default non-throwing decoder, so a legacy-encoded export succeeds, stores replacement characters in every accented value, and reports no error at all |
| Columns are bound by name, not by position | 12 | Fields are consumed in file order, so a file with the same columns in a different order imports cleanly into the wrong fields, and only a human reading the data afterwards can detect it |
| The file is bounded before it is parsed | 9 | The whole upload is read into memory and parsed inside the request, so a large file exhausts the runtime or times out, and what was written before that is undefined |
| A rehearsal exists | 7 | The only way to discover what a file will do is to let it do it, because nothing validates the whole file and reports before anything is written |
Three weighting decisions worth arguing about now
Why partial application outranks diagnosis. Signals 1 and 2 are close, and a reasonable person would swap them. We put partial application first because its failure is the only one on the list that leaves the database in a state no party can describe. A missing row number is expensive and recoverable: the data is consistent and somebody can go and look. A half-applied import with no record of where it stopped is a state from which the correct next action is genuinely unknown, including to the person who wrote the code.
Why re-submission safety is worth 16 and not more. It is the signal most likely to be argued up, because duplicate records are the visible, embarrassing failure that a customer reports. We hold it at 16 because it is partly downstream of signal 1: an importer with true all-or-nothing semantics has nothing to duplicate, which is the interaction described under limitations.
Why rehearsal is last. A dry run is the most satisfying feature on this list and the least load-bearing. It is a convenience layered on top of signals 1 and 2; if those are right, a failed import is cheap and reversible, and a preview is a nicety. If they are wrong, a preview is a second code path that can disagree with the real one, which is its own failure mode. Seven points is deliberate and we would defend it, but it is the weight we most expect to be challenged.
Four postures
For legibility, the weighted score maps to one of four levels.
Level 0, optimistic. Rows are parsed and written in a loop. The first failure ends the loop wherever it happens to be. Whatever was written stays written. There is no count, no row number and no record of the stopping point, so the state of the table after a failed import is not derivable from anything the application kept.
Level 1, reported. The import still applies partially, but it reads the whole file rather than stopping at the first problem, and it returns a summary: this many rows were created, this many were rejected. The user learns that something failed. This is the rung PostgreSQL calls ignore with no reject limit, and it is where most working importers sit.
Level 2, bounded. The outcome is one a person can act on. Either the import is all-or-nothing, so a failure leaves the table exactly as it was, or it is resumable and records a position it can continue from. Rejected rows are identified individually, by line number and field. A ceiling exists above which the import is treated as the wrong file rather than as a file with some bad rows.
Level 3, rehearsed and re-runnable. Level 2 plus a validation pass that reports every failing row before anything is written, an encoding and header contract that is determined rather than assumed, and a natural key that makes re-submitting a corrected file converge on the right state rather than doubling it.
The rungs are deliberately uneven, and the uneven one is 1 to 2. Going from level 0 to level 1 is a loop change and an accumulator: real work, but a single sitting. Going from level 1 to level 2 requires deciding what the unit of atomicity is, which is a design question about the product rather than about the code, and no amount of care inside the parsing loop answers it.
The measurement protocol
Ten steps. The point of writing them down is that two people running them on the same export should get the same answer.
- Generate a reference app from the fixed prompt suite, one generation per builder, no manual edits. The prompt asks for a way to load records from a file. It does not mention validation, encodings, duplicates or rollback, because a prompt that named the defences would measure instruction-following rather than default posture.
- Export the project and record the import path verbatim: the route that receives the file, the parser, the loop or statement that writes, and every transaction boundary in between. Record whether a transaction is opened at all.
- Import a clean file of 500 valid rows. Record the response body, the response time and the row count. This establishes that the feature exists and gives the baseline every later step is compared against.
- Import the file the application's own export produced. This step carries no information about the axis and is included so that nobody mistakes it for evidence. It is the file the importer is guaranteed to handle, and running it is how the trap below stays invisible. We record the result and score nothing from it.
- Import a 500-row file in which row 400 has a value that cannot be converted to its column's type. Record three things separately: the HTTP status, the response body, and the number of rows now in the table. The row count is the measurement; the status is not.
- Import a 500-row file in which row 400 violates a uniqueness constraint rather than a type. Compare against step 5. A build that handles one and not the other has implemented the layer underneath without implementing the rest, which is the scope limit documented above.
- Correct the offending row in the step 5 file and upload the corrected file. Count the rows again. The result is one of three: the table converges on 500, it reaches 899, or the upload is refused. All three are legitimate postures and they score differently.
- Import a file encoded in a legacy single-byte encoding containing accented characters. Record whether anything failed, and separately whether the stored values are correct. These come apart, and the interesting case is the one where nothing failed.
- Import a file whose columns are the expected columns in a different order, with the header row intact. Record whether the data lands in the right fields.
- Import a file large enough to exceed the platform's documented request or execution ceiling. Record what the user sees and, separately, how many rows were written before it ended.
Steps 5, 7 and 8 are the three that separate the postures, and in each the response status and the database state are recorded as two independent observations. That separation is the whole method, because on this axis they routinely disagree, and every posture above level 0 is defined by what the table contains rather than by what the response said.
The named trap: the round-trip illusion
Every axis we publish names the illusion that keeps its defect invisible, because the defect is rarely hidden by difficulty. It is hidden by a test that cannot fail.
Here the trap is that the only file the importer has ever been given is one the exporter wrote.
It is worth being slow about why that is so hard to escape. You build the export first, because it is easier and because somebody asked for it. Then you build the import, and you need a file to test with, and there is exactly one file in existence: the one your own application just produced. It is UTF-8, because your runtime emits UTF-8. Its columns are in your order, because your code wrote the header. Its dates are in your format. Every row is valid, because every row came out of your own validated table. There are forty of them, because that is what you seeded.
The file cannot fail. Not because you were careless, but because the producer and the consumer are two halves of one artefact, and every assumption the importer makes was satisfied by construction before it was ever written down. The test passes, and it was never capable of doing anything else.
This is the first of our named traps where the illusion is actively strengthened by a correct remedy on a neighbouring axis. Our export work rewards a project for round-tripping its own export through a conformant reader, and that is good advice which makes a project's exports genuinely better. It also means the better a project scores there, the more confident it is in exactly the one input that proves nothing here.
The file a real person uploads was written by a different program, on a different operating system, by somebody who exported it out of the tool they are leaving. It has ten thousand rows. It has a column your app does not have and is missing one it needs. Row 7,412 has a date written the other way round. And the first time anybody discovers what your importer does with a file like that, it is doing it to production data belonging to a customer who is, at that exact moment, in the middle of migrating onto your product.
The trap joins a family we have been naming across this series, in which every observation available at test time is generated by the same actor. This one has a property the others do not: the test input is manufactured by the system under test, so the instrument and the subject are two halves of the same artefact. There is no adversary to model and no timing to get unlucky with. The input simply cannot exercise the failure, and no amount of running it more times changes that.
Two documented tensions worth stating
Neither is a contradiction. Both are places where two pieces of good advice pull against each other, and in both the resolution is scope rather than choice.
The byte order mark overrides the exporter's declaration. Our export axis weights encoding declared rather than assumed at 16 points, and the remedy is to declare the character set on the media type. On the receiving side that declaration can be overruled by three bytes at the front of the file. The Encoding Standard says so, and says what it costs: "For compatibility with deployed content, the byte order mark is more authoritative than anything else. In a context where HTTP is used this is in violation of the semantics of the Content-Type header."
So an exporter can do everything our own rubric asks and still have its declaration ignored by a conformant reader, because the file itself outranks the header. The same specification also notes that standards are strongly discouraged from using these legacy hooks except as needed for compatibility, which is a fair description of the entire situation. We resolve it by scope: the export axis scores what you declare, this axis scores what you determine on the way in, and an importer that trusts a Content-Type charset without checking for a mark has trusted the less authoritative of the two.
All-or-nothing fights resumability. Signal 1 rewards a build for making partial application a decision, and the cleanest way to do that is a single transaction: nothing is applied unless everything can be. But a single transaction over a very large file holds one connection and one open transaction for the duration, which is precisely the pressure our database-connection work describes, and it makes the import unresumable by construction, since there is no partial state to resume from. The opposite design, batched commits with a recorded position, is friendlier to large files and gives up the property that made it safe.
We do not resolve this by preferring one. We resolve it by classification, which is why level 2 admits both: all-or-nothing and resumable-with-a-recorded-position are two correct answers to the same question, and the failing posture is the one that is neither.
What this axis is not
- It is not a performance measure. Whether a million-row import finishes is a capacity question that belongs to our database-connection and background-work axes. Signal 6 is about whether the ceiling is handled, not about where it sits.
- It is not a judgement on any builder. No builder is scored, named or assigned a posture on this page, and none will be until the rubric has been through a community-edit window. The cohort is stated so that the eventual scope is predictable, not because anything has been measured.
- It is CSV-shaped as drafted, and should not be. Signals 4 and 5 assume a delimited text file with a header row. A JSON import has no delimiter and no column order, and a spreadsheet-native format carries types rather than guessing at them, which makes signal 4 close to free and signal 5 meaningless. That is a limitation of this draft rather than of the axis, and it is the first thing we would want a contributor to argue with.
- Signal 3 can punish the better design, and the rubric must say so. A build with true all-or-nothing semantics can never have a partial application to duplicate, so re-submission is safe for free. Scoring it as having no idempotency mechanism would penalise the design that made one unnecessary. Where signal 1 is satisfied by atomicity, signal 3 scores a full pass rather than a blank. We are stating this in the rubric rather than leaving it to a scorer's judgement, because a scorer working row by row would get it wrong.
- Signal 7 may be worth less than seven points, or nothing. The same argument applies one step further: if signals 1 and 2 are both satisfied, a failed import is cheap, reversible and fully diagnosed, and a rehearsal adds very little. There is a coherent version of this rubric with six signals in which the dry run is folded into signal 2 as a way of satisfying it. We have not taken it because a preview reports before writing and a diagnosis reports after, and on a migration those are different products. But the argument against us is good.
- A build with no import feature at all scores nothing here, and that is a finding rather than a null. The reference prompt asks for the capability. A builder that emits none has answered the prompt, and the answer is recorded as such rather than as an absence of evidence.
- Single-generation variance. One export is one sample. A builder that is inconsistent between generations needs a multi-run design before scoring is fair.
Comments, counter-rubrics and reproduction attempts are welcome. The single most useful thing you can send us is a step 5 and step 7 result from your own generated application: how many rows were in the table after the failure, and how many were in it after you uploaded the corrected file.
References
- PostgreSQL 17,
COPY, the default failure behaviour, theON_ERRORclause and itsstopandignorevalues, the type-conversion scope, and theLOG_VERBOSITYsplit between an ignored-row count and a per-row line and column, https://www.postgresql.org/docs/17/sql-copy.html (accessed September 30, 2026) - PostgreSQL 18,
COPY, theREJECT_LIMITclause, the unlimited default tolerance underON_ERROR = ignore, and thesilentverbosity value, https://www.postgresql.org/docs/18/sql-copy.html (accessed September 30, 2026) - PostgreSQL, Tutorial 3.4, Transactions, the all-or-nothing definition and the guarantee that a failed transaction leaves no step affecting the database, https://www.postgresql.org/docs/current/tutorial-transactions.html (accessed September 30, 2026)
- WHATWG, Encoding Standard, the problem statement on legacy encodings, the
decodealgorithm's byte-order-mark precedence and its note onContent-Typesemantics, and the legacy-hooks guidance, https://encoding.spec.whatwg.org/ (accessed September 30, 2026) - MDN,
TextDecoder(), theutf-8label default, thefataldefault offalseand its replacement-character behaviour, and theignoreBOMdefault, https://developer.mozilla.org/en-US/docs/Web/API/TextDecoder/TextDecoder (accessed September 30, 2026)
Written by
BuilderProof editorial teamCite this benchmark
BuilderProof editorial team. "What happens to the rows that were fine? A bulk-import correctness axis proposal (September 2026)". BuilderProof, September 2026. https://www.builderproof.org/benchmarks/what-happens-to-the-rows-that-were-fine-bulk-import-axis-september-2026.
@misc{builderproof-what-happens-to-the-rows-that-were-fine-bulk-import-axis-september-2026,
title = {{What happens to the rows that were fine? A bulk-import correctness axis proposal (September 2026)}},
author = {{BuilderProof editorial team}},
year = {2026},
month = {sep},
howpublished = {\url{https://www.builderproof.org/benchmarks/what-happens-to-the-rows-that-were-fine-bulk-import-axis-september-2026}},
note = {BuilderProof, builderproof.org}
}Frequently asked questions
What is bulk-import correctness?
It is a proposed BuilderProof benchmark axis, drafted 30 September 2026, that scores what an AI app builder's generated import path does when a multi-record file fails partway through: whether partial application is a deliberate decision or an accident, whether the failing row is identified rather than merely counted, whether re-submitting a corrected file is safe, and whether the encoding is determined rather than assumed. It weights seven signals and maps them to four postures. No builder is scored on the proposal page.
Why is this not covered by the input-validation axis?
Because the unit changed, and the rubric inverts. The input-validation axis scores a 400 or 422 as a pass and a 200 that persists rows as a fail, which is correct for one request carrying one record. On a ten-thousand-row file, rejecting the request discards the 9,999 rows that were fine, and applying what you can while reporting what you cannot is close to the right answer. On a bulk-import path that neighbouring rubric scores the correct behaviour as a failure. Neither axis needs amending; they need scoping.
What does the database itself do when an imported row is bad?
PostgreSQL's COPY documentation states that by default COPY will fail if it encounters an error during processing, and that ON_ERROR can be set to ignore to discard the offending row and continue. Two things are opt-in beyond that. If no REJECT_LIMIT is given, ON_ERROR = ignore allows an unlimited number of errors, so an entirely bad file loads nothing without raising. And the line number and column name of each discarded row are emitted only when LOG_VERBOSITY is set to verbose; by default you get a count and not an identity.
Can an import fail without anything reporting an error?
Yes, and encoding is the common case. MDN documents that a TextDecoder's label defaults to utf-8 and that its fatal option defaults to false, which means the decoder will substitute malformed data with a replacement character. A file written by a spreadsheet in a legacy single-byte encoding therefore decodes without raising, every accented character becomes a replacement character, and the import reports complete success. There is no row to skip and no count to report.
What is the round-trip illusion?
It is the named trap of this axis: the only file an importer is given during development is one the same application's export produced, so it shares every assumption the importer makes and cannot fail. The producer and the consumer are two halves of one artefact. It is reinforced by good advice, because our data-export axis rewards a project for round-tripping its own export through a conformant reader, which makes a project most confident in precisely the input that proves nothing about importing a stranger's file.
Related benchmarks
Input-Validation and Data-Integrity Posture: A Proposed Axis for Whether AI App Builders Validate Untrusted Input at the Boundary (August 2026)
A candidate BuilderProof benchmark axis that scores whether the code AI app builders emit validates untrusted input at the server boundary, or trusts whatever the client sends. Rubric, four posture levels, an adversarial-payload reproduction protocol, and an open call for comment.
File-Upload and Media-Handling Posture: A Proposed Axis for What an AI Builder's Upload Path Actually Permits (August 2026)
A proposed BuilderProof benchmark axis measuring whether the file-upload path an AI app builder emits is owner-scoped, content-verified, bounded and durable, or whether it only works because the developer testing it is the only account in the app. Six weighted signals, four posture levels, a reproducible protocol.
Does the Export Open Anywhere Else? A Proposed Axis for Data-Export Correctness (September 2026)
Every export is checked by the one reader whose settings match the writer's. A pre-registered seven-signal axis for whether a generated application's export survives arriving anywhere else.