33,415 Apex tests from 76 open-source projects.
98.2% produce identical results locally.
Nimbus is measured against real Salesforce codebases, cloned unmodified, pinned to a commit, and run with the same tests and assertions their authors wrote. This page is the record of how that goes, including the part that does not flatter us: every test that still fails, and how many of those failures were re-run on a real org and confirmed as ours.
@isTest method in those projects, none excludedScorecard run Thu Oct 8 20:05:43 CEST 2026 · sha256 4119e2ea662bb6da
What identical means
Here is the whole method behind the number.
The projects are unmodified clones.
Each row is a public repository checked out at a specific commit, with no source changes. No test is skipped, rewritten, or annotated away, and no project is dropped from the corpus because it scores badly. The worst-scoring rows are in the table below.
Nimbus runs them with no org attached.
The whole suite runs from source on one machine: the Apex interpreter, SOQL translated to PostgreSQL, an embedded database carrying the Salesforce data model, transaction and savepoint semantics, triggers, and record-triggered flows in their order-of-execution slots. A test passes only when its own assertions pass. Nothing is stubbed to make a number look better.
Every failure is re-run on a real org.
A test failing locally proves nothing on its own: the test may be broken, or the org may be configured in a way the project assumes. So each failure is deployed source-identical to a real Salesforce org and run there. The deploy is the identity proof: it is atomic, so if the source had not compiled identically, nothing would have deployed.
Three verdicts, one of which counts against us.
The org passes and Nimbus fails: a confirmed divergence, counted against us and filed as a bug. Both fail: a shared failure (the test is order-dependent, or asserts something only true inside a packaged namespace). The org cannot run it at all: environmental. The last two are reported here and stay in the denominator.
Every project, every number
Largest suite first. The last column is how many of that project's current failures have been individually re-run on a Salesforce org and confirmed as a Nimbus bug.
| Project | Tests | Passing | Failing | Pass rate | Org-confirmed |
|---|---|---|---|---|---|
| NPSP19646faa | 3,985 | 3,979 | 6 | 99.8% | 6 |
| nppatchfc34ed1f | 3,674 | 3,659 | 15 | 99.6% | not classified |
| EDAe5939577 | 2,357 | 2,355 | 2 | 99.9% | 0 |
| Portwood62183cae | 2,097 | 2,094 | 3 | 99.9% | not classified |
| nebula-logger05ce16c5 | 1,354 | 1,354 | — | 100.0% | — |
| aiAgentStudio57533303 | 1,090 | 1,075 | 15 | 98.6% | not classified |
| ApexEloquentb9109567 | 1,010 | 1,010 | — | 100.0% | — |
| fflib-apex-extensions368978bf | 892 | 892 | — | 100.0% | — |
| FormulaShare-DXdc69c49d | 876 | 876 | — | 100.0% | — |
| dlrs5dbd186b | 830 | 830 | — | 100.0% | — |
| forceeae7713d24 | 830 | 827 | 3 | 99.6% | 3 |
| rflibac3032b0 | 808 | 799 | 9 | 98.9% | 9 |
| at4dx8c109e4f | 746 | 746 | — | 100.0% | — |
| fflib86bd8791 | 746 | 746 | — | 100.0% | — |
| apex-rollup6b8ae4ca | 744 | 742 | 2 | 99.7% | 0 |
| CVMA20-79ffc7e64 | 740 | 224 | 516 | 30.3% | not classified |
| apex-stream066bc579 | 735 | 735 | — | 100.0% | — |
| expressionacbd1374 | 710 | 710 | — | 100.0% | — |
| Salesforce-Moxygen51d085e4 | 683 | 683 | — | 100.0% | — |
| fflib-apex-mocks24deeebc | 471 | 471 | — | 100.0% | — |
| soql-lib0fdaeb9d | 466 | 466 | — | 100.0% | — |
| PMM4c5f3c79 | 433 | 429 | 4 | 99.1% | 3 |
| apex-mocks-sfdx81ce160b | 428 | 428 | — | 100.0% | — |
| amoss3c054ed3 | 424 | 424 | — | 100.0% | — |
| dml-lib9c02feb3 | 372 | 372 | — | 100.0% | — |
| MoH-SAT1cb74490 | 371 | 371 | — | 100.0% | — |
| forcedotcom-enterprise-architecture-second-edition655fabfc | 363 | 361 | 2 | 99.4% | not classified |
| Apex-Opensource-Library57d09684 | 347 | 347 | — | 100.0% | — |
| ApexKitc2660a33 | 334 | 334 | — | 100.0% | — |
| apex-recipesd2ec0b33 | 322 | 322 | — | 100.0% | — |
| salt-yastf457731fb | 303 | 303 | — | 100.0% | — |
| nebula-core0a3a20cc | 294 | 286 | 8 | 97.3% | 0 |
| sf-bedrock70847609 | 230 | 230 | — | 100.0% | — |
| salesforce-isv-cockpitd38f5f1d | 185 | 170 | 15 | 91.9% | not classified |
| apex-fp61769f99 | 179 | 179 | — | 100.0% | — |
| FluentQueryb5dc8483 | 175 | 175 | — | 100.0% | — |
| Volunteers-for-Salesforce40c3d311 | 165 | 159 | 6 | 96.4% | 4 |
| Apex-GraphQL-Clientb29350eb | 162 | 162 | — | 100.0% | — |
| trigger-actions-framework5c3793d2 | 155 | 155 | — | 100.0% | — |
| ActionPlansV4ef5a417a | 152 | 151 | 1 | 99.3% | not classified |
| ada-wallet-for-salesforce81cf66f2 | 145 | 145 | — | 100.0% | — |
| apex-mockery320d3a73 | 133 | 133 | — | 100.0% | — |
| sfdc-xml-parser377631b7 | 131 | 131 | — | 100.0% | — |
| apex-utilities3d9b8182 | 129 | 128 | 1 | 99.2% | not classified |
| box-salesforce-sdk9d41fdf0 | 124 | 124 | — | 100.0% | — |
| sobject-fabricatorbdf66b2e | 124 | 124 | — | 100.0% | — |
| automation-componentsa43d9416 | 122 | 122 | — | 100.0% | — |
| apex-evalex30a3f011 | 105 | 103 | 2 | 98.1% | not classified |
| apex-dml-mocking5daf499b | 95 | 95 | — | 100.0% | — |
| force-dot-com-esapib6eddab9 | 89 | 89 | — | 100.0% | — |
| my-org-butler8a0994f3 | 88 | 88 | — | 100.0% | — |
| apex-test-kit607e6577 | 87 | 87 | — | 100.0% | — |
| TestDataFactory22487c80 | 83 | 82 | 1 | 98.8% | 1 |
| async-lib5486e899 | 82 | 82 | — | 100.0% | — |
| org-error-inbox3707fa37 | 75 | 74 | 1 | 98.7% | not classified |
| force-di574d0509 | 61 | 61 | — | 100.0% | — |
| http-mock-lib42bdd066 | 50 | 50 | — | 100.0% | — |
| lightweight-soap-utild1321cee | 44 | 44 | — | 100.0% | — |
| lightweight-xlsx-util6b0c2e1a | 42 | 42 | — | 100.0% | — |
| apex-xpath11d5c22c | 40 | 40 | — | 100.0% | — |
| jsonparse69da320c | 39 | 39 | — | 100.0% | — |
| test-lib82ebbc9c | 36 | 36 | — | 100.0% | — |
| NebulaQueryAndSearch269a787b | 34 | 34 | — | 100.0% | — |
| apex-domainbuilder26ca2d52 | 34 | 34 | — | 100.0% | — |
| approvalsOnFlow6579d8f9 | 32 | 28 | 4 | 87.5% | not classified |
| apex-validatee0bf0c35 | 29 | 29 | — | 100.0% | — |
| ecarsa5e9cc01 | 23 | 23 | — | 100.0% | — |
| OutboundFunds8acc96b4 | 21 | 21 | — | 100.0% | — |
| cache-manager2e5e55ab | 18 | 18 | — | 100.0% | — |
| ApexValidationRules6a2aa3e9 | 13 | 13 | — | 100.0% | — |
| sfdc-trigger-frameworkb7e36c76 | 13 | 13 | — | 100.0% | — |
| dreamhouse-lwc9810695a | 11 | 11 | — | 100.0% | — |
| streaming-monitor0d35b300 | 11 | 11 | — | 100.0% | — |
| ItemsToApprovea7292e81 | 6 | 6 | — | 100.0% | — |
| apex-constsff260c18 | 4 | 4 | — | 100.0% | — |
| ebikes-lwcfd29e9fe | 4 | 4 | — | 100.0% | — |
| All projects | 33,415 | 32,799 | 616 | 98.2% | 26 |
Failing is measured locally: how many of that project's tests Nimbus does not currently reproduce. Org-confirmed divergences is a stricter number: it counts only failures that were deployed to a real Salesforce org, run there, observed to pass, and recorded by name. A test drops out of that column the moment it stops failing, so the column shrinks when we fix something rather than staying at whatever a past verification run measured.
Where the two differ, the gap is not a hidden failure. It is either a failure the org reproduced too, or one the org could not run, or one that started failing after the last verification pass. Those are all counted in the failing column and none of them are counted as confirmed.
Rows measured against a private repository are left off the published scorecard; they contribute no tests to the totals above.
How much of this is verified
Each of these failures was executed on Salesforce, by name, and its verdict written down.
The failing column is the backlog.
A confirmed divergence is a work item with a reproduction, a real-org expectation and a name. That is the reason to run a corpus this size against an oracle: it produces a queue of the bugs that matter, ranked by how much real code trips over them. 26 of them are open right now, tracked as issues rather than filtered out of the scorecard.
576 of the failures on this page carry no classification yet: they began failing after the most recent verification pass. They are counted against us in every figure above until an org says otherwise.
Verification passes
- 2026-09-18T00:43:59+02:000 failures classified against the scorecard as it stood then, 0 confirmed as divergences.
- 2026-09-1860 failures classified against the scorecard as it stood then, 60 confirmed as divergences.
- 2026-09-17T19:44:11+02:000 failures classified against the scorecard as it stood then, 0 confirmed as divergences.
- 2026-09-1740 failures classified against the scorecard as it stood then, 38 confirmed as divergences (1 of the projects it covered no longer fail at all, so nothing from them appears in the table).
- 2026-09-170 failures classified against the scorecard as it stood then, 0 confirmed as divergences.
- 2026-08-25638 failures classified against the scorecard as it stood then, 628 confirmed as divergences (8 of the projects it covered no longer fail at all, so nothing from them appears in the table).
- 2026-08-19993 failures classified against the scorecard as it stood then, 969 confirmed as divergences (7 of the projects it covered no longer fail at all, so nothing from them appears in the table).
- 0 failures classified against the scorecard as it stood then, 0 confirmed as divergences.
- 400 failures classified against the scorecard as it stood then, 44 confirmed as divergences (1 of the projects it covered no longer fail at all, so nothing from them appears in the table).
Reproduce it
Every row above is a public repository at a published commit, and the command is the same one we ran.
brew install nimbus-solution/nimbus/nimbus
# or: curl -fsSL https://testnimbus.dev/install.sh | shnebula-logger: 1,354 tests, all of them reproduced locally.
git clone https://github.com/jongpie/NebulaLogger
cd NebulaLogger
git checkout 05ce16c5
nimbus test "*"forceea: 830 tests, 3 of them still failing. Check that number too; it is the one we would most like to be wrong about.
git clone https://github.com/nmitrakis/Forceea
cd Forceea
git checkout e7713d24
nimbus test "*"The manifest lists every project on this page with its repository URL, the exact commit it was measured at, and the counts to expect.
Expect a couple of tests of movement on the largest suites: a handful of record-merge and bulk-DML tests are sensitive to how much of the machine they get, so a run can land a test or two either side of the number here. That is why the manifest carries per-project counts and commits rather than one headline percentage: a difference can be localised to a project and then to a test name, which is how we track it internally too.
A local runtime approximates the platform. This is the measurement of how closely.
Nimbus is not Salesforce and does not claim to be. It is an implementation of the parts of the platform that Apex tests exercise, and like every implementation it is right until it is found to be wrong. The corpus exists to find that out on real code, continuously, before a customer does, and the results are published whether they moved in our favour or not.
The org still owns:
- Sharing rules (org-wide defaults, the role hierarchy, groups and share rows are enforced)
- Flow-based approvals, time-dependent workflow actions, scheduled process actions and outbound messages
- Lightning UI / browser testing
- Final pre-deployment validation (your org is still the source of truth)
Failures are tracked as public issues. See the issue tracker.
Why publish this
Fidelity is easy to claim and hard to prove. "Runs your Apex exactly like the platform" is a sentence anyone can type, and we have never found a local Apex runtime (ours included, until now) that published a record you could check it against. So here is ours: named repositories, pinned commits, failing counts, and a verdict from a real org on each failure.
The point of the format is that it can be falsified. We would rather publish 616 failures with their receipts than a round number with none. If a row does not reproduce on your machine, that is a bug report we want.
nimbus test "*"