Property-Based Testing · Lesson 02
Relate Two Calls
Lesson 01 predicted every response from a model. Sometimes you can't predict the answer, because the data isn't yours or the logic is too complicated to model. You can still say how two answers must relate. That is metamorphic testing, and list endpoints are full of it.
The oracle problem
Every test needs an oracle: something that says whether the output is right. An example test hard-codes it. A model-based test computes it from the model. Both stop working when the right answer is expensive to know:
- Search with relevance ranking, or a query served by Elasticsearch/Algolia instead of the database.
- Permission-filtered views, where “what can this user see” is the complicated logic itself.
- A staging environment with real, messy data that you don't control and can't reset.
Metamorphic testing avoids the problem. You don't ask “is this output right?” You ask “are these two outputs consistent with each other?”
Terms
- Source test case
- The first call, with input x. Its output S is not checked by itself.
- Follow-up test case
- A second call, with input x′ derived from x in a known way, e.g. by adding a filter. Output F.
- Metamorphic relation (MR)
- The rule tying them together: how x′ is built from x, and what must hold between S and F.
Hughes's summary is “related calls return related results.” Wlaschin's “different paths, same destination” is one kind of metamorphic relation: the two paths are the two calls, and the relation is equality.
Six relations that fit almost every list endpoint
Segura et al. (IEEE TSE, 2018) observed that REST APIs use filtering, ordering and pagination in very similar ways, so the same relations apply across APIs. They named six metamorphic relation output patterns (MROPs), each a set relation between S and the follow-ups Fi:
| MROP | Relation | In your GET /projects |
|---|---|---|
| Equivalence | Same items, any order | ?sort=name and ?sort=createdAt return the same set |
| Equality | Same items, same order | Omitting limit = passing the default; all pages at size 20 concatenated = all pages at size 50 |
| Subset | F ⊆ S | Adding &ownerId=u1 to ?status=active returns a subset |
| Disjoint | S ∩ F = ∅ | ?status=active and ?status=archived share no rows |
| Complete | S = F1 ∪ … ∪ Fn, and |S| = Σ|Fi| | The union over every status value = the unfiltered list, and the counts add up |
| Difference | F differs from S in exactly a known set of fields | PATCH {name} returns the previous GET body with only name and updatedAt changed |
Two details from the paper are easy to miss. Complete requires disjoint, but not the other way round: partitions can be disjoint and still not cover everything. And the count check |S| = Σ|Fi| is what catches duplicates, because a set union hides them.
What it found
Segura et al. identified 60 relations in the Spotify and YouTube APIs and found 11 issues. Ten were confirmed by the developers or reproduced by other users. Every one of them is a list-endpoint bug of the kind a CRUD app can have:
- Equality, pagination. A Spotify album search returned 21 results when paged by 20 and 27 when paged by 30. Spotify fixed it within two weeks; it was a regression from another fix.
- Subset, filter. Adding
market=ESto a Spotify search increased the results from 35 to 45. A filter that adds rows. - Equivalence, sorting. A YouTube search returned 46 items; the same search with
order=datereturned 4. - Complete, query operators. “mended” returned 27 albums; “mended AND with” returned 1 and “mended NOT with” returned 25. 1 + 25 ≠ 27.
None of these needed anyone to know the correct result of a search.
As a property
The relation is fixed and the inputs are generated. The generator produces a random filter, and the test derives the follow-up from it:
const filter = fc.record({
status: fc.option(fc.constantFrom('active', 'archived', 'draft'), { nil: undefined }),
ownerId: fc.option(fc.constantFrom(...seededUserIds), { nil: undefined }),
q: fc.option(fc.constantFrom('roadmap', 'q3', 'a'), { nil: undefined }),
}, { requiredKeys: [] });
// Follows nextCursor until the end. Never compare single pages.
const all = async (f: Filter, pageSize: number) => { /* … */ };
const ids = (rows: Project[]) => rows.map((r) => r.id);
test('page size does not change the result (equality)', () =>
fc.assert(fc.asyncProperty(filter, fc.constantFrom(1, 7, 50), async (f, size) => {
assert.deepEqual(ids(await all(f, size)), ids(await all(f, 20)));
})));
test('adding an owner filter gives a subset', () =>
fc.assert(fc.asyncProperty(filter, fc.constantFrom(...seededUserIds), async (f, owner) => {
const s = new Set(ids(await all(f, 50)));
for (const id of ids(await all({ ...f, ownerId: owner }, 50))) assert.ok(s.has(id));
})));
test('statuses partition the list (complete)', () =>
fc.assert(fc.asyncProperty(filter, async (f) => {
const source = ids(await all({ ...f, status: undefined }, 50));
const parts = await Promise.all(STATUSES.map((st) => all({ ...f, status: st }, 50)));
const union = parts.flatMap(ids);
assert.equal(union.length, source.length); // catches duplicates
assert.deepEqual(new Set(union), new Set(source));
})));
These tests don't reset the database and don't need a model. They run against whatever data is there, which is why Segura could run them against live Spotify. There are three conditions:
- No writes between source and follow-up. If someone creates a project between the two calls, the relation fails without a bug. Against a shared environment, use a read-only dataset or retry once before reporting.
- Equality needs a total order. If the sort has ties and no tie-breaker, the order between tied rows isn't defined, and offset pagination can drop or repeat them. Here, a failing equality relation is reporting a real bug, not a flaky test.
- Writing the relation forces a decision on the spec. Hughes's first metamorphic property, “insert
order doesn't matter”, failed immediately: inserting the same key twice with different values does depend on
order. The property wasn't wrong to fail. The spec had never said which value wins. Expect the same from your
relations, e.g. “does
?q=match archived projects?”
Model or metamorphic?
The two are not exclusive. Lesson 01's model could also predict GET /projects?status=active,
because filtering a Map is easy. A rough rule:
- Model when you control the state and the intended behaviour is simple enough to write down: writes, lifecycle, uniqueness, soft-delete.
- Metamorphic when the answer is expensive to predict or the data isn't yours: search, ranking, permission views, reporting endpoints, anything served from a second store. Also when you want to test in staging against realistic data.
Exercise 1 · Interleaved recognition
Name the relation
Type the MROP (equivalence, equality, subset, disjoint, complete, difference) or, for the items from lesson 01, the pattern (model, invariant, inverse, idempotence, commutativity). Press Enter to check.
Exercise 2 · Bug hunt
Find the defect in the test
Click the defective line, or press no defect.
Exercise 3 · Execution
Write the relations
An endpoint you don't have a model for:
GET /reports/revenue?customerId=…&from=2026-01-01&to=2026-04-01¤cy=EUR
→ { total: 18240.50, invoiceCount: 37, invoices: [ … ] }
Write two metamorphic relations. For each one, give the follow-up input(s) and the relation between outputs.
Complete, by splitting the date range. Choose any mid between from and to.
[from, mid) and [mid, to) must together contain exactly the invoices of
[from, to), and invoiceCount and total must add up. This finds bugs where
the boundaries are inclusive on both ends (an invoice dated mid is counted twice) or exclusive on both ends
(it is lost), and timezone bugs at midnight.
Subset, by narrowing the range. A range inside [from, to) returns a subset of invoices and a
total no larger, provided totals can't be negative. Credit notes break the “no larger” part. If you
weren't sure whether they exist, that's the spec question this relation makes you ask.
Also valid: equality between omitting currency and passing the default; across
customerId values, results are disjoint, and complete if every invoice has a customer. A
currency relation (EUR total × rate = USD total) is only valid if conversion happens per report and not per
invoice. That is the kind of thing worth finding out.
Exercise 4 · Recall
From memory
Write before revealing. Question 3 is from lesson 01.
To keep it
Pick the list endpoint in your codebase with the most query parameters. Without looking at the lesson, write one relation per parameter and name its MROP. Parameters you can't write a relation for are often ones whose behaviour nobody has specified. In a week, do the same for a second endpoint, then compare with the table above.
Sources: Segura et al., Metamorphic Testing of RESTful Web APIs (TSE 2018) · Wayne: Metamorphic Testing · Hughes: How to Specify It! §4.3 · all resources
← 01 · Model the Resource · Status · Next: 03 · Five Ways to Specify It →