How would you design for a customer with 500 million records on a single object?
Suggested answer
At that scale I stop asking how to tune the object and start asking whether it should exist in that form:
1. Challenge the shape: Is this one object because the data is genuinely one entity, or because nobody separated hot transactional data from cold history? Usually it is the latter, and splitting it is the highest-value change available.
2. Tier the storage: Recent, actively used records in the custom object; historical records in a Big Object or an external store; aggregates in a small summary object that carries the reporting load.
3. Design access patterns first: At this volume you design the queries before the schema. Every access path needs a selective, indexed filter — typically a date range plus an owner or an External Id. Ad hoc querying is not a supported use case.
4. Control skew rigorously: No parent or owner concentration anywhere near the thresholds, enforced by design and monitored, not left to chance.
5. Plan operations: Deferred sharing calculation for loads, skinny tables for the hot reports, PK chunking for extracts, and a documented archive job that runs continuously rather than as an annual event.
6. Set expectations: I would be explicit with stakeholders that some conveniences — unfiltered list views, arbitrary cross-object reporting, casual full exports — are not available at this scale. Agreeing that up front prevents a lot of disappointment later.
Practice content for interview preparation; not an official vendor answer. Verify details against current product documentation.
Community comments (0)
No comments yet.
Sign in or create a free account to add a comment. Comments are moderated before they appear.