A client's nightly integration is failing intermittently with UNABLE_TO_LOCK_ROW. How would you troubleshoot it?
Suggested answer
I work from evidence rather than guessing, roughly in this order:
1. Confirm the pattern: Pull the failed batches and look for what the failing records have in common. Almost always they share a parent record, an owner, or a common lookup target.
2. Check for skew: Query the child object grouped by the parent Id and by OwnerId to see whether any single value is far above the rest. That usually identifies the contention point in minutes.
3. Look at concurrency: Is the job running Bulk API in parallel mode? Parallel batches touching the same parent will contend. Serial mode is slower but often the fastest path to a job that actually completes.
4. Sort the input: Ordering the file by parent Id so all children of a parent land in the same batch removes most cross-batch contention without giving up parallelism entirely.
5. Look for automation: A trigger or flow that updates a shared parent or a shared summary record on every child insert will serialise the whole load. That is a design fix, not a tuning fix.
6. Then tune: Only after the above would I adjust batch size, defer sharing calculation for the load window, or split the job. Reducing batch size first is a common instinct and often just makes a slow job slower without fixing the cause.
Practice content for interview preparation; not an official vendor answer. Verify details against current product documentation.
Community comments (0)
No comments yet.
Sign in or create a free account to add a comment. Comments are moderated before they appear.