Separate tracks were the speaker attribution. The lock is what takes them away — not by failing, but by quietly leaving one stream unrecognised.
I needed two concurrent speech recognition streams. Apple's docs don't mention you can't have them — the second start silently kills the first.