Run explorer

Filter the corpus, tick two runs, get a diff. Comparisons the evidence cannot support are refused rather than shown with a caveat.

Run SamplesTPS FPSWorstErrTests

Crashed runs are kept and labelled, never dropped. The run that ended badly is usually the interesting one.