Between July 11 and July 29 I created 100 GPU instances on a rental marketplace. 62 never delivered a working GPU. Total...

Between July 11 and July 29 I created 100 GPU instances on a rental marketplace. 62 never delivered a working GPU. Total damage for all 62 failures: $4.37.The worst host booted fine, passed nvidia-smi, showed healthy VRAM, and could not run a single tensor operation. It idle-burned 105 minutes before a human noticed.A health check that asks a component to describe itself is not a health check.https://roamingpigs.com/r/ml6t#machinelearning #gpu #selfhosting

Read Original

Related