Three years ago we started publishing a model card with every model, including a section on where the model does not work. It is worth reporting what that has cost.
It has cost us at least two competitive evaluations that I know of, where a procurement team compared our stated recall under dense canopy against a competitor's unqualified headline figure and drew the obvious conclusion. That is a real cost and I am not going to pretend it is not.
It has returned three things. First, it has almost entirely eliminated a category of support case where a customer discovers a limitation in production and feels misled. Second, it has made our own engineering better, because a failure mode you have to write down publicly is a failure mode someone will try to fix. Third, it has been decisive in at least four enterprise deals where a technically sophisticated buyer recognised what the document meant.
The net is positive but it is not overwhelmingly positive, and anyone adopting this practice because it sounds virtuous should know the trade is real.
What has changed in the practice: we now publish the evaluation dataset composition, after a customer pointed out — correctly — that a recall figure is uninterpretable without knowing what it was measured on.