Trust, Risk & Safety advanced

Interpretability

Understanding how a model works internally, as opposed to just describing its outputs.

L’interpretabilità è la domanda meccanicistica: che cosa stanno effettivamente calcolando questi pesi e queste attivazioni? È più difficile della spiegabilità e più preziosa, perché può rivelare i modi di fallire prima che compaiano negli output. I progressi sono reali ma molto indietro rispetto alle capacità.

In pratica: Individuare le feature interne che un modello usa per rappresentare un concetto.