optimus/trainer.py · 4ab91bcf57f01d37ee29bbd98459a04129df73aa · NetSys / Optimus Prime

1 year ago

Ignore last batches when calculating final train loss · 4ab91bcf

Alexandru-Mihai GHERGHESCU authored 1 year ago

Visual change. This only changes what the trainer reports as the final
training loss.

Not quite sure if the value before was accurate anyway, since gradient
accumulation would not let the optimizer step every batch anyway.

For a big enough dataset, this should not have any impact at all.

The final loss value will be reported based on the last calculation of
the loss, correctly taking into consideration gradient accumulation as
well.

Unverified

4ab91bcf

History

Ignore last batches when calculating final train loss

Alexandru-Mihai GHERGHESCU authored 1 year ago

Visual change. This only changes what the trainer reports as the final
training loss.

Not quite sure if the value before was accurate anyway, since gradient
accumulation would not let the optimizer step every batch anyway.

For a big enough dataset, this should not have any impact at all.

The final loss value will be reported based on the last calculation of
the loss, correctly taking into consideration gradient accumulation as
well.

Code owners

Assign users and groups as approvers for specific file changes. Learn more.