Question on training time with Resnet18

I am intrigued by your work that demonstrates the effectiveness of ZO optimization in training large-scale models. Your main experiments on CIFAR-10 using ResNet20 show that it takes approximately 60 minutes per epoch (a result I have successfully replicated).
![image](https://github.com/user-attachments/assets/ba18e20b-47eb-42a3-9f1b-ff67e37370f4)
However, your framework, Deepzero, utilizes CGE, which causes the inference time to increase linearly with the model size. In Appendix D, Table A3, you reported training ResNet18, whose model size is approximately ten times larger than ResNet20. I am curious about how long it took to train ResNet18 using the Deepzero framework. 



Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Question on training time with Resnet18 #4

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Question on training time with Resnet18 #4

Description

Metadata

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Issue actions