Skip to content

More reward models: Train a DPO model and autoencoders for hh-hrlf #10

Description

@amirabdullah19852020

For now, do this just for gpt-125m-neo to make this fast.

Make the pipeline configurable while doing this, so we can easily run this for a range of other models quicker.

Metadata

Metadata

Labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions