Skip to main content
These models output a scalar reward score or classification result, often used in reinforcement learning or content moderation tasks. They are executed with --is-embedding and some may require --trust-remote-code.

Example launch Command

Supported models