Documentation

Documentation

See the documentation for a technical overview of the platform and train your first agent

Quick Start

1. Install uv (Python package manager)

# macOS/Linux:
$ curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows:
PS> powershell -c "irm https://astral.sh/uv/install.ps1 | iex"

2. Install ReinforceNow

uv init && uv venv --python 3.11
source .venv/bin/activate  # Windows: .\.venv\Scripts\Activate.ps1
uv pip install rnow

3. Authenticate

rnow login

4. Create & Run Your First Project

rnow init --template sft
rnow run

That's it! Your training run will start on ReinforceNow's infrastructure. Monitor progress in the dashboard.

Core Concepts

Go from raw data to a reliable AI agent in production. ReinforceNow gives you the flexibility to define:

1. Reward Functions

Define how your model should be evaluated using the @reward decorator:

from rnow.core import reward, RewardArgs

@reward
async def accuracy(args: RewardArgs, messages: list) -> float:
    """Check if the model's answer matches ground truth."""
    response = messages[-1]["content"]
    expected = args.metadata["answer"]
    return 1.0 if expected in response else 0.0

→ Write your first reward function

2. Tools (for Agents)

Give your model the ability to call functions during training:

from rnow.core import tool

@tool
def search(query: str, max_results: int = 5) -> dict:
    """Search the web for information."""
    # Your implementation here
    return {"results": [...]}

→ Train an agent with custom tools

3. Training Data

Create a train.jsonl file with your prompts and reward assignments:

{"messages": [{"role": "user", "content": "Balance the equation: Fe + O2 → Fe2O3"}], "rewards": ["accuracy"], "metadata": {"answer": "4Fe + 3O2 → 2Fe2O3"}}
{"messages": [{"role": "user", "content": "Balance the equation: H2 + O2 → H2O"}], "rewards": ["accuracy"], "metadata": {"answer": "2H2 + O2 → 2H2O"}}
{"messages": [{"role": "user", "content": "Balance the equation: N2 + H2 → NH3"}], "rewards": ["accuracy"], "metadata": {"answer": "N2 + 3H2 → 2NH3"}}

→ Learn about training data format

Contributing

We welcome contributions! ❤️ Please open an issue to discuss your ideas before submitting a PR

Name		Name	Last commit message	Last commit date
Latest commit History 45 Commits
.github/workflows		.github/workflows
assets		assets
rnow		rnow
.env		.env
.gitignore		.gitignore
.pre-commit-config.yaml		.pre-commit-config.yaml
.python-version		.python-version
LICENSE		LICENSE
Makefile		Makefile
README.md		README.md
pyproject.toml		pyproject.toml
tox.ini		tox.ini
uv.lock		uv.lock

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

Documentation

Quick Start

1. Install uv (Python package manager)

2. Install ReinforceNow

3. Authenticate

4. Create & Run Your First Project

Core Concepts

1. Reward Functions

2. Tools (for Agents)

3. Training Data

Contributing

About

Uh oh!

Releases

Contributors 2

Uh oh!

Languages

License

ReinforceNow/reinforcenow-cli

Folders and files

Latest commit

History

Repository files navigation

Documentation

Quick Start

1. Install uv (Python package manager)

2. Install ReinforceNow

3. Authenticate

4. Create & Run Your First Project

Core Concepts

1. Reward Functions

2. Tools (for Agents)

3. Training Data

Contributing

About

Resources

License

Uh oh!

Stars

Watchers

Forks

Releases

Contributors 2

Uh oh!

Languages