Skip to content
@Toloka

Toloka

Data labeling platform for ML

Pinned Loading

  1. tolokaforge tolokaforge Public

    Universal LLM benchmarking harness for tool use, browser, mobile, coding, and long-horizon evals

    Python 18 10

  2. tendem-evaluation tendem-evaluation Public

    Tendem hybrid AI+Human system benchmarking

    Python 3

  3. beemo beemo Public

    Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

    11 2

  4. u-math u-math Public

    Official evaluation code for the U-MATH and μ-MATH benchmarks. These datasets are designed to test the mathematical reasoning and meta-evaluation capabilities of LLMs on university-level problems.

    Python 11 3

  5. crowd-kit crowd-kit Public

    Control the quality of your labeled data with the Python tools you already know.

    Python 254 21

Repositories

Showing 10 of 34 repositories