Pinned Loading
Repositories
Showing 10 of 34 repositories
- tolokaforge Public
Universal LLM benchmarking harness for tool use, browser, mobile, coding, and long-horizon evals
- toloka-model-name-normalizer Public
One identity and rule-derived descriptive fields for every LLM model reference. Every parsing rule is configuration.
- n8n-nodes-tendem Public
- template-builder Public
-
- u-math Public
Official evaluation code for the U-MATH and μ-MATH benchmarks. These datasets are designed to test the mathematical reasoning and meta-evaluation capabilities of LLMs on university-level problems.
