Fancy Gym: Unifying interface for various RL benchmarks with support for Black Box approaches.
-
Updated
Apr 17, 2024 - Python
Fancy Gym: Unifying interface for various RL benchmarks with support for Black Box approaches.
Cloud Computing
LLM extraction eval: 40 hand-labeled job postings, 8 fields, 3 models, scored in CI with a quality gate that fails the build on regression. Measures hallucination on absent fields, not just accuracy.
To associate your repository with the bechmarks topic, visit your repo's landing page and select "manage topics."