Cross-provider AI code review for Claude Code — evidence-based confidence scoring with Codex, Gemini & Claude
-
Updated
Sep 16, 2026 - Shell
Cross-provider AI code review for Claude Code — evidence-based confidence scoring with Codex, Gemini & Claude
Uncertainty based selection of compatible inputs
Extract structured data from any document — PDF, DOCX, HTML, CSV, plain text — using LLMs with Pydantic schema validation, per-field confidence scores, and source grounding.
Open-source LLM evaluation engine with statistical confidence scoring
Runtime reliability intelligence designed specifically for OpenClaw frameworks and agents.
Zero-Noise utilities for safer product research and review signal analysis.
Multi-agent AI task delegation architecture for n8n: orchestrator routes natural-language commands to specialist agents with confidence scoring and human-in-the-loop gates.
RAG-based document intelligence system with semantic retrieval, grounded Q&A, citations, confidence scoring, and guardrails.
System that aggregates outputs from multiple Large Language Models (GPT-4, Claude-3, custom models) to generate reliable, high-confidence results through consensus-based reasoning evaluation. Demonstrates sophisticated AI orchestration with 92.7% accuracy improvement over single-model.
Research-grade Self-Correcting RAG agent built with LangGraph that retrieves knowledge, generates answers, evaluates grounding/relevance/completeness, and iteratively self-improves with confidence scoring and memory.
Deterministic structured extraction from noisy LLM/OCR output. Zero LLM round-trips, microsecond latency, confidence score on every result. msgspec · Pydantic · dataclasses.
This project is a support ticket classifier that uses machine learning to classify support tickets into different categories. It uses FastAPI for the backend and Next.js for the frontend.
AI-powered concierge that normalises guest messages from WhatsApp, Booking.com, Airbnb, Instagram and direct channels, drafts a reply with Claude, and routes responses through a deterministic confidence-scoring pipeline. Built with FastAPI + Claude Sonnet 4.
Production-grade HR document intelligence system built on n8n, Pinecone, OpenAI, and PostgreSQL. Automatically detects and processes multiple files from a Google Drive folder, then answers natural-language queries against your HR documents with cited, confidence-scored responses — complete with query logging, caching, and error handling.
Verification system that catches coding agents falsely claiming task completion. Runs 4 parallel checks (file integrity, test quality, scope narrowing, optional LLM judge) over task+claim+diff and returns a weighted 0-100 confidence score with evidence.
Smart Document Conversion for the AI Era - CPU-only, fast, with confidence scoring. Converts PDF, DOCX, PPTX, HTML, EPUB to Markdown, JSON, HTML, Text.
RAG service with a policy engine that decides whether to answer, ask a clarifying question, or refuse before generation ever runs. Lexical retrieval, SQLite-backed chunk storage, page-level citations, and a mocked generation layer shaped for a drop-in real LLM call.
Enterprise-grade Confidence-Driven State Reconciliation Platform leveraging Kafka Streams, Redis, PostgreSQL, and Spring Boot to preserve competing truths, compute evidence-backed consensus, deterministic replay, and operational analytics.
PDF word-level confidence extraction skill with OCR modality detection, bounding box mapping, Gemini 3.7 Flash context-aware auto-correction, and interactive review dashboard.
Self-healing and drift recovery for agents. A zero-runtime-dependency TypeScript library and CLI that scores output confidence, detects ungrounded claims with GSAR-style typed grounding, monitors behavioral drift, and orchestrates recovery (rollback, retry, escalate, ask a human), with a tamper-evident audit trail.
To associate your repository with the confidence-scoring topic, visit your repo's landing page and select "manage topics."