CLEAR-S Eval Harness
A layered evaluation system for testing agent outputs, trajectories, safety, latency, and cost across prompts, frameworks, and models.
- MLflow
- agents
- evaluation
- traces
AI systems, field notes, and workshop material.
Open source · demos · experiments
Open-source tools, reference implementations, and research demos built to make an AI question concrete enough to test.
A layered evaluation system for testing agent outputs, trajectories, safety, latency, and cost across prompts, frameworks, and models.
An autonomous AI and ML engineer that researches, trains, measures, reproduces, and serves what it builds using Databricks-native primitives.
A Python package that converts PDF, PowerPoint, and Word files into Markdown and layout-aware JSON using multimodal models.
A reference implementation for streaming agent events, tool activity, approvals, and errors through FastAPI and WebSockets.
An experiment in parallel hypothesis generation and measured iteration using Databricks serverless jobs.
A governed web crawling and question-answering application that follows site links, builds a searchable corpus, and exposes a conversational interface.