Skip to content
View rctruta's full-sized avatar

Block or report rctruta

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rctruta/README.md

Ramona C. Truta

Hi, I'm Ramona, an Independent Researcher, Solutions Architect and Educator. I bring over 20 years of rigorous database engineering experience to the world of AI; my current focus is on System Integrity—applying the strict determinism of data engineering to the probabilistic nature of AI Agents. More to the point, I build instruments that measure whether database and AI agents do, and then publish the experiment alongside the result. In that way, you can reproduce it or show me where I'm wrong. This site is built the same way: generated from data, and the build fails on an unknown tag, a broken anchor, or an unregistered number.

Portfolio · Writing · LinkedIn · ORCID


sqlbenchdag — a reproducible SQL benchmarking laboratory

sql-benchmarks-dagster · pip install sqlbenchdag · MCP server in the official registry

Every experiment is a capsule addressed by an 8-character SHA-256 fingerprint over its config, SQL, and every line of measurement-relevant Python. Change the method, the ID changes. Cold-cache execution, hard row-count assertions, an integrity seal on every capsule — and OpenTimestamps proofs anchored to Bitcoin on the four published Quack capsules.

Finding Capsule
DuckDB-over-Quack (pushdown) beats PostgreSQL at every scale past the noise floor — 4.4× @100K, 6.1× @1M, 13.2× @10M rows 902d1277
Attach-mode overhead grows with scan size (2.6× @100K → 9.5× @10M); pushdown stays flat at ~2× b8e2bfaf
That flat ~2× residual is reduced server-side parallelism, not protocol transport 25b0e134
Attach mode cannot execute multi-table joins at all; pushdown holds ~1.8× on a 3-way TPC-H join b198363e

Full published-capsule index →



Projects

Data modeling

rctruta.github.io The repository powering this website is built as a data modeling project. →

Benchmarking

lakehouse-semantics A suite of 23 correctness probes executed across DuckDB, DuckDB+Delta, PostgreSQL, and Databricks SQL. →

AI evaluation

harness-bench I engineered harness-bench to evaluate agent deployments before they reach production. →
malloy-publisher-agent-study The Malloy Publisher ships agent skills alongside its MCP server. →
bauplan-agent-parse-study Per-turn agent traces against a commercial lakehouse platform. Surfaced three defects I reported upstream — #369, #370, #371 — including an error message that actively sends an agent in the wrong direction. →
agent-telemetry A local analytics package that parses raw JSONL session transcripts from coding agents (Claude Code) into un-opinionated process metrics: tool_calls_per_turn, reads_before_first_edit, repeated_commands, error_rate, seconds_to_first_write, and edit_revisits. →
podcast-rag Ask a question about a podcast, get an answer and a link to the exact second someone said it. LanceDB, local-first, no cluster and no API key. Built a labelled eval set and found it answering 17 of 20 questions its sources couldn't support; recalibrated precision 0.56 → 0.87. →
music_recommendation_system Top-10 song recommendation over the Million Song Dataset Taste Profile Subset: two million listening events, ten thousand songs. →

AI security

adversarial-judgement-research Metatracing engine and execution traces for structural failure modes in multi-agent LLM consensus pipelines: status bias, persona bleed, frame break, axiomatic refusal, consensus contagion, semantic camouflage. →
ai-security-testbed A deterministic matrix engine for adversarial audit of multi-agent pipelines: temperature 0, replication, status and thinking-budget sweeps, rule-based grading. 416 traces across frontier and local models. →
ai-agent-utils Boilerplate and security guidelines for collaborating safely with autonomous coding agents, with the gates already enforced rather than written down and hoped for. →

Writing

Everything →

Speaking

  • Neo4j NODES — Structure Is Not Security: Poisoning Graph-Based Agent Memory Through the Extraction Pipeline · virtual · November 12, 2026
  • Data in the D — Measure What Matters · Detroit · October 16, 2026
  • Canadian Women in Cybersecurity — Consensus Contagion: Status Bias and the RLHF Tax in Agentic AI Routing Nodes · Toronto · May 26, 2026
  • AI Tinkerers Toronto — Semantic Laundering: Agentic Memory · Toronto · February 26, 2026

Everything →

Pinned Loading

  1. sql-benchmarks-dagster sql-benchmarks-dagster Public

    SQL Benchmarking Laboratory using Dagster for Orchestration

    Python 2

  2. harness-bench harness-bench Public

    Before an organization deploys LLM agents over MCP servers / tool APIs, harness-bench measures what that deployment will actually cost and where it will actually fail.

    Python

  3. ai-agent-utils ai-agent-utils Public

    A collection of boilerplate scripts and security guidelines for safely collaborating with autonomous coding agents

    Shell 1

  4. podcast-rag podcast-rag Public

    Ask a question about a podcast. Get an answer, and get sent to the exact second someone said it.

    Python

  5. music_recommendation_system music_recommendation_system Public

    Capstone Project for the Applied Data Science Program, MIT Professional Education

    Jupyter Notebook 3

  6. awesome-pedantic-medallion awesome-pedantic-medallion Public

    A community-curated collection of playbooks for building shared understanding. This is the official repository for the Pedantic Medallion framework.

    1