Ali Toygar Abak PRO
phionyx
AI & ML interests
AI governance, AI safety, multi-agent systems, reproducible LLM evaluation, agent protocols, open-source AI, local inference, and trustworthy AI
Recent Activity
updated a dataset about 16 hours ago
phionyx/airep-embedded-evaluation-profile updated a dataset about 16 hours ago
phionyx/airep-evidence-cases posted an update 3 days ago
AI evals can be reproducible, signed, and still overclaim what the evidence actually establishes.
A tool-call denial proves a denial. It does not prove that a safeguard prevented an incident.
For AI evaluation, agent evaluation, and AI safety assurance, provenance alone is not enough. We also need to preserve the scope, context, and limits of the claim.
That is the case for a claim-preserving evidence contract.
📄 Access Is Not Yet Verifiability
https://huggingface.co/blog/phionyx/access-is-not-yet-verifiabilityOrganizations
None yet