AI Engineer Ā· Agentic Systems in Production Ā· Toronto
I build LLM agent systems that run on real infrastructure -- multi-agent orchestration, tool use, guardrails, and the compliance paths where a bug is a legal problem. Two are live right now: an AI voice receptionist answering real calls for real clients, and a nine-agent SMS outreach system with a 124-test suite behind it.
Alongside that I'm building DeepAnalytic, a retrieval system over the Stanford Encyclopedia of Philosophy -- an attempt to make a hard corpus genuinely answerable, and to find out by measuring rather than assuming. Five PhDs in the field are testing it now, with 50+ philosophers and PhD students in the coming months, on the theory that the people who know the corpus best will break it fastest.
Click any card below for the full write-up ā
Featured Work
Two systems built for live traffic -- one on the phone, one over SMS. Both are mine end to end: architecture, agent design, guardrails, deployment, and the failures that only show up once real numbers are on the other end.
Live -- answering real customer calls
Hear a real call
Built and shipped solo Ā· Real clients Ā· 2025ā26
A real-time voice agent that answers the call, qualifies the job in conversation, checks live calendar availability and books the slot before hanging up. Live with real clients, and now running across three verticals: home services, restaurants and clinics.
The market it's built for
A call forwards in when the business can't pick up. The agent answers with their greeting, works out whether it's an emergency or routine, collects and reads back the caller's name, number, address and the job itself, checks the real calendar for open slots, books one on the call, and texts both the owner and the caller a confirmation. Every call lands in Supabase with its transcript and disposition.
Configured per business rather than forked per business: greeting, voice, hours, service area, booking length, transfer number and scope all come from the customer record at call time.
Stack: Vapi Ā· Twilio Ā· Node/Express Ā· Supabase Ā· Google Calendar API Ā· Railway.
I've called a number of restaurant voice agents around Toronto and got several of them to tell me what model they run, what tools they hold and roughly what their system prompt says -- just by asking in the right way. Presumably each of those prompts contains an instruction not to. A rule written into a prompt gets followed most of the time, which is not the same as being enforced, and no amount of rewording closes that gap: the instruction is one more thing in the context, competing with everything else in it.
Part of the answer is to keep things out of the model's reach in the first place -- it can't disclose or misuse what it was never handed. Three places where the server decides instead:
That narrows the blast radius but doesn't close it. None of it stops an agent naming its provider or paraphrasing its instructions, because that's the model talking rather than the model acting, and there's no parameter to remove. The only real handle on it is detection -- assertions over transcripts and end-of-call reports -- which turns disclosure from something I hope doesn't happen into a rate I can measure. That's part of what the eval work below is for, and it's what I'm building now.
Live -- sending to real prospects
Sole engineer Ā· 2026
Three agents draft every cold message in parallel and a picker agent selects one against explicit criteria -- no single model settling for its first attempt. Replies go to a stateful SDR agent, and the rules that carry legal weight are enforced in code, not prompts.
Measured on real sends
A lead scraper pulls home-service businesses from Google Places into Supabase, and a hook agent reads each business profile to find a specific angle worth opening on -- that hook is what the three drafters compete over. Once the picker has chosen, the draft passes through code-level checks -- compliance footer, GSM-7 sanitisation, suppression list, line-type guard -- before Twilio sends it.
When a reply lands on the FastAPI webhook, a triage agent classifies it as interested, objection, not-now, or opt-out. Opt-outs are handled in code and never reach a model. Everything else goes to a stateful SDR agent that holds the conversation, grounded in a YAML knowledge base, and hands off to me the moment it hits something outside its scope. An operator console shows every thread with its full trace, and I can take over any conversation mid-flight.
Two guardrails wrap the reply path. A cheap keyword scope check runs first and catches the obvious out-of-bounds cases for nothing; an LLM classifier catches the rest. A separate grounding judge checks that every claim in a drafted reply traces back to the knowledge base, because the expensive failure here isn't rudeness -- it's an agent inventing a price or a capability to a real prospect.
An MCP server publishes the pipeline as nine tools -- prospects, threads, send logs, suppression state, delivery status and the rest -- so the whole system is inspectable from an MCP client without opening the database. Every one of them is read-only. Sending stays behind the same code-level checks it always did, because the point of the server is observability, not a second path to the send button: a tool that can text a real business is not a tool worth exposing for convenience.
Output guardrails run after tool calls, so an agent holding send_sms can deliver an ungrounded message before any check sees it -- and a text can't be unsent. So the agent returns structured output and code decides whether to send. The same reasoning drives retrieval: the knowledge base is lexical rather than vector, because an off-topic question must retrieve nothing, and similarity search always returns a nearest neighbour.
Five rules migrated out of prompts during development: the CASL opt-out footer, GSM-7 punctuation, the sub-4.0 ratings the hook agent kept building insulting angles from, and two others. An instruction followed 98% of the time isn't a guarantee, and twice, tightening the instruction caused the next failure instead of fixing it. The compliance tests assert the paths where a bug is a legal problem -- a suppressed number can never be sent to, opt-outs survive prospect deletion, a blocked message never returns silence.
app/agents/ hook Ā· 3 drafters Ā· picker Ā· triage Ā· draft Ā· SDR Ā· guardrails app/kb/ YAML knowledge base + lexical retrieval app/tools/ send_sms with suppression + line-type guards app/db/ Supabase client, schema, migrations app/routes/ FastAPI webhook + operator console app/mcp/ MCP server, 9 read-only tools over the pipeline tests/ 124 tests, 20 compliance
Retrieval & Applied ML
One retrieval system, built on a corpus hard enough that the retrieval strategy decides the answer -- and measured until it said things I didn't want to hear.
Live -- philosophers testing it nowš Access by request -- email me
Solo build Ā· Full corpus indexed Ā· 2026
111,048 chunks across the full encyclopedia, now in front of real philosophers -- and an evaluation habit that has contradicted me three times: the reranker hurt, the chunking strategy I built the thing around didn't matter, and scaling the corpus twentyfold improved retrieval when the literature said it should break. Being pointed next at a harder product -- a referee that checks a philosophy paper's claims against the literature.
Measured, against my own hypotheses
The app is password-protected while the pilot runs -- the link opens a login screen. Email me and I'll send you credentials.
A retrieval-augmented QA system over the full 2024 Stanford Encyclopedia of Philosophy, scraped and indexed myself -- 111,048 chunks across 1,803 articles. The corpus is a deliberately hard case: dense, heavily cross-referenced, full of terms whose meaning shifts between sub-fields.
The product is now pivoting toward a paper-referee tool: decompose a philosophy paper into individual claims, check each against the encyclopedia, and report which are supported, which are contradicted, and which the literature simply doesn't address. The chatbot becomes the companion feature for discussing that feedback rather than the thing itself. The retrieval work below is what makes the referee possible, and it's also what told me the referee was the better product.
Five PhDs in the field already have access. The wider round -- 50+ professional philosophers and PhD students -- is phase two, in the coming months. Deliberately staged: five people who know the corpus will surface more in a week than fifty will in a day, and I'd rather fix what they find before the larger group arrives than collect fifty reports of the same bug.
The app serves one pipeline with no mode switching, because testers should be judging whether answers are useful rather than evaluating my architecture. Two interface decisions came straight out of the retrieval findings. Sources are grouped under the sub-question that retrieved them rather than shown as a flat list, since multiquery produces no global ranking across sub-questions and a flat list would imply one that doesn't exist. Passages under 100 characters are hidden, which removes section headings that landed alone at a chunk boundary and would otherwise be cited as if they were evidence.
Conversation memory spans three turns and lives entirely in the app layer, leaving the pipeline stateless. Testers are told not to lean on it: self-contained questions retrieve better, and self-contained claims are the unit the referee will work with. Interactions and feedback log anonymously to three joined sheets -- the interactions, per-answer ratings split between answer quality and source quality, and the full text of every retrieved passage for later analysis. Emails are hashed to short pseudonymous ids, so a tester's questions can be grouped without storing who asked them. Cost is capped twice over: a hard budget limit on the API project, and a per-session query cap so one tester can't exhaust it in a sitting and so exhaustion produces a clear message instead of raw errors.
The published literature predicted degradation: Xiang et al. measured vector RAG accuracy on complex reasoning dropping 26% relative across a twentyfold expansion, attributed to retrieval capturing high-similarity but irrelevant noise as the search space grows. My expansion was 19.5Ć, near-identical in scale. Recall held completely -- recall@5, @10 and @20 unchanged, no question dropping out of the top 10, nothing falling out of the top 50.
One metric did appear to show crowding, rising 47%, and that reading was wrong. It counted chunks from other articles rather than wrong ones, and the 100-article baseline was an alphabetical slice that excluded most relevant material by construction. Head-query testing settled it: asked about supervenience, the small corpus returned nine chunks of Anomalous Monism, while the full corpus returned six of Supervenience -- the entry actually named after the concept. Four of five generic concepts behaved the same way. Paraphrase testing found rank moving at most one position under expansion, and rewording a question moving rank more than expanding the corpus did.
Two measurement bugs turned up in my own analysis along the way. Both are written up alongside the result, because a scaling finding that contradicts the literature is only worth anything if the method that produced it is inspectable.
The premise was that chunking along real section boundaries would beat fixed token windows. A five-way comparison -- section-aware, fixed-window, paragraph, sentence-window, semantic -- says otherwise: on straightforward questions the first three retrieved near-identical text, because paragraph breaks in a well-edited article already track topic breaks closely enough. Semantic chunking was the interesting failure, the only strategy that ever clearly won a question and the only one that ever clearly lost one, which disqualifies it as a default however good its ceiling. Section-aware stays because nothing beat it, not because it won.
Parse quality did something I didn't expect: it improved with scale rather than degrading. 95% clean at 100 articles, 96.3% at 300, 97.5% across all 1,803, with a single article failing outright.
What all of this points at is that the bottleneck sits downstream of chunking. On composite questions nearly every strategy retrieved both halves of the evidence and then declined to synthesize them, reporting no direct comparison in a context that plainly contained one. A chunk can mention or presuppose content it doesn't itself contain, and nothing notices the gap -- which is what multiquery decomposition is aimed at, now built and in the pipeline as multiquery.py, topic-first and schema-enforced.
Methodology follows current RAG-eval practice: LLM-assisted test-set generation paired with LLM-as-judge scoring and human verification. Claude generated the article selection, questions and rubric, built the automation script, and did first-pass scoring; I scored independently afterward, spot-checking the highest-stakes rows against the retrieved chunk text directly. Both score sets sit side by side in tests/eval_scored.xlsx, including a row where my pass corrected Claude's. The pilot replaces this with something better: real questions from people who already know what a right answer looks like.
app.py Streamlit pilot: auth, query cap, memory, feedback logging ingest.py SEP corpus ā section split ā chunk ā embed ā Pinecone rag_pipeline.py retrieve ā rerank ā generate, plus naive and multiquery paths multiquery.py topic-first, schema-enforced question decomposition section_parser.py TOC-aware chunking with logged fallback chain run_eval.py runs the eval set through both modes, logs results check_chunks.py inspect what actually got retrieved for a query config.py environment-driven settings, no secrets in code tests/ question set, raw results, rubric, scored output notebooks/ four experiments, each with its own findings log
Phase two in the coming months, then the referee work: objection-and-response checking against documented SEP objections, argument reconstruction into premises and conclusion for per-premise checking, and user-uploadable documents so a draft can be checked against a philosopher's own citation list rather than only the encyclopedia.
Further out: subsection indexing to enable multi-hop retrieval, a FastAPI service, and an open-model option alongside OpenAI. One research question is parked rather than dropped: the scaling literature treats corpus size as the variable while composition goes unisolated, and adding canonically relevant documents is not the same operation as adding adjacent ones. My own result fits the composition hypothesis. Testing it properly needs a second corpus and a much larger question set.
Solo build Ā· Experiment
Built to answer one question: does a hand-rolled planner/executor beat the SDK's native tool-calling loop? It doesn't, and the write-up says why.
I started with a custom planner/executor, then switched to native @function_tool with tool_choice="required". The reason is specific: in the custom version each tool call was planned before prior results existed, whereas the SDK's conversation loop puts real results in context before the next call is chosen. Planning ahead of your own outputs is planning blind.
Forced tool use means the model must pick something, so on trivial inputs it reached for whichever tool was nearest and produced nonsense. Adding DirectAnswerTool as an explicit "no tool applies" option fixed it -- the constraint needed an escape hatch, not a better prompt.
Solo build Ā· 2023
Give it a recipe you like and it returns the five closest matches from 520k scraped from food.com. My BrainStation capstone, extended past the assignment and deployed -- still running three years later.
Embeddings over ingredients, instructions and themes, so the query can be a whole recipe rather than a keyword. Scraped with Selenium, embedded, deployed. Stack: bertopic Ā· spacy Ā· nltk Ā· pandas / numpy Ā· openai Ā· tiktoken Ā· scikit-learn Ā· umap Ā· selenium Ā· streamlit.
Single-vector cosine similarity, which was a reasonable 2023 choice and is commodity in 2026 -- DeepAnalytic is where the retrieval thinking lives now. Kept here because it has been deployed and usable for three years, and because scraping and embedding half a million records end to end is the part most tutorials skip.
Team of 4 Ā· 24 hours Ā· 1st place
Won a 24-hour industry hackathon designing "Insights" -- an LLM widget giving users plain-language explanations of how Google's AI features use their data.
Repository āA one-day hackathon hosted by BrainStation for Google, July 2023. Our team tackled growing public mistrust of AI: no transparency into how AI-driven results are produced, and no way for a user to verify anything.
"Insights," a widget giving users an on-demand explanation of how AI is using their data. A hover reveals a summary; a click opens a chatbot -- envisioned on an LLM fine-tuned on Google's product docs -- that answers deeper questions in the moment, turning a black box into something users can interrogate. Engagement with the widget doubles as a trust signal, analyzed with ANOVA and A/B testing to feed back into product decisions.
Distributed big-data wrangling, SQL and dashboarding, NLP classification, public-health prediction, and hospital readmission modelling -- the statistical and data-engineering groundwork under the agent work above.
Big Data Wrangling ā Google Books Ngrams
Loading, filtering, and visualizing a dataset covering roughly 4% of all printed books (1800sā2000s) in a cloud distributed-computing environment with Hadoop, Spark, Hive, and S3. Repository ā
Bixi Bike Usage ā Trend Analysis
SQL-driven exploration of Montreal's bike-share system across the 2016ā2017 seasons, paired with an interactive drill-down Tableau dashboard on usage drivers and station demand. SQL ā
Sentiment Analysis ā Hotel Reviews
NLP pipeline from EDA through feature engineering to sentiment classification -- comparing Logistic Regression, PCA, KNN, and Decision Trees, refined with cross-validation. Repository ā
West Nile Virus ā Predictive Analysis
Statistical and predictive modeling of Chicago mosquito trap data for the city's public health monitoring program -- chi-square testing, then linear and logistic regression on WNV prevalence. Repository ā
U.S. Diabetes Readmission
Explored ~100k U.S. hospital records (1999ā2008) to identify sub-groups at higher readmission risk, then modeled readmission with logistic regression. Repository ā
Research & Public Voice
AI safety, alignment, and what it takes for a system to actually follow a rule.
Conversations at the intersection of AI and philosophy with researchers in AI safety, ethics, and emerging tech. 13 episodes and counting, on YouTube and Spotify.
A short, sharp book on the mathematics underneath data science -- linear algebra, calculus, statistics, and the math behind core ML algorithms, with Python code.
Amazon ā Book PDF āMost books teaching the mathematics under data science run ~350 pages, cost a lot, lean heavily on code, and teach surprisingly little actual math. This one is built to fix that: short, sharp, and math-first.
Part I covers the core mathematics -- linear algebra, calculus, statistics, and probability theory. Part II builds on it to cover the mathematics underneath regression, tree methods, support vector machines, clustering, PCA, and neural networks, with Python implementations throughout.
Peer-reviewed paper in Synthese providing a unified semantics for propositional logics of impure ground using structured propositions.
Read paper āI show that the assumption of highly structured propositions can be leveraged to provide a unified semantics for various propositional logics of impure ground in a very expressive and flexible way. It is shown, in particular, that the induced models are capable of capturing an infinitude of grounding facts that follow from unrestricted logics of ground, but, due to certain artificial restrictions, are left unaccounted for by the existing semantics in the literature. It is also shown that our models, unlike the ones in the literature, are easily extendable to capture certain distinct views about iterated as well as identity grounding.
Open-source guide to descriptive and inferential statistics for real-world analytics, paired with reproducible Python code.
Repository āAn open-source, hands-on guide to the descriptive and inferential statistics needed in real analytics and ML pipelines. Each concept is paired with clear intuition, minimal proofs, and fully reproducible Python code.
Covers measures of central tendency and dispersion, hypothesis testing, A/B testing, and regression models, illustrated with end-to-end examples using pandas, NumPy, SciPy, and statsmodels.
University of Calgary (2023) Ā· University of Tartu (2019)
After years of exposure to mathematical logic in my undergraduate years, I grew interested in the philosophical ideas behind math and logic and did my MA in Philosophy at the University of Tartu in Estonia.
I then continued into a PhD at the University of Calgary, where over 4 years I engaged in research in Logic and Metaphysics.
My MA thesis, Perception, Abductive Methodology, and Compositional Universalism, argued for compositional universalism -- the view that any plurality of objects constitutes an object. Read the thesis ā
My PhD research focused on questions of fundamentality, with a secondary interest in IT ethics, AI safety, and machine consciousness that I still work in today. Read the dissertation ā
One dissertation chapter was later published in Synthese. Read the paper ā
Taught a graduate-level logic course on modern modal logics to 10+ masters and PhD students across Philosophy, Mathematics, and Semiotics during my MA.
During my PhD, taught hundreds of students on AI, IT, and data ethics -- deepening my grounding in AI safety debates. I write on these topics at substack.com/@llminds and occasionally for Wikipedia and the Centre for Social Impact Technology.
MA scholarships (over $24,000):
PhD scholarships (over $50,000):
Tehran Polytechnic (Amirkabir University of Technology), 2017
Undergraduate in Mathematics and Applications at Amirkabir University of Technology in Iran. My final thesis was on mathematical logic, with broad coursework across pure and applied math.
Thesis on coalgebraic logics -- Logics for Coalgebras, and a Final Coalgebra Theorem (written in Persian). Read the thesis ā
3-month intensive data science bootcamp
Cloud & ML infrastructure certification
Always glad to talk about agent systems that have to survive real traffic -- evaluation, guardrails, and the failures that only surface once something is live. Open to agentic AI and applied ML work in Toronto or remote.
Ā© 2026 Amir Kiani