AI Engineer Ā· Agentic Systems in Production Ā· Toronto

Amir Kiani

I build LLM agent systems that run on real infrastructure -- multi-agent orchestration, tool use, guardrails, and the compliance paths where a bug is a legal problem. Two are live right now: an AI voice receptionist answering real calls for real clients, and a nine-agent SMS outreach system with a 124-test suite behind it.

Alongside that I'm building DeepAnalytic, a retrieval system over the Stanford Encyclopedia of Philosophy -- an attempt to make a hard corpus genuinely answerable, and to find out by measuring rather than assuming. Five PhDs in the field are testing it now, with 50+ philosophers and PhD students in the coming months, on the theory that the people who know the corpus best will break it fastest.

See the agent systems → GitHub
Amir Kiani

Click any card below for the full write-up ↓

Featured Work

Agentic AI Systems in Production

Two systems built for live traffic -- one on the phone, one over SMS. Both are mine end to end: architecture, agent design, guardrails, deployment, and the failures that only show up once real numbers are on the other end.

VoiceCaptures -- AI voice receptionist that answers, qualifies and books calls Live -- answering real customer calls

Hear a real call

Built and shipped solo Ā· Real clients Ā· 2025–26

VoiceCaptures — AI Voice Receptionist

A real-time voice agent that answers the call, qualifies the job in conversation, checks live calendar availability and books the slot before hanging up. Live with real clients, and now running across three verticals: home services, restaurants and clinics.

The market it's built for

14%
missed-call rate in home services (CallRail, 2025)
85%
of callers whose call goes unanswered never call back (CallRail, 2025)
$500–$1,200
typical value of one missed service call (answeringagent.com, 2026)
Vapi + Twilio Node/Express Ā· Supabase Tool use (Calendar API) RAG over per-business KB Server-side guardrails

What it does

A call forwards in when the business can't pick up. The agent answers with their greeting, works out whether it's an emergency or routine, collects and reads back the caller's name, number, address and the job itself, checks the real calendar for open slots, books one on the call, and texts both the owner and the caller a confirmation. Every call lands in Supabase with its transcript and disposition.

Configured per business rather than forked per business: greeting, voice, hours, service area, booking length, transfer number and scope all come from the customer record at call time.

Stack: Vapi Ā· Twilio Ā· Node/Express Ā· Supabase Ā· Google Calendar API Ā· Railway.

Constraints that don't live in the prompt

I've called a number of restaurant voice agents around Toronto and got several of them to tell me what model they run, what tools they hold and roughly what their system prompt says -- just by asking in the right way. Presumably each of those prompts contains an instruction not to. A rule written into a prompt gets followed most of the time, which is not the same as being enforced, and no amount of rewording closes that gap: the instruction is one more thing in the context, competing with everything else in it.

Part of the answer is to keep things out of the model's reach in the first place -- it can't disclose or misuse what it was never handed. Three places where the server decides instead:

  • The model never supplies a phone number for an SMS. It picks a role -- caller, dispatch, or both -- and the server resolves the real numbers from the customer record and the call. A misheard callback number can't redirect a text, because there's no number parameter to get wrong.
  • Availability comes back from the server as the sentence to read out, so open slots can't be paraphrased into something untrue.
  • The business is identified from the call payload, never from anything the model asserts, and there's no single-tenant fallback to guess with.

That narrows the blast radius but doesn't close it. None of it stops an agent naming its provider or paraphrasing its instructions, because that's the model talking rather than the model acting, and there's no parameter to remove. The only real handle on it is detection -- assertions over transcripts and end-of-call reports -- which turns disclosure from something I hope doesn't happen into a rate I can measure. That's part of what the eval work below is for, and it's what I'm building now.

What's still open

  • Disclosure testing. No scripted probes yet for what the agent will say about itself under pressure -- its model, its tools, the contents of its instructions. It's the failure I went looking for in other people's agents and I don't currently measure it in my own.
  • Structured outputs. Calls still produce free-form model behaviour rather than a defined schema. A fixed shape for what a call must yield makes tool calls deterministic and makes evaluation possible at all -- right now there's nothing to diff against.
  • An automated eval harness. The loop today is: place a call, read the transcript, fix by hand. The target is scripted caller scenarios per vertical -- emergency, routine booking, no availability, hangup mid-flow, garbled address -- fired automatically, asserting on outcomes: did the row appear, did both texts fire, did the calendar event land, did the call end only when it should. Pass/fail per scenario, so a prompt change is regression-tested in minutes.
  • Latency and cost instrumentation. No stage-level p50/p95 split across ASR, LLM and TTS yet, and no cost per call. I'd rather have the breakdown before trading anything for it.
  • Flow control. Vapi Workflows may replace prompt-driven control with explicit graph nodes, removing a class of fragility instead of patching it. Worth evaluating before adding more prompt scaffolding.
Three drafting agents write competing messages in parallel; a picker agent selects one; Twilio sends it Live -- sending to real prospects

Sole engineer Ā· 2026

SMS Outreach & SDR Agent System

Three agents draft every cold message in parallel and a picker agent selects one against explicit criteria -- no single model settling for its first attempt. Replies go to a stateful SDR agent, and the rules that carry legal weight are enforced in code, not prompts.

Measured on real sends

44%
of scraped numbers could actually receive SMS
$0.020
true cost per delivered message, not $0.009
124
tests, 20 covering CASL compliance paths
OpenAI Agents SDK Multi-agent orchestration MCP server Guardrails FastAPI Twilio Supabase

What it does, end to end

A lead scraper pulls home-service businesses from Google Places into Supabase, and a hook agent reads each business profile to find a specific angle worth opening on -- that hook is what the three drafters compete over. Once the picker has chosen, the draft passes through code-level checks -- compliance footer, GSM-7 sanitisation, suppression list, line-type guard -- before Twilio sends it.

When a reply lands on the FastAPI webhook, a triage agent classifies it as interested, objection, not-now, or opt-out. Opt-outs are handled in code and never reach a model. Everything else goes to a stateful SDR agent that holds the conversation, grounded in a YAML knowledge base, and hands off to me the moment it hits something outside its scope. An operator console shows every thread with its full trace, and I can take over any conversation mid-flight.

The two agents that only exist to say no

Two guardrails wrap the reply path. A cheap keyword scope check runs first and catches the obvious out-of-bounds cases for nothing; an LLM classifier catches the rest. A separate grounding judge checks that every claim in a drafted reply traces back to the knowledge base, because the expensive failure here isn't rudeness -- it's an agent inventing a price or a capability to a real prospect.

The pipeline is exposed over MCP, read-only

An MCP server publishes the pipeline as nine tools -- prospects, threads, send logs, suppression state, delivery status and the rest -- so the whole system is inspectable from an MCP client without opening the database. Every one of them is read-only. Sending stays behind the same code-level checks it always did, because the point of the server is observability, not a second path to the send button: a tool that can text a real business is not a tool worth exposing for convenience.

The autopilot agent holds no tools

Output guardrails run after tool calls, so an agent holding send_sms can deliver an ungrounded message before any check sees it -- and a text can't be unsent. So the agent returns structured output and code decides whether to send. The same reasoning drives retrieval: the knowledge base is lexical rather than vector, because an off-topic question must retrieve nothing, and similarity search always returns a nearest neighbour.

Rules that must always hold live in code, not prompts

Five rules migrated out of prompts during development: the CASL opt-out footer, GSM-7 punctuation, the sub-4.0 ratings the hook agent kept building insulting angles from, and two others. An instruction followed 98% of the time isn't a guarantee, and twice, tightening the instruction caused the next failure instead of fixing it. The compliance tests assert the paths where a bug is a legal problem -- a suppressed number can never be sent to, opt-outs survive prospect deletion, a blocked message never returns silence.

Layout

app/agents/   hook Ā· 3 drafters Ā· picker Ā· triage Ā· draft Ā· SDR Ā· guardrails
app/kb/       YAML knowledge base + lexical retrieval
app/tools/    send_sms with suppression + line-type guards
app/db/       Supabase client, schema, migrations
app/routes/   FastAPI webhook + operator console
app/mcp/      MCP server, 9 read-only tools over the pipeline
tests/        124 tests, 20 compliance

Retrieval & Applied ML

RAG & Semantic Search

One retrieval system, built on a corpus hard enough that the retrieval strategy decides the answer -- and measured until it said things I didn't want to hear.

DeepAnalytic retrieval pipeline Live -- philosophers testing it nowšŸ”’ Access by request -- email me

Solo build Ā· Full corpus indexed Ā· 2026

DeepAnalytic — RAG over the Stanford Encyclopedia of Philosophy

111,048 chunks across the full encyclopedia, now in front of real philosophers -- and an evaluation habit that has contradicted me three times: the reranker hurt, the chunking strategy I built the thing around didn't matter, and scaling the corpus twentyfold improved retrieval when the literature said it should break. Being pointed next at a harder product -- a referee that checks a philosophy paper's claims against the literature.

Measured, against my own hypotheses

19.5Ɨ
corpus expansion, recall@5/10/20 unchanged
4.11 → 3.74
overall score fell when reranking was added
97.5%
of 1,803 articles parsed on their own section boundaries
Section-aware chunking Multiquery decomposition Streamlit pilot Pinecone Retrieval evaluation

The app is password-protected while the pilot runs -- the link opens a login screen. Email me and I'll send you credentials.

What it is, and where it's going

A retrieval-augmented QA system over the full 2024 Stanford Encyclopedia of Philosophy, scraped and indexed myself -- 111,048 chunks across 1,803 articles. The corpus is a deliberately hard case: dense, heavily cross-referenced, full of terms whose meaning shifts between sub-fields.

The product is now pivoting toward a paper-referee tool: decompose a philosophy paper into individual claims, check each against the encyclopedia, and report which are supported, which are contradicted, and which the literature simply doesn't address. The chatbot becomes the companion feature for discussing that feedback rather than the thing itself. The retrieval work below is what makes the referee possible, and it's also what told me the referee was the better product.

The pilot, running now

SEP Assistant answering a question about induction, abduction and van Fraassen's bad lot argument, with grouped sources and a feedback form

Five PhDs in the field already have access. The wider round -- 50+ professional philosophers and PhD students -- is phase two, in the coming months. Deliberately staged: five people who know the corpus will surface more in a week than fifty will in a day, and I'd rather fix what they find before the larger group arrives than collect fifty reports of the same bug.

The app serves one pipeline with no mode switching, because testers should be judging whether answers are useful rather than evaluating my architecture. Two interface decisions came straight out of the retrieval findings. Sources are grouped under the sub-question that retrieved them rather than shown as a flat list, since multiquery produces no global ranking across sub-questions and a flat list would imply one that doesn't exist. Passages under 100 characters are hidden, which removes section headings that landed alone at a chunk boundary and would otherwise be cited as if they were evidence.

Conversation memory spans three turns and lives entirely in the app layer, leaving the pipeline stateless. Testers are told not to lean on it: self-contained questions retrieve better, and self-contained claims are the unit the referee will work with. Interactions and feedback log anonymously to three joined sheets -- the interactions, per-answer ratings split between answer quality and source quality, and the full text of every retrieved passage for later analysis. Emails are hashed to short pseudonymous ids, so a tester's questions can be grouped without storing who asked them. Cost is capped twice over: a hard budget limit on the API project, and a per-session query cap so one tester can't exhaust it in a sitting and so exhaustion produces a clear message instead of raw errors.

I expected scaling to break retrieval. It fixed it.

The published literature predicted degradation: Xiang et al. measured vector RAG accuracy on complex reasoning dropping 26% relative across a twentyfold expansion, attributed to retrieval capturing high-similarity but irrelevant noise as the search space grows. My expansion was 19.5Ɨ, near-identical in scale. Recall held completely -- recall@5, @10 and @20 unchanged, no question dropping out of the top 10, nothing falling out of the top 50.

One metric did appear to show crowding, rising 47%, and that reading was wrong. It counted chunks from other articles rather than wrong ones, and the 100-article baseline was an alphabetical slice that excluded most relevant material by construction. Head-query testing settled it: asked about supervenience, the small corpus returned nine chunks of Anomalous Monism, while the full corpus returned six of Supervenience -- the entry actually named after the concept. Four of five generic concepts behaved the same way. Paraphrase testing found rank moving at most one position under expansion, and rewording a question moving rank more than expanding the corpus did.

Two measurement bugs turned up in my own analysis along the way. Both are written up alongside the result, because a scaling finding that contradicts the literature is only worth anything if the method that produced it is inspectable.

I was wrong about chunking too

The premise was that chunking along real section boundaries would beat fixed token windows. A five-way comparison -- section-aware, fixed-window, paragraph, sentence-window, semantic -- says otherwise: on straightforward questions the first three retrieved near-identical text, because paragraph breaks in a well-edited article already track topic breaks closely enough. Semantic chunking was the interesting failure, the only strategy that ever clearly won a question and the only one that ever clearly lost one, which disqualifies it as a default however good its ceiling. Section-aware stays because nothing beat it, not because it won.

Parse quality did something I didn't expect: it improved with scale rather than degrading. 95% clean at 100 articles, 96.3% at 300, 97.5% across all 1,803, with a single article failing outright.

What all of this points at is that the bottleneck sits downstream of chunking. On composite questions nearly every strategy retrieved both halves of the evidence and then declined to synthesize them, reporting no direct comparison in a context that plainly contained one. A chunk can mention or presuppose content it doesn't itself contain, and nothing notices the gap -- which is what multiquery decomposition is aimed at, now built and in the pipeline as multiquery.py, topic-first and schema-enforced.

Known open issues

  • The Oruka bug. One question consistently retrieves a wrong-but-adjacent section. Diagnosed: the correct chunk ranks 22nd by pure vector similarity, from content dilution plus crowding by thematically adjacent chunks elsewhere. A TF-IDF lexical blend moved it to rank 8 with zero regression on two controls -- but that fix lives in a notebook and is not wired into the pipeline yet.
  • Refusal detection is a stopgap. When a question falls outside the corpus, vector search still returns nearest neighbours, and showing them beside a refusal makes the app look like it found something relevant and ignored it. Refusals are currently detected by matching phrasing in the answer and gating on length, which is brittle -- it depends on wording the prompt encourages but can't guarantee, and it broke silently once already when the prompt changed. The principled fix is a similarity threshold, which needs scores surfaced through the pipeline.
  • Multiquery rerank collapse. At 18 subqueries the reranked path returns nothing at all, despite correct retrieval behind it. The non-reranked path handles the same case. Found, not yet diagnosed -- and largely moot once reranking stops being the default.
  • Subsections aren't indexed. Parsing stops at top-level headings, so a pointer like "as we saw in section 1.2" has nothing to resolve against, and within-article failures can't be measured automatically -- only found by hand. Prerequisite for the cross-reference work.

How the evaluation was built

Methodology follows current RAG-eval practice: LLM-assisted test-set generation paired with LLM-as-judge scoring and human verification. Claude generated the article selection, questions and rubric, built the automation script, and did first-pass scoring; I scored independently afterward, spot-checking the highest-stakes rows against the retrieved chunk text directly. Both score sets sit side by side in tests/eval_scored.xlsx, including a row where my pass corrected Claude's. The pilot replaces this with something better: real questions from people who already know what a right answer looks like.

Write-ups

What's in the repo

app.py              Streamlit pilot: auth, query cap, memory, feedback logging
ingest.py           SEP corpus → section split → chunk → embed → Pinecone
rag_pipeline.py     retrieve → rerank → generate, plus naive and multiquery paths
multiquery.py       topic-first, schema-enforced question decomposition
section_parser.py   TOC-aware chunking with logged fallback chain
run_eval.py         runs the eval set through both modes, logs results
check_chunks.py     inspect what actually got retrieved for a query
config.py           environment-driven settings, no secrets in code
tests/              question set, raw results, rubric, scored output
notebooks/          four experiments, each with its own findings log

What's next

Phase two in the coming months, then the referee work: objection-and-response checking against documented SEP objections, argument reconstruction into premises and conclusion for per-premise checking, and user-uploadable documents so a draft can be checked against a philosopher's own citation list rather than only the encyclopedia.

Further out: subsection indexing to enable multi-hop retrieval, a FastAPI service, and an open-model option alongside OpenAI. One research question is parked rather than dropped: the scaling literature treats corpus size as the variable while composition goes unisolated, and adding canonically relevant documents is not the same operation as adding adjacent ones. My own result fits the composition hypothesis. Testing it properly needs a second corpus and a much larger question set.

Smaller builds & experiments

Solo build Ā· Experiment

Agentic Task Runner

Built to answer one question: does a hand-rolled planner/executor beat the SDK's native tool-calling loop? It doesn't, and the write-up says why.

FastAPI OpenAI Agents SDK tool_choice=required
Repository →

The decision that changed the design

I started with a custom planner/executor, then switched to native @function_tool with tool_choice="required". The reason is specific: in the custom version each tool call was planned before prior results existed, whereas the SDK's conversation loop puts real results in context before the next call is chosen. Planning ahead of your own outputs is planning blind.

The failure that needed an extra tool

Forced tool use means the model must pick something, so on trivial inputs it reached for whichever tool was nearest and produced nonsense. Adding DirectAnswerTool as an explicit "no tool applies" option fixed it -- the constraint needed an escape hatch, not a better prompt.

Live on Streamlit -- try it

Solo build Ā· 2023

DeepChef — Semantic Recommender

Give it a recipe you like and it returns the five closest matches from 520k scraped from food.com. My BrainStation capstone, extended past the assignment and deployed -- still running three years later.

OpenAI embeddings Semantic search BERTopic Ā· UMAP Streamlit

Walkthrough

What it does

Embeddings over ingredients, instructions and themes, so the query can be a whole recipe rather than a keyword. Scraped with Selenium, embedded, deployed. Stack: bertopic Ā· spacy Ā· nltk Ā· pandas / numpy Ā· openai Ā· tiktoken Ā· scikit-learn Ā· umap Ā· selenium Ā· streamlit.

Where it sits now

Single-vector cosine similarity, which was a reasonable 2023 choice and is commodity in 2026 -- DeepAnalytic is where the retrieval thinking lives now. Kept here because it has been deployed and usable for three years, and because scraping and embedding half a million records end to end is the part most tutorials skip.

Team of 4 Ā· 24 hours Ā· 1st place

Google Ɨ BrainStation Hackathon — 1st Place

Won a 24-hour industry hackathon designing "Insights" -- an LLM widget giving users plain-language explanations of how Google's AI features use their data.

Repository →

The problem we picked

A one-day hackathon hosted by BrainStation for Google, July 2023. Our team tackled growing public mistrust of AI: no transparency into how AI-driven results are produced, and no way for a user to verify anything.

What we designed

"Insights," a widget giving users an on-demand explanation of how AI is using their data. A hover reveals a summary; a click opens a chatbot -- envisioned on an LLM fine-tuned on Google's product docs -- that answers deeper questions in the moment, turning a black box into something users can interrogate. Engagement with the widget doubles as a trust signal, analyzed with ANOVA and A/B testing to feed back into product decisions.

Google Insights concept slide 1st Place Industry Project certificate

Earlier data science & analytics

Five analytics & ML studies

Distributed big-data wrangling, SQL and dashboarding, NLP classification, public-health prediction, and hospital readmission modelling -- the statistical and data-engineering groundwork under the agent work above.

Spark / PySpark Hadoop Ā· Hive Ā· S3 SQL Ā· Tableau Regression Ā· Classification

Big Data Wrangling — Google Books Ngrams

Loading, filtering, and visualizing a dataset covering roughly 4% of all printed books (1800s–2000s) in a cloud distributed-computing environment with Hadoop, Spark, Hive, and S3. Repository →

Bixi Bike Usage — Trend Analysis

SQL-driven exploration of Montreal's bike-share system across the 2016–2017 seasons, paired with an interactive drill-down Tableau dashboard on usage drivers and station demand. SQL →

Sentiment Analysis — Hotel Reviews

NLP pipeline from EDA through feature engineering to sentiment classification -- comparing Logistic Regression, PCA, KNN, and Decision Trees, refined with cross-validation. Repository →

West Nile Virus — Predictive Analysis

Statistical and predictive modeling of Chicago mosquito trap data for the city's public health monitoring program -- chi-square testing, then linear and logistic regression on WNV prevalence. Repository →

U.S. Diabetes Readmission

Explored ~100k U.S. hospital records (1999–2008) to identify sub-groups at higher readmission risk, then modeled readmission with logistic regression. Repository →

Research & Public Voice

Podcast & Talks

AI safety, alignment, and what it takes for a system to actually follow a rule.

Virtuous Machines — the podcast

Conversations at the intersection of AI and philosophy with researchers in AI safety, ethics, and emerging tech. 13 episodes and counting, on YouTube and Spotify.

virtuousmachines.com → 13 episodes Ā· YouTube & Spotify

All 13 episodes →

Speaking at the Mindstone AI Meetup

Public Engagements

More talks at virtuousmachines.com/talks →
Working through logic notation on a whiteboard

Research & Writings

More at virtuousmachines.com/research-consulting →

Publications

No Bullshit Math for Data Science cover

No Bullshit Math for Data Science

A short, sharp book on the mathematics underneath data science -- linear algebra, calculus, statistics, and the math behind core ML algorithms, with Python code.

Amazon → Book PDF →

Most books teaching the mathematics under data science run ~350 pages, cost a lot, lean heavily on code, and teach surprisingly little actual math. This one is built to fix that: short, sharp, and math-first.

Part I covers the core mathematics -- linear algebra, calculus, statistics, and probability theory. Part II builds on it to cover the mathematics underneath regression, tree methods, support vector machines, clustering, PCA, and neural networks, with Python implementations throughout.

Sample page: Vectors Sample page: table of contents Sample page: matrix inverse
Structured propositions paper first page

Structured Propositions & Logics of Ground

Peer-reviewed paper in Synthese providing a unified semantics for propositional logics of impure ground using structured propositions.

Read paper →

Abstract

I show that the assumption of highly structured propositions can be leveraged to provide a unified semantics for various propositional logics of impure ground in a very expressive and flexible way. It is shown, in particular, that the induced models are capable of capturing an infinitude of grounding facts that follow from unrestricted logics of ground, but, due to certain artificial restrictions, are left unaccounted for by the existing semantics in the literature. It is also shown that our models, unlike the ones in the literature, are easily extendable to capture certain distinct views about iterated as well as identity grounding.

Practical Statistics for Data Practitioners cover

Practical Statistics for Data Practitioners

Open-source guide to descriptive and inferential statistics for real-world analytics, paired with reproducible Python code.

Repository →

An open-source, hands-on guide to the descriptive and inferential statistics needed in real analytics and ML pipelines. Each concept is paired with clear intuition, minimal proofs, and fully reproducible Python code.

Covers measures of central tendency and dispersion, hypothesis testing, A/B testing, and regression models, illustrated with end-to-end examples using pandas, NumPy, SciPy, and statsmodels.

Academic CV ↓

Education & Certificates

Philosophy studies

MA & PhD, Philosophy (Logic, Semantics & Metaphysics)

University of Calgary (2023) Ā· University of Tartu (2019)

Overview

After years of exposure to mathematical logic in my undergraduate years, I grew interested in the philosophical ideas behind math and logic and did my MA in Philosophy at the University of Tartu in Estonia.

I then continued into a PhD at the University of Calgary, where over 4 years I engaged in research in Logic and Metaphysics.

Research

My MA thesis, Perception, Abductive Methodology, and Compositional Universalism, argued for compositional universalism -- the view that any plurality of objects constitutes an object. Read the thesis →

My PhD research focused on questions of fundamentality, with a secondary interest in IT ethics, AI safety, and machine consciousness that I still work in today. Read the dissertation →

One dissertation chapter was later published in Synthese. Read the paper →

Teaching

Taught a graduate-level logic course on modern modal logics to 10+ masters and PhD students across Philosophy, Mathematics, and Semiotics during my MA.

During my PhD, taught hundreds of students on AI, IT, and data ethics -- deepening my grounding in AI safety debates. I write on these topics at substack.com/@llminds and occasionally for Wikipedia and the Centre for Social Impact Technology.

Awards

MA scholarships (over $24,000):

  • Best MA Thesis Prospectus of The Year Award, University of Tartu (2018)
  • Tuition Waiver Scholarship, University of Tartu -- CA$6,300/yr (2017 & 2018)
  • DoraPlus Scholarship, University of Tartu -- CA$5,700/yr (2017 & 2018)
  • Achievement Stipend, University of Tartu (2018)
  • Excellence Scholarship for Travel and Accommodation, University of Tartu

PhD scholarships (over $50,000):

  • Alberta Graduate Excellence Scholarship (AGES), International -- CA$15,000 (2021, 2019)
  • Department of Philosophy Graduate Essay Award -- CA$2,500 (2021)
  • Full Funding Package (4 years) -- CA$30,000/yr (2019–2023)
Amirkabir University of Technology

BSc, Mathematics & Applications

Tehran Polytechnic (Amirkabir University of Technology), 2017

Overview

Undergraduate in Mathematics and Applications at Amirkabir University of Technology in Iran. My final thesis was on mathematical logic, with broad coursework across pure and applied math.

Coursework

Statistics & Probability Theory Numerical Analysis Optimization Theory Linear & Abstract Algebra Real & Complex Analysis Differential Equations Differential & Algebraic Geometry Logic, Set Theory & Category Theory

Research

Thesis on coalgebraic logics -- Logics for Coalgebras, and a Final Coalgebra Theorem (written in Persian). Read the thesis →

BrainStation certificate

BrainStation — Data Science Diploma

3-month intensive data science bootcamp

Google Cloud Platform certificate

Google Cloud Platform

Cloud & ML infrastructure certification

Get in touch

Always glad to talk about agent systems that have to survive real traffic -- evaluation, guardrails, and the failures that only surface once something is live. Open to agentic AI and applied ML work in Toronto or remote.

amirkianitech@gmail.com LinkedIn GitHub

Ā© 2026 Amir Kiani