An AI receptionist that answers the workshop phone and texts
A Python server behind a Twilio number answers calls and texts, quotes from my price list, books jobs and puts people through, with the hard rules in code.
- Difficulty
- Advanced
- Parts cost
- About US$64 a month in Twilio fees, plus a computer that's always on
- Build time
- A working version first, then months of tuning on real calls
- Skills
- Python, FastAPI, Twilio, Linux services, Prompt writing
The problem
A small repair workshop I ran had one phone number, and I was the one answering it. I spent whole days fielding calls instead of fixing things. Most of them were the same handful of questions: how much for a tyre, what are your hours, can I bring it in. Every one of them meant putting a tool down and wiping my hands.
So I built a receptionist. It answers the phone and the texts, quotes from my own price list, books simple jobs straight into the job system, and only puts a call through to a person when it should. It ran on my own server for months, on real customers.
The most useful part of this guide is the lessons at the bottom, and one of them most of all: a business rule I’d only written into the AI’s instructions got skipped, and a customer drove to the workshop for a job we’d said we weren’t taking. The rules belong in the code.
How it works
The workshop number is a Twilio number (Twilio rents you phone numbers and lets your own code answer them). When someone rings, Twilio plays a greeting in an Australian voice, listens, turns what they say into text, and posts it to my server as a web request. My server answers with TwiML, Twilio’s little XML language: “say this, then listen again”, or “dial this number”. So a call is a run of turns: they talk, it goes quiet, my server thinks, Twilio speaks the answer.
The server is one Python program (FastAPI) running as a service on a computer that’s always on. Twilio has to be able to reach it from the internet, so I used Tailscale Funnel, which gives one port on the machine a public https address without opening the router. A Cloudflare Tunnel does the same job.
Plain code runs first, the AI model last. Every turn goes through simple checks before any model sees it: did they say goodbye, did they ask for a person, is there a rule that says no (like “we’re not taking new jobs”), are they halfway through a booking. Only what’s left goes to the model, along with my price list and a short set of rules: quote only what’s on the list, never diagnose a fault, never promise a finish date, never give a discount. Calls use Claude Haiku 4.5 through OpenRouter, because a caller can’t wait long. Texts aren’t in a hurry, so they ran on a local Qwen3 8B model on the same server.
Texts work a bit differently. Twilio gives up on a web request after 15 seconds, and a model can take longer than that. So the server answers Twilio straight away with an empty reply, works out the answer in the background, and sends it as a fresh text through Twilio’s API.
Two of us took turns running the workshop on a fortnightly roster, so “put me through” rings whoever is in that day, worked out from the date, with a manual override for swaps. Bookings collect a name and what they need, confirm it back, and create a job card. The phone number comes from Twilio, so it never has to ask. A call that ends with no booking, no transfer and no goodbye goes on a list for one summary email a day, and the caller gets a text a couple of minutes later inviting them to reply.
What you need
- A Twilio account and a mobile number that can do voice and SMS. Twilio: https://www.twilio.com
- A computer that’s always on, with Python 3.10 or newer. Packages:
pip install fastapi uvicorn httpx python-multipart twilio - An OpenRouter account and API key, for the model. OpenRouter is one account that gives you many AI
models: https://openrouter.ai. Model ID for calls:
anthropic/claude-haiku-4.5. - Optional: Ollama (https://ollama.com) if you want texts answered by a local model, like mine.
- A public https address for Twilio to post to: Tailscale Funnel (https://tailscale.com) or a Cloudflare Tunnel.
- Somewhere for bookings to land. Mine went into my own job system through a “create a job” web address. If you don’t have one, start by having bookings emailed to you.
- Your price list and business details written out as a plain text file.
Wiring
There are no wires. Everything happens in Twilio’s settings for your number, which point at your server:
| Twilio setting | Points at | What it’s for |
|---|---|---|
| A call comes in (webhook, POST) | https://your-public-address/voice/incoming |
the greeting |
| Call status changes (POST) | https://your-public-address/voice/status |
spotting calls that went nowhere |
| Primary handler fails | a TwiML Bin that forwards the call to a mobile | your server is down |
| A message comes in (webhook, POST) | https://your-public-address/sms/incoming |
texts |
The other two addresses, /voice/respond and /voice/no-answer, are set by the server itself, inside
the TwiML it sends back.
The fallback matters. A TwiML Bin is a bit of TwiML that Twilio hosts for you, so it still works when your server or your internet is down. Mine just forwarded the call to my mobile.
Firmware and config
There’s no firmware. It’s a settings file, a price list, one Python file and a service file.
The settings file
Secrets live here, not in the code. Mine was ~/.config/receptionist/env. Make it readable only by you
(chmod 600). Phone numbers go in the international format Twilio uses, +61 and then the number
without its leading zero:
TWILIO_ACCOUNT_SID=your-account-sid
TWILIO_AUTH_TOKEN=your-auth-token
TWILIO_FROM_NUMBER=+614XXXXXXXX
PUBLIC_URL=https://your-public-address
OPENROUTER_API_KEY=your-openrouter-key
# who takes transfers, in roster order, comma separated
STAFF_NUMBERS=+614XXXXXXXX,+614XXXXXXXX
JOBS_URL=http://localhost:3001/api/jobs/create
# 1 = not taking new jobs. Bookings are refused in code, on every path.
AT_CAPACITY=0
# 0 = stop the automatic callback texts without touching the code
CALLBACK_SMS_ENABLED=1
The price list
This is what the model quotes from. Plain text, in sections. The prices, turnaround and policies below are lines from mine; the hours are made up for the example:
OPENING HOURS:
- Monday to Friday: 9am - 5pm
- Saturday: 9am - 12pm
SERVICES:
- Tyre change: $65 fitting plus the tyre (varies by model)
- General service: $65 (1 hour labour)
- Electrical fault diagnosis: $65/hr
- Controller replacement: quote after inspection
TURNAROUND TIMES:
- Simple jobs like a tyre: same day, while you wait
- Waiting on parts: 1-4 weeks depending on availability
POLICIES:
- Free initial assessment and quotes
- 3 months warranty on parts and labour
When the workshop went to “no new jobs”, I also put a status banner at the very top of this file, so the model would stop inviting bookings in conversation. That banner is the polite part. The code is what actually stops a booking.
The capacity gate
This is the bit to copy even if you copy nothing else. One flag, one helper, and the helper is asked on every path that can reach a booking, before the booking flow runs:
AT_CAPACITY = os.environ.get("AT_CAPACITY", "0") == "1"
def capacity_blocks_booking() -> bool:
"""True when no booking may be STARTED or CONTINUED. Every code path that can
reach a booking asks this first. The prompt says it too, but the prompt is
not what stops a booking: the booking flow runs before the model is asked."""
return AT_CAPACITY
# in the SMS handler, BEFORE the booking flow, and covering a half-finished booking too:
if phone in bookings or wants_booking(said):
if capacity_blocks_booking():
bookings.pop(phone, None)
return CAPACITY_SMS
On mine it sat in three places: the text handler, the start of a booking on a call, and the next step of a booking already under way on a call. Search your code for everything that can create a job, and put it in front of each one.
The server
This is the shape of what ran on my server, tidied up for reuse and cut down to the parts that matter. It
answers calls and texts, transfers by roster, runs a simple booking flow, and sends the callback text.
Save it as ~/receptionist/receptionist.py, with your price list beside it as knowledge_base.txt.
#!/usr/bin/env python3
"""receptionist.py: an AI phone and SMS receptionist on Twilio.
Calls: Twilio speaks and listens (<Gather>), posts what the caller said here,
and we answer with TwiML (Twilio's XML). Texts: we answer Twilio straight away
with an empty <Response/>, then send the reply through Twilio's REST API.
Order matters. Plain code runs first: goodbye, "talk to a human", business
rules, a booking already under way. The model only sees what's left.
uvicorn receptionist:app --host 127.0.0.1 --port 8500
"""
import asyncio, json, os, re, time
from datetime import datetime, timedelta
from pathlib import Path
from xml.sax.saxutils import escape
import httpx
from fastapi import FastAPI, HTTPException, Request, Response
from twilio.request_validator import RequestValidator
# ---------- settings (secrets come from the environment, never the code) ----------
TWILIO_SID = os.environ["TWILIO_ACCOUNT_SID"]
TWILIO_TOKEN = os.environ["TWILIO_AUTH_TOKEN"]
TWILIO_FROM = os.environ["TWILIO_FROM_NUMBER"] # your Twilio number
PUBLIC_URL = os.environ["PUBLIC_URL"].rstrip("/") # the https address Twilio posts to
OPENROUTER_KEY = os.environ["OPENROUTER_API_KEY"]
VOICE_MODEL = os.environ.get("VOICE_MODEL", "anthropic/claude-haiku-4.5")
SMS_MODEL = os.environ.get("SMS_MODEL", VOICE_MODEL)
STAFF = os.environ["STAFF_NUMBERS"].split(",") # who takes transfers, in roster order
JOBS_URL = os.environ.get("JOBS_URL", "") # your job system's "create a job" endpoint
VALIDATE = os.environ.get("TWILIO_VALIDATE", "1") == "1" # 0 only for local tests
VOICE, LANG = "Google.en-AU-Wavenet-C", "en-AU" # an Australian voice, no extra charge
HERE = Path(__file__).resolve().parent
PRICE_LIST = (HERE / "knowledge_base.txt").read_text()
# ---------- the trading state: ONE flag, enforced in code ----------
AT_CAPACITY = os.environ.get("AT_CAPACITY", "0") == "1"
CAPACITY_VOICE = ("We've reached capacity and we're not taking new jobs right now, sorry. If we've "
"already got something of yours, ask to talk to a human and I'll put you through.")
CAPACITY_SMS = ("We've reached capacity and we're not taking new jobs right now, sorry. If we've "
"already got something of yours, reply HUMAN and someone will get back to you.")
def capacity_blocks_booking() -> bool:
"""True when no booking may be STARTED or CONTINUED. Every code path that can
reach a booking asks this first. The prompt says it too, but the prompt is
not what stops a booking: the booking flow runs before the model is asked."""
return AT_CAPACITY
# ---------- opening hours and who's on ----------
HOURS = {0: (9, 17), 1: (9, 17), 2: (9, 17), 3: (9, 17), 4: (9, 17), 5: (9, 12)} # weekday: (open, close)
ROSTER_START = datetime(2026, 1, 1) # day 0 of a 14-day cycle: first person 7 days, second 7 days
def on_duty(day=None):
"""The number to transfer to today, or None when nobody's in."""
day = day or datetime.now()
if day.weekday() not in HOURS:
return None
cycle = (day.date() - ROSTER_START.date()).days % 14
return STAFF[0] if cycle < 7 or len(STAFF) == 1 else STAFF[1]
def fmt(h: float) -> str:
hh, mm = int(h), round((h % 1) * 60)
return f"{hh % 12 or 12}{':%02d' % mm if mm else ''}{'am' if hh < 12 else 'pm'}"
def drop_off_line(now=None) -> str:
"""When can they actually bring it in? Checks the clock and the day."""
now = now or datetime.now()
hour_now = now.hour + now.minute / 60
for add in range(8):
hours = HOURS.get((now + timedelta(days=add)).weekday())
if not hours or (add == 0 and hour_now >= hours[1]):
continue
opens, closes = hours
if add == 0 and hour_now >= opens:
return f"We're open until {fmt(closes)} today, so drop it in any time before then."
when = "today" if add == 0 else "tomorrow" if add == 1 else (now + timedelta(days=add)).strftime("%A")
return f"We open {when} at {fmt(opens)}, drop it in any time after that."
return "Give us a call before you come down."
# ---------- plain-code checks: these run before any model ----------
# These are regexes over ORDINARY SENTENCES. "Can we book it in?" fires \bbook\b.
BOOKING_PATTERNS = [r"\bbook\b", r"\bbooking\b", r"\bbring it in\b", r"\bdrop it off\b", r"\bdrop off\b"]
HUMAN_PHRASES = ["talk to a human", "speak to a human", "talk to someone", "speak to someone",
"talk to a person", "real person", "can i speak", "can i talk", "want to speak",
"want to talk", "like to speak", "like to talk", "put me through", "transfer me"]
GOODBYES = {"no thanks", "no thank you", "thats all", "that's all", "all good", "nothing else",
"nah thanks", "nah cheers", "cheers", "bye", "goodbye", "see ya"}
YES = re.compile(r"\b(yes|yeah|yep|sure|correct|sounds good)\b")
NAME_PREFIX = re.compile(r"^(my name is|my name's|it's|its|i'm|im|this is)\s+", re.I)
def wants_booking(said: str) -> bool:
return any(re.search(p, said.lower()) for p in BOOKING_PATTERNS)
def wants_human(low: str) -> bool:
return low in ("human", "human please") or any(p in low for p in HUMAN_PHRASES)
# ---------- the booking flow: name -> work -> confirm -> job ----------
bookings = {} # CallSid (voice) or phone (SMS) -> {"state": ..., "name": ..., "work": ...}
def looks_like_a_name(s: str) -> bool:
words = s.split()
return 1 <= len(words) <= 3 and all(w.replace("-", "").replace("'", "").isalpha() for w in words)
def start_booking(key: str) -> str:
bookings[key] = {"state": "name"}
return "No worries, I can book that in. What name should I put it under?"
async def continue_booking(key: str, said: str, caller: str, source: str):
"""Returns (reply, booked). The phone number is the caller's, from Twilio. Never ask for it."""
b = bookings[key]
if b["state"] == "name":
name = NAME_PREFIX.sub("", said.strip()).rstrip(".!")
if not looks_like_a_name(name):
return "Sorry, just your name is fine. What should I put the job under?", False
b.update(name=name.title(), state="work")
return f"Thanks {b['name']}. What do you need done?", False
if b["state"] == "work":
b.update(work=said.strip(), state="confirm")
return f"Just to confirm, you need: {b['work']}. Sound good?", False
bookings.pop(key, None) # state == "confirm"
if not YES.search(said.lower()):
return "No problem, I haven't booked anything. What else can I help with?", False
await create_job(b["name"], caller, b["work"], source)
return f"All booked in, {b['name']}. {drop_off_line()}", True
async def create_job(name: str, phone: str, work: str, source: str):
if not JOBS_URL:
print("JOB (no JOBS_URL set):", name, phone, work, source)
return
async with httpx.AsyncClient(timeout=10) as c:
r = await c.post(JOBS_URL, json={"customer_name": name, "customer_phone": phone,
"description": work, "source": source})
r.raise_for_status()
# ---------- the model ----------
def system_prompt(channel: str) -> str:
"""ONE prompt builder for both channels, so voice and SMS can't drift apart."""
style = ("This is a phone call. Answer in one or two short sentences." if channel == "voice" else
"This is a text message. No emojis. Two or three sentences at most.")
status = ("CURRENT STATUS: at capacity, not taking any new jobs. Do not offer to book anyone in. "
"You can still quote prices if asked." if AT_CAPACITY else "")
return (f"You are the AI receptionist for a small repair workshop. {style}\n"
"Only quote prices that are in the price list below. Never diagnose a fault, never promise "
"a finish date, never give a discount. Don't tell people how to reach a human; that's "
f"handled for you.\n{status}\n\nPRICE LIST AND DETAILS:\n{PRICE_LIST}")
async def ask_model(messages: list, model: str, max_tokens: int, temperature: float) -> str:
async with httpx.AsyncClient(timeout=15) as c:
r = await c.post("https://openrouter.ai/api/v1/chat/completions",
headers={"Authorization": f"Bearer {OPENROUTER_KEY}"},
json={"model": model, "messages": messages,
"max_tokens": max_tokens, "temperature": temperature})
if r.status_code != 200:
print("model error", r.status_code, r.text[:500]) # the body says WHY. Log it.
raise RuntimeError(f"model returned {r.status_code}")
return r.json()["choices"][0]["message"]["content"].strip()
# ---------- Twilio plumbing ----------
app = FastAPI()
validator = RequestValidator(TWILIO_TOKEN)
async def twilio_form(request: Request) -> dict:
"""The posted form, but only if Twilio really sent it (it signs every request)."""
form = dict(await request.form())
url = PUBLIC_URL + request.url.path # the address Twilio used, not the local one
if VALIDATE and not validator.validate(url, form, request.headers.get("X-Twilio-Signature", "")):
raise HTTPException(status_code=403)
return form
def twiml(xml: str) -> Response:
return Response(content=xml, media_type="text/xml")
def say(words: str) -> str:
return f'<Say voice="{VOICE}" language="{LANG}">{escape(words)}</Say>'
def gather(words: str) -> str:
"""Say something, then listen. Twilio posts back even on silence, so /voice/respond
counts the silences and gives up after three."""
return (f'<Response><Gather input="speech" action="{PUBLIC_URL}/voice/respond" method="POST" '
f'speechTimeout="3" timeout="10" actionOnEmptyResult="true" language="{LANG}">'
f'{say(words)}</Gather></Response>')
def hang_up_after(words: str) -> str:
return f"<Response>{say(words)}</Response>"
async def send_sms(to: str, body: str):
async with httpx.AsyncClient(timeout=10) as c:
r = await c.post(f"https://api.twilio.com/2010-04-01/Accounts/{TWILIO_SID}/Messages.json",
auth=(TWILIO_SID, TWILIO_TOKEN), data={"To": to, "From": TWILIO_FROM, "Body": body})
if r.status_code not in (200, 201):
print("SMS send failed", r.status_code, r.text[:300])
def log(phone: str, role: str, words: str, channel: str):
with (HERE / "conversation_log.jsonl").open("a") as f:
f.write(json.dumps({"ts": datetime.now().isoformat(timespec="seconds"), "phone": phone,
"role": role, "content": words, "channel": channel}) + "\n")
# ---------- voice ----------
calls = {} # CallSid -> {"caller", "history", "outcome", "silences"}
@app.post("/voice/incoming")
async def voice_incoming(request: Request):
f = await twilio_form(request)
calls[f["CallSid"]] = {"caller": f.get("From", ""), "history": [], "outcome": None, "silences": 0}
log(f.get("From", ""), "system", "INCOMING CALL", "voice")
# Keep the greeting SHORT. Anything they need to know goes in the answer to what they asked.
greeting = (CAPACITY_VOICE if AT_CAPACITY else
"Thanks for calling. You're speaking with our AI assistant. What can I help you with?")
return twiml(gather(greeting))
@app.post("/voice/respond")
async def voice_respond(request: Request):
f = await twilio_form(request)
sid = f["CallSid"]
call = calls.setdefault(sid, {"caller": f.get("From", ""), "history": [], "outcome": None, "silences": 0})
said = f.get("SpeechResult", "").strip()
low = said.lower().rstrip(".!")
if not said: # silence: ask again, but not forever
call["silences"] += 1
if call["silences"] >= 3:
return twiml(hang_up_after("Sorry, I can't hear you. Text us on this number and we'll get back to you."))
return twiml(gather("Sorry, I didn't catch that. Could you say that again?"))
call["silences"] = 0
log(call["caller"], "user", said, "voice")
# 1. Plain code first.
if low in GOODBYES:
call["outcome"] = "goodbye"
return twiml(hang_up_after("No worries. Thanks for calling, have a good one."))
if wants_human(low):
number = on_duty()
if not number:
return twiml(gather("There's nobody in today, sorry. I can help with prices or our hours. Anything else?"))
call["outcome"] = "transferred"
return twiml(f'<Response>{say("No worries, putting you through now. Just hold on a sec.")}'
f'<Dial timeout="30" action="{PUBLIC_URL}/voice/no-answer">'
f'<Number>{escape(number)}</Number></Dial></Response>')
# 2. The hard rule, on BOTH booking paths: one already under way, and a new one.
if sid in bookings or wants_booking(said):
if capacity_blocks_booking():
bookings.pop(sid, None)
log(call["caller"], "system", "BOOKING REFUSED - AT CAPACITY", "voice")
return twiml(gather(CAPACITY_VOICE + " Is there anything else I can help with?"))
if sid not in bookings:
return twiml(gather(start_booking(sid)))
reply, booked = await continue_booking(sid, said, call["caller"], "ai-phone")
if booked:
call["outcome"] = "booked"
log(call["caller"], "system", "BOOKING CREATED", "voice")
return twiml(gather(reply))
# 3. Only now, the model. Hold a LOCAL reference to the history: the call can end
# (and /voice/status clears it) while we're waiting on the model.
history = call["history"]
history.append({"role": "user", "content": said})
try:
answer = await ask_model([{"role": "system", "content": system_prompt("voice")}] + history[-10:],
VOICE_MODEL, max_tokens=80, temperature=0.7)
except Exception:
answer = "Sorry, I'm having a bit of trouble. Text us on this number and we'll get back to you."
history.append({"role": "assistant", "content": answer})
log(call["caller"], "assistant", answer, "voice")
return twiml(gather(answer + " ... Is there anything else I can help with?"))
@app.post("/voice/no-answer")
async def voice_no_answer(request: Request):
f = await twilio_form(request)
if f.get("DialCallStatus") in ("no-answer", "busy", "failed"):
return twiml(hang_up_after("Sorry, nobody can get to the phone right now. Text this number with "
"your name and what you need, and we'll get back to you."))
return twiml("<Response/>")
@app.post("/voice/status")
async def voice_status(request: Request):
f = await twilio_form(request)
if f.get("CallStatus") in ("completed", "failed", "canceled", "no-answer", "busy"):
call = calls.pop(f["CallSid"], None)
bookings.pop(f["CallSid"], None)
if call and call["outcome"] is None: # no booking, no transfer, no goodbye
with (HERE / "missed_followups.jsonl").open("a") as q: # one summary email a day reads this
q.write(json.dumps({"caller": call["caller"], "ts": time.time(),
"last": call["history"][-4:]}) + "\n")
if should_send_callback(call["caller"]):
asyncio.create_task(send_callback(call["caller"]))
return Response("OK")
# ---------- the callback text, for calls that went nowhere ----------
CALLBACK_ON = os.environ.get("CALLBACK_SMS_ENABLED", "1") == "1" # kill switch, no code change
COOLDOWN_H, DAILY_CAP, DELAY_S = 48, 25, 120
callback_sent, callbacks_today = {}, {}
def is_au_mobile(phone: str) -> bool:
p = phone.replace(" ", "")
if p.startswith("+61"):
p = "0" + p[3:]
return re.fullmatch(r"04\d{8}", p) is not None
def should_send_callback(phone: str) -> bool:
if not CALLBACK_ON or not is_au_mobile(phone) or phone in STAFF:
return False
last = callback_sent.get(phone)
if last and time.time() - last < COOLDOWN_H * 3600:
return False
return callbacks_today.get(datetime.now().date().isoformat(), 0) < DAILY_CAP
def callback_text() -> str:
"""Kept under 160 characters, so it's billed as one SMS. Follows the trading flag."""
if AT_CAPACITY:
return ("Hi, you called us but we didn't get to finish. We're not taking new jobs right now, "
"sorry. Got a job with us already? Reply HUMAN.")
return ("Hi, you called us but we didn't get to finish. Reply to this text with what you need, "
"happy to quote or book you in.")
async def send_callback(phone: str):
await asyncio.sleep(DELAY_S) # don't text them while they're redialling
if not should_send_callback(phone): # check again: they may have called back
return
await send_sms(phone, callback_text())
today = datetime.now().date().isoformat()
callback_sent[phone] = time.time()
callbacks_today[today] = callbacks_today.get(today, 0) + 1
log(phone, "system", "AUTO CALLBACK SMS SENT", "voice")
# ---------- SMS ----------
texts = {} # phone -> recent messages
@app.post("/sms/incoming")
async def sms_incoming(request: Request):
f = await twilio_form(request)
asyncio.create_task(handle_text(f.get("From", ""), f.get("Body", "").strip()))
return twiml("<Response/>") # answer Twilio now: it gives up on a webhook after 15 seconds
async def handle_text(phone: str, said: str):
try:
reply = await text_reply(phone, said)
except Exception as e:
print("SMS error", phone, repr(e))
reply = "Sorry, something went wrong at our end. Please try again shortly."
log(phone, "assistant", reply, "sms")
await send_sms(phone, reply)
async def text_reply(phone: str, said: str) -> str:
low = said.lower().rstrip(".!")
log(phone, "user", said, "sms")
if not said:
return CAPACITY_SMS if AT_CAPACITY else "Thanks for your message. How can we help?"
if low in GOODBYES:
return "No worries, thanks for getting in touch. Have a good one."
if wants_human(low):
log(phone, "system", "HUMAN HANDOFF REQUESTED", "sms") # your monitor page shows these
return "No worries, someone from the workshop will get back to you as soon as they can."
# The hard rule sits BEFORE the booking flow, and covers a half-finished booking too.
if phone in bookings or wants_booking(said):
if capacity_blocks_booking():
bookings.pop(phone, None)
log(phone, "system", "BOOKING REFUSED - AT CAPACITY", "sms")
return CAPACITY_SMS
if phone not in bookings:
return start_booking(phone)
reply, booked = await continue_booking(phone, said, phone, "ai-sms")
if booked:
log(phone, "system", "BOOKING CREATED", "sms")
return reply
history = texts.setdefault(phone, [])
history.append({"role": "user", "content": said})
answer = await ask_model([{"role": "system", "content": system_prompt("sms")}] + history[-10:],
SMS_MODEL, max_tokens=300, temperature=0.4)
history.append({"role": "assistant", "content": answer})
return answer
A few things in there that matter more than they look:
- It checks every request really came from Twilio. Twilio signs each one with your auth token, and
twilio_form()refuses anything that doesn’t match. The check usesPUBLIC_URL, because the address Twilio signed is the public one, not the local one the server sees behind the tunnel. - One prompt builder for both channels. On mine, the calls and the texts each had their own copy of the business description, plus the price list file. Change one and forget another, and the phone and the texts tell people different things.
- The booking flow checks the clock before it says “drop it in”. Mine didn’t. See the lessons.
- Voice answers are capped at 80 tokens (roughly 60 words). Long answers are slow to say, and slow to listen to.
The service
So it starts on boot and restarts if it falls over. Save as /etc/systemd/system/receptionist.service
and change YOUR_USER to your login. If you installed the packages in a virtual environment, use that
environment’s python3 in ExecStart:
[Unit]
Description=AI phone and SMS receptionist
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=YOUR_USER
WorkingDirectory=/home/YOUR_USER/receptionist
EnvironmentFile=/home/YOUR_USER/.config/receptionist/env
ExecStart=/usr/bin/python3 -m uvicorn receptionist:app --host 127.0.0.1 --port 8500
Restart=always
RestartSec=5s
[Install]
WantedBy=multi-user.target
What it costs
My Twilio bill for one 30-day stretch was about US$64: SMS $44.34, speech recognition $9.00, incoming
calls $9.54, incoming texts $0.83. Texts were about 70% of it. The Australian voice
(Google.en-AU-Wavenet-C) didn’t show up as a separate charge. The model for calls costs cents per call,
and the local model for texts costs nothing per message.
Step by step
- Buy a Twilio number that does voice and SMS.
- Write your price list into
knowledge_base.txt. Hours, prices, turnaround, policies. - Make the settings file. Leave
AT_CAPACITY=0. - Copy
receptionist.pyto~/receptionist/and install the packages. - Install the service:
sudo systemctl daemon-reload, thensudo systemctl enable --now receptionist. - Give it a public address. With Tailscale:
sudo tailscale funnel --bg 8500. - In Twilio, point the number’s call, call status and message webhooks at your address (the table in Wiring), and set the fallback TwiML Bin.
- Ring it and text it from your own mobile. Read
conversation_log.jsonlafterwards. - Run the replay test below, with the flag on and off.
- Put it on the real number when you trust it, and read the log every day for the first couple of weeks.
Testing it
- Replay the messages that broke it. When something goes wrong, take the real messages from the log,
turn them into a test, and assert the thing that shouldn’t happen happens zero times. Then flip the
flag off and assert it still works, so switching back on later isn’t quietly broken. Here’s the test I’d
start with, reworded from the real incident. Save it beside
receptionist.py,pip install pytest, and run it with the settings loaded:set -a; . ~/.config/receptionist/env; set +a; TWILIO_VALIDATE=0 python3 -m pytest -q test_capacity.py
# test_capacity.py: replay the messages that caused the incident, with the flag on and off.
import asyncio
import receptionist as r
jobs = []
async def fake_create_job(*args):
jobs.append(args)
async def fake_model(*args, **kwargs):
return "(model reply)"
async def fake_sms(*args):
pass
r.create_job, r.ask_model, r.send_sms = fake_create_job, fake_model, fake_sms
# Reworded from the real texts. Write yours from your own log.
INCIDENT = ["Can we book it in for a service please, it won't turn on",
"Jo Example", "It won't turn on at all", "yes"]
def replay(messages, phone="test-caller"):
r.bookings.clear()
jobs.clear()
return [asyncio.run(r.text_reply(phone, m)) for m in messages]
def test_at_capacity_nothing_gets_booked():
r.AT_CAPACITY = True
replies = replay(INCIDENT)
assert jobs == []
assert "not taking new jobs" in replies[0]
def test_a_booking_already_under_way_is_stopped_too():
r.AT_CAPACITY = False
replay(INCIDENT[:2]) # half-way through a booking...
r.AT_CAPACITY = True # ...when the flag goes on
rest = [asyncio.run(r.text_reply("test-caller", m)) for m in INCIDENT[2:]]
assert jobs == []
assert "not taking new jobs" in rest[0]
def test_with_the_flag_off_bookings_still_work():
r.AT_CAPACITY = False
replies = replay(INCIDENT)
assert len(jobs) == 1
assert replies[-1].startswith("All booked in, Jo Example.")
def test_a_sentence_is_not_a_name():
r.AT_CAPACITY = False
replies = replay(["I'd like to book in", "Just got to the workshop and nobody is here"])
assert "just your name" in replies[1]
- Count outcomes, not just errors. For every call, write down how it ended: booked, transferred, goodbye, or nothing. The “nothing” pile is where the real problems were hiding (see the lessons).
- Watch it live. Mine had a monitor page that showed conversations as they happened, with badges for bookings and handoffs. You don’t need that on day one, but you do need to read what it’s saying to people.
- Test the fallback. Mine forwarded to my mobile if the server was down. I never actually drilled it. Do: stop the service and ring the number.
Lessons learnt
A rule that only lives in the prompt will get skipped. When the workshop got too far behind, I switched the receptionist to “at capacity, no new jobs” by writing it into the AI’s instructions for calls and for texts. Three days later a customer texted asking if they could book it in for a service because it wouldn’t turn on. The word “book” matched the booking check, and the booking check sends the message straight to the booking flow, before any model sees it. So the instructions never came into it. The customer was told “all booked in, drop it off anytime during business hours”, turned up eleven minutes before we opened, found it shut, and emailed me. The fix was the gate above, in three places, and a test that replays their actual messages and expects zero bookings. If a rule has to hold, trace every path from the webhook inward and find each one that reaches the action without asking the model. Gate it there.
“It only triggers on the command word” usually isn’t true. I thought of the booking check as a command. It was a regex over ordinary sentences, so normal sentences fired it. Look at what actually matches, not what you called the feature.
Every way it talks to customers needs the same rule. After the gate went in, the automatic callback text was still saying “happy to quote or book you in”. So a caller heard “we’re at capacity” and two minutes later got a text inviting a booking. 26 people got that in four days, and 12 replied, several of them asking for a person. Then I went through every customer-facing sentence in the code and found two more: the closed-day “talk to a human” reply still said they could reply “book”, and a note to the model on closed days still said it could help them book, contradicting the rule in the same prompt. When you flip a flag like this, list every message the system can send first. Don’t find them one complaint at a time.
The long greeting lost more calls than anything else. Over four months, 742 calls: 71% ended in a transfer to a person, only 6% in a booking, and 172 ended with the caller never understood at all. 63 of those stayed on the line for over 20 seconds, and 78 hung up between 5 and 20 seconds in. The greeting was 77 words, about 30 seconds of talking before anyone could say a word. I cut it to 16 words and moved the “here’s what we’re taking on” message into the answer, where it matches what they asked. I haven’t measured whether the never-heard calls dropped after that. If you build this, keep the greeting to one breath.
My log lied by leaving things out. The call log showed one AI turn for most of those never-heard calls, so it looked like it never asked twice. It did. The “are you still there?” retry was inside the same TwiML, so Twilio played it without ever calling my server, and nothing got logged. Know what Twilio does on its own before you trust your own log.
Speed wasn’t the problem I thought it was. It took a median 2.0 seconds to answer (3.0 seconds at the slow end), where the streaming voice services quote 0.4 to 0.9. That’s because each turn waits for silence, then the whole answer, then the whole voice. I looked hard at rebuilding it as streaming and dropped it: the local voice software I’d have used had no Australian voice, the Twilio voice I had cost nothing extra, and it would have saved about $9 a month. When 71% of calls want a person and 6% book, a faster switchboard is still a switchboard. Fix the conversation first.
Don’t read phone numbers out. Text them. Callers kept asking it to say a number slower, or saying they didn’t have a pen. Those were people trying to reach me and failing. Send numbers, addresses and prices by text.
Follow-up that depends on me doesn’t happen. It sent me an alert about every call that went nowhere, 23 of them in one month, and I was on the tools and rarely followed them up. So now the caller gets the text instead, two minutes after hanging up, with guards: Australian mobiles only, once per 48 hours per number (one number dead-ended eight times), at most 25 a day, never to staff, an off switch in the settings, and under 160 characters so it’s billed as one text. My first draft was 176 characters, which is two.
The call can end while the model is still thinking. Twilio sometimes sends “call ended” in the same
second as the caller’s last words. My status handler tidied that call’s history away while the model was
still answering, the next line crashed looking for it, and the server sent back an error, which killed
the live call. Four calls in fourteen days. Hold a local reference to anything you need after an await.
It failed quietly for two days. The local model started rejecting the text replies, and the code only logged “400 Bad Request”. So every texter got “sorry, I’m having a bit of trouble” and I had no idea why. On one day that was 7 messages out of 8. The model’s error body says what’s wrong. Log it.
Two booking bugs I never got to fix on mine. The name step took anything, so “Just arrived at
workshop and no one here.” became a customer name on a real job card. And the confirmation said “drop it
off anytime during business hours” without checking the clock or the day, which is how a customer ended
up at a shut door before opening. Capacity mode switched bookings off before I fixed either, so the
looks_like_a_name() and drop_off_line() in the code above are what I’d put in, not something that ran
on real customers.
Small ones. Write numbers the way you want them said: the model turned “two week to two month” into “two to two month”, and “anywhere from 2 weeks up to 2 months” came out right. And if your local model is a “thinking” model, turn the thinking off for this. With it on, a reply took over two minutes.