Driving a UBTECH AlphaMini robot with an LLM agent
alphamini-demo is a short experiment in putting an LLM agent in charge of a UBTECH AlphaMini, a small humanoid robot with a speaker, microphone and a library of built-in motions. It is a single-commit Python project, and it is meant as a demo rather than a product.
Talking to the robot
robot.py wraps the AlphaMini Python SDK. It finds the robot on the local network by a suffix of its serial number, connects, and exposes three async helpers: speak text with on-board TTS, play a named action, and stop all actions.
The SDK's WebSocket client checks a closed attribute that newer versions of the websockets library no longer have. Rather than pin an old dependency, the demo patches the client's alive property at import time to fall back to close_code when closed is missing.
The repo also vendors the SDK's sample scripts (face detection, touch, infrared, posture, speech recognition) under mini_demo/ for reference.
The agent
llm.py defines a single agent with the OpenAI Agents SDK, using a small GPT-4.1 model and a short system prompt to keep answers brief. Conversation history is stored in a SQLite session, and token usage for each run is written to the same database.
The agent has one tool, robot_actions. It maps about thirty readable names (hello, bow, hug, handshake, sit_down, turn_left, laugh, and so on) to the robot's internal action IDs. The model picks an action that fits the moment, and the tool plays it.
Three ways to drive it
main.pyis a text REPL: type a request, the agent answers, the robot says the answer aloud and may act on it.server.pyis a Flask app with aPOST /ask-evalendpoint that takes a message and a user name. The robot first announces who asked what, then speaks the agent's reply. Because Flask is synchronous and the SDK is async, requests are scheduled onto an asyncio loop running in a background thread.speech_recognition.pyis an early test of the robot's own speech recognition observer, which would let it listen without the REPL.
What we took from it
Off-the-shelf consumer robots already have good motion libraries. Exposing those motions as one enumerated tool was enough for a general model to pick sensible gestures, which is the same pattern we use for device-side MCP tools on our ESP32 hardware.
