What it is
Panda is a proactive, on-device AI agent for Android. It takes natural language commands and operates the phone's UI by tapping, swiping and typing to complete multi-step tasks across apps. It uses a multi-agent architecture written in Kotlin: the Android Accessibility Service reads the screen and performs gestures, and LLMs handle reasoning and planning. The project is marked work-in-progress.
Who it's for
- Android developers who want to build or experiment with LLM-driven UI automation
- Hobbyists and students using it for personal, educational or non-commercial purposes
- Contributors interested in agentic phone-operator projects
Requirements
Requirements
- Android Studio (latest version recommended)
- An Android device or emulator with API level 26+
- Gemini API keys, or any backend that accepts the documented request payload (GCLOUD_PROXY_URL)
- Accessibility permission granted to the Panda service
- API key written in local.properties
- Personal Use License; commercial use requires a separate license
Setup
Build & run
Open the project in Android Studio, let Gradle sync all dependencies, then run the app on your selected device or emulator. Write your API key in local.properties; the README says more keys give better speed.
Enable Accessibility Service
On first run the app prompts for Accessibility permission. Click "Grant Access" and enable the "Panda" service in your phone's settings. This is required for the agent to see and control the screen.
Examples
Backend proxy environment variables
python# the name of these keys donot mean you need google cloud, you can use any servers that can accept requests, i will improve the developer experience in the future by making openapi compatible
GCLOUD_PROXY_URL=<url-of-any-backend-that-accept-responses-like-below-payload>
GCLOUD_PROXY_URL_KEY=<any-password-you-wanna-set-or-leave-empty>What it does: Points the app at any backend that accepts the documented request payload, so Google Cloud is not required.
Expected request payload
json{
"modelName": "model-name",
"messages": [
{
"role": "user",
"parts": [
{
"text": "Hello, what can you do?"
}
]
},
{
"role": "model",
"parts": [
{
"text": "I can help you with a variety of tasks. What do you need assistance with today?"
}
]
}
]
}What it does: The message format your custom backend must accept.
Using Gemini keys directly
bashGEMINI_API_KEYS=What it does: Alternative to the proxy: add Gemini keys to play around.
View logs in real time
bashadb logcat | grep GeminiApiWhat it does: Filters device logs to Gemini API activity for debugging.
Pros & cons
Pros
- Pro:Operates the phone UI through the Android Accessibility Service, so it can perform multi-step tasks across different apps
- Pro:Backend is flexible: use any server accepting the documented payload, or Gemini keys directly
- Pro:Free for personal, educational and non-commercial use, including modification and distribution
Cons
- Con:Project status is work-in-progress and the roadmap is marked as not updated
- Con:Persistent memory is temporarily disabled
- Con:Commercial use requires a separate license under the Personal Use License