Project Overview
AmbientAI is an edge-AI music companion built around the PHYTEC phyBOARD-Atlas RT1170 platform and the NXP i.MX RT1176 crossover MCU. The system combines visual context, environmental sensing, local inference, and locally stored music to select a suitable playlist for the current situation.
The idea is to make music selection more adaptive without requiring the user to manually choose a playlist every time or send camera data to a cloud service. The system uses a camera-based facial-expression model together with temperature, humidity, ambient-light, gas-resistance, and optional weather information. These inputs are converted into context scores, fused locally, and used to select a tagged playlist.
The application is designed around four broad states: positive, neutral, tired, and stressed. These states are intentionally coarse so that the system remains practical for an embedded MCU and produces stable playlist decisions rather than attempting to make a complex psychological assessment.
Hardware
The primary processing platform is the PHYTEC phyBOARD-Atlas based on the NXP i.MX RT1176. I use the Cortex-M7 for the main application, including network communication, image preprocessing, FER inference, sensor fusion, playlist control, SD-card access, and audio playback. The Cortex-M4 remains available for future workload partitioning, but the current integrated application runs its main processing flow on the M7.
The environmental sensing subsystem contains:
- Bosch BME688 for temperature, relative humidity, pressure, and gas-resistance measurements
- Vishay VEML7700 for ambient-light measurement
- I2C communication between the sensors and the RT1170 carrier board
The camera is implemented as a separate ESP32 camera satellite. The ESP32 handles the camera driver, image capture, grayscale conversion, and Wi-Fi transport. It sends the captured frames over the local Wi-Fi network to the RT1170, which is connected to the same local network through Ethernet.
This separate-camera architecture provides two practical advantages. First, the ESP32 handles the camera hardware interface and driver, allowing the RT1170 to receive a defined frame format instead of directly handling the camera sensor timing. Second, the camera can be placed independently from the larger carrier board. For example, I can position the camera on a desk or at another location with a clear view of the user while keeping the RT1170 near the Ethernet connection, SD card, and audio output. This makes the system easier to use in different rooms and for different users.
I initially investigated the Raspberry Pi Camera Module v2 using the Sony IMX219 sensor and achieved substantial progress with the sensor control interface. However, the remaining native CSI path requires reliable handling of the camera's raw MIPI CSI-2 image stream. I also considered an OV5640 CSI camera module, but its cost was not practical for this prototype. The ESP32 camera satellite therefore became the selected implementation for the current system.
For local audio, I use the microSD card on the RT1170 unit to store music files and tagged playlists. Audio is played through the existing 3.5 mm headphone jack on the RT1170 carrier board. I connect wired headphones or a powered speaker directly to this jack. The current design does not require a separate external audio-codec module.
Software
I use Zephyr RTOS as the embedded software platform. Zephyr provides the device-tree configuration, RTOS scheduling, networking, I2C, SD-card, filesystem, audio, and peripheral-driver support needed for this application.
The main software components are:
- Zephyr RTOS application running on the RT1170
- ESP32 camera firmware for image capture and local Wi-Fi transport
- TensorFlow Lite Micro for embedded FER inference
- Quantized INT8 facial-expression model
- BME688 and VEML7700 sensor acquisition
- Ethernet and Wi-Fi local-network communication
- FAT filesystem and microSD music storage
- SAI audio playback through the carrier board's existing audio path
- Rule-based context fusion and playlist selection
The project was developed using the Zephyr SDK and PHYTEC board support, with WSL/Ubuntu, MCU-Link/LinkServer, and a serial terminal used during development and debugging.
AI and Visual Processing
The supplied facial-expression model produces seven classes:
- Angry
- Disgust
- Fear
- Happy
- Sad
- Surprise
- Neutral
The ESP32 camera provides a 320 × 240 grayscale frame to the RT1170. The RT1170 validates the frame, extracts the central face region used by the embedded pipeline, reduces it to a 64 × 64 grayscale image, and passes it to the quantized FER model running locally on the Cortex-M7.
The seven model outputs are mapped into four application states:
- Happy and surprise contribute to positive
- Neutral maps to neutral
- Sad maps to tired
- Angry, disgust, and fear contribute to stressed
The model output is filtered with confidence and timing rules so that the playlist does not change because of one brief or uncertain frame. The application uses smoothing, a confidence threshold, a candidate dwell period, and a minimum hold time before changing the selected playlist.
Environmental Context
The BME688 provides temperature, relative humidity, pressure, and gas-resistance data. The VEML7700 provides ambient-light data. These values are acquired over I2C and passed to the environment-scoring layer.
Temperature and humidity are used to calculate a comfort-related preference score. Ambient light contributes information about the surrounding brightness. The gas-resistance value is retained as a sensor measurement and is not presented as a calibrated CO2 or air-quality concentration. Optional weather information can contribute cloud-cover context when configured, but the camera, sensor, inference, and audio pipeline remain local to the system.
Context Fusion and Playlist Selection
The visual and environmental stages produce separate score vectors. When both are valid, the visual context has the stronger contribution and the environmental context provides a smaller preference adjustment. If environmental data is unavailable, the visual result can still be used. If the visual result is invalid or below the confidence threshold, the current playlist is retained instead of inventing a new mood.
The final state selects one of four playlists:
- Uplifting
- Focus
- Gentle motivation
- Calm
The playlists and music files are stored on the local microSD card. The audio worker reads the selected playlist, validates the WAV file format, streams the audio through DMA-safe buffers, and sends the playback data through the carrier board's onboard audio path to the 3.5 mm headphone jack.
Privacy and Local Operation
Privacy is a central design objective. The ESP32 camera module contains only the camera capture and local transport functionality. It does not run cloud analytics or send the camera feed to an external inference service.
The ESP32 connects to the local Wi-Fi network, while the RT1170 uses Ethernet on the same local network.
Camera frames are exchanged locally between the two devices. The RT1170 uses the internet only for optional services such as weather retrieval and time synchronization when that feature is enabled. The camera feed and FER inference remain local to the system.
This arrangement also allows the camera to be physically separated from the base unit while keeping the data path under the control of the local network.
Completed Work
I have completed and integrated the main application modules:
- Zephyr RT1170 application and board configuration
- ESP32 camera capture and Wi-Fi frame transport
- FER model integration using TensorFlow Lite Micro
- Local image preprocessing and facial-expression classification
- BME688 sensor acquisition
- VEML7700 ambient-light acquisition
- Local context fusion and playlist-selection logic
- microSD filesystem and playlist handling
- WAV audio streaming
- Audio output through the existing carrier-board 3.5 mm jack
- Ethernet communication between the RT1170 and the local network
- Local camera-to-base-unit data exchange
- Privacy-oriented local processing architecture
The source package contains the embedded source code, ESP32 camera firmware, model files, sensor drivers, audio and SD-card modules, test utilities, playlists, and configuration files.
I independently verified the camera transport, sensor acquisition, FER inference path, context-fusion logic, SD-card playlist handling, and audio output path during development. The only remaining hardware path is direct native CSI image capture on the RT1170 carrier board.
Remaining Development
The remaining camera-related development is native CSI capture directly on the RT1170 carrier board. The IMX219 investigation established communication with the sensor, but the complete raw MIPI CSI-2 frame-acquisition path still requires additional driver, capture, and image-format work. An OV5640-based native camera path is another possible option if the hardware becomes available.
The separate ESP32 camera remains a useful architecture even after native CSI support is developed because it allows flexible camera placement. A future version could also replace the current Wi-Fi transport with a dedicated local RF camera link so that camera frames are exchanged without depending on a routed network connection.
Future versions may also add user-specific playlist preferences, a display for status feedback, more environmental inputs, and improved model calibration. The central design remains the same: use local edge processing to combine visual and environmental context and create a responsive music experience without continuous cloud-based camera processing.
Demonstration and Supporting Files
The submission includes photographs of the RT1170 carrier board, sensor connections, ESP32 camera module, SD-card/audio setup, and development environment. It also includes the source repository and the system architecture and processing diagrams.
The source package includes:
- RT1170 Zephyr firmware
- ESP32 camera firmware
- Supplied AI source and model files
- BME688 and VEML7700 integration
- Camera frame protocol and transport code
- Context-fusion and playlist logic
- microSD and WAV playback code
- Configuration and build instructions
- Test and validation utilities
Demo Video: https://drive.google.com/drive/folders/1DVAtAM-3pC48sySoNEBxiPSq8nYHx7tI?usp=sharing
Source Code: https://drive.google.com/drive/folders/1zorzZZtVhAIaqvKWOb6jwIr2sH1IOWgD?usp=sharing
Block Diagram/Flow Chart: https://drive.google.com/drive/folders/1OLV1So9ZWN9NbbfAfGWxTNtdrP46WEht?usp=sharing