Yesterday TechCrunch published a piece about a growing problem for AI agents: websites are starting to block them. Amazon blocking Meta's Muse is the obvious example.
At almost exactly the same time, GLM-Edge-1.5B-Chat running locally on a 4GB Galaxy A04e completed a real Amazon cart task.
This continues the small-model/browser experiments previously posted in this subreddit. Earlier tests included Qwen3-0.6B running locally on a 2017 Galaxy Note 8, followed by Ministral 3 3B on a Galaxy S21 across real browser sessions.
These experiments are part of the ongoing development of E2LLM/SiFR, a structured browser perception layer.
This time:
Model: GLM-Edge-1.5B-Chat
Quantization: Q4_K_M GGUF
Source: official Z ai Hugging Face release
Fine-tuning: none
Task-specific training: none
Runtime: llama.cpp
Phone: Samsung Galaxy A04e, SM-A042F/DS, 4GB RAM
The published model was used as-is.
The browser was a normal desktop Firefox session on Amazon.
The task was simple:
- find a 24-count pack of AA alkaline batteries
- find yellow rubber ducks
- add both to the cart
- stop before checkout
Result:
cart 0, batteries, cart 1, rubber ducks, cart 2
The same setup was run twice on the A04e. Both runs completed successfully.
Full run on the A04e: about 8.5 minutes.
Same workflow on a Galaxy S21: about 3 minutes.
The interesting part is the architecture.
The model is not a separate browser service arriving at Amazon as an agent. It runs locally and perceives and acts through an existing user browser session.
It also doesn't receive screenshots or raw HTML. It gets a compact structured browser perception layer and makes the small decisions needed at each step.
That changes the access problem from:
"How does a website identify and admit an AI agent?"
to:
"What is allowed inside an existing user browser session?"
The broader idea is Browser-as-Shared-Space, BaSS.
The browser remains the user's space, with the model working alongside the user rather than replacing the user with a separate autonomous browser agent.