We Tested 5 AI App Builders With the Same Prompt. Here's What Actually Happened.
Five AI builders got the same two-word brief. The fastest was not the best experience, and the deepest platform could not publish.
Azhab NS
Five agents. One deliberately vague prompt. The differences showed up before the code did.
Part 1 of 3: The Benchmark
We wanted to understand what a genuinely better AI app-building experience should feel like. So instead of comparing feature pages, we gave five products the same simple prompt: "Create a portfolio."
Why such a simple prompt?
Because a vague prompt exposes the product's defaults.
A detailed specification can hide a weak product experience. If the user tells the system the page structure, content, visual direction, workflow and technical requirements, the user has already done much of the thinking.
"Create a portfolio" forces the builder to reveal what it does when information is missing.
Does it ask?
Does it plan?
Does it guess?
Does it build immediately?
Can the user edit the result without returning to chat?
Can it actually publish?
And what happens to credits while all of this is happening?
That was the point of the test.
Method note: This was a hands-on, experience-led benchmark, not a universal ranking. We used the same prompt across five agents and free access wherever possible. Results reflect this controlled test and the product behaviour observed during it. A simple portfolio does not prove backend depth, production-scale reliability, uptime, or heavy-concurrency performance.
The test: one prompt, eight things to watch
We tested:
- Clarification - how the agent asks for missing information.
- Planning - whether the user sees a plan before generation.
- Speed - how quickly the first usable output appears.
- UI quality - how convincing the first result looks.
- Editing - whether the user can refine the result without endless prompting.
- Publishing - whether the path to a live result actually works.
- Reliability - whether the generated experience breaks during the test.
- Cost friction - how quickly credits, limits or upgrade pressure enter the experience.
We also looked at developer depth, because a no-code experience still needs somewhere to go when the application stops being simple.
These criteria line up almost exactly with the internal FloNeo evaluation framework we had already written down: prompt clarity, plan/spec creation, UI and usability, visual editing, developer depth, editing without token loss, user control, token visibility and publishing - followed by deeper tests for reliability, accuracy, performance and predictable cost.
The competitive benchmark gave us the surface experience.
The handwritten framework told us what sits underneath it.
Emergent: the most comfortable guided experience
Observed build time: about 7 minutes to preview
Publishing: paid publishing in this test
Emergent was the most comfortable agent to use.
It did not treat the prompt as complete. It asked smart clarification questions and offered an auto-answer path, which matters more than it sounds. The user could stay involved without being forced to manually answer every small question.
Its planning behaviour also felt strong. The system appeared willing to make autonomous decisions, but it did so inside a guided flow rather than jumping straight into code.
The result was visually striking, and the experience had a testing mindset.
What worked
- smart clarification flow;
- auto-answer for questions;
- strong planning;
- autonomous judgement without making the flow feel chaotic;
- visually strong output;
- comfortable overall agent experience.
What held it back
- free publishing was blocked in the test;
- credits burned quickly;
- upgrade pressure became obvious.
This is an important combination.
Emergent scored highly on questions, planning, UI, reliability and developer depth in our experience matrix, but weakly on publishing and cost friction.
A builder can feel intelligent and still create anxiety at the last mile.
Lovable: the most balanced overall
Observed build time: about 5 minutes
Publishing: published successfully
Lovable was the most balanced experience in the test.
It asked guided questions, offered design choices and showed a formal plan before the build. That plan approval step matters because it gives the user one cheap moment to correct direction before generation begins.
The product did not feel as autonomous as Emergent, but it created a good balance of speed, control and output quality.
What worked
- guided questions;
- design choices;
- formal plan approval before generation;
- strong visual output;
- strong visual editing;
- publishing worked;
- clear, balanced flow.
What held it back
- the simple prompt produced fairly generic placeholder content;
- the test did not prove real backend depth.
That second point is important.
A portfolio is a good test of experience. It is a weak test of business logic.
The UI can look finished while deeper application behaviour remains unknown.
Lovable therefore came out strongest across the complete flow, not because the test proved it can do everything.
Base44: fastest build, strongest first-look surprise
Observed build time: about 2 minutes
Publishing: published successfully
Base44 was the fastest by a wide margin.
It also produced one of the strongest first-look interfaces in the test, including generated visual assets.
That combination is powerful: fast output plus a polished first screen.
It also had useful direct visual editing for things such as text and fonts without making the user prompt again.
But the speed came with a trade-off.
What worked
- fastest result;
- original, polished UI;
- generated images;
- useful visual edit mode;
- successful publishing;
- strong first impression.
What held it back
- no clarification questions in the initial flow;
- less user control before first generation.
Base44 is the clearest example of why speed and control are different product qualities.
The product was fast.
The user had less opportunity to shape the direction before the first build.
For simple tasks that may be acceptable. For applications where data, permissions or workflows matter, the cost of a wrong assumption increases quickly.
Kimi Websites: ambitious ideas, but reliability changed the experience
Observed build time: about 10 minutes
Publishing: published after a fix
Kimi had modern visual ideas, animation ambition, simple source files and free publishing in the test.
But the first preview was blank because GSAP failed.
The agent identified the problem, yet the repair could not finish automatically because the tool limit was reached.
That single event exposed something the other scores cannot capture:
The quality of the diagnosis does not matter if the system cannot complete the recovery.
What worked
- free publishing;
- simple source files;
- modern visual direction;
- animation ideas;
- useful multimodal potential.
What held it back
- blank first preview after a dependency failure;
- tool limit interrupted the repair;
- weaker reliability in the observed experience;
- weaker visual editing and developer depth in this specific test.
Kimi eventually published after the issue was fixed.
But the experience had already crossed from "build an app" into "repair the build system."
Replit Agent: deepest developer environment, weakest test outcome
Observed build time: about 6 minutes
Publishing: failed after three attempts and more than 30 minutes
Replit was the most technically capable environment in the group.
The generated project was a real React/Vite project with editable code. It offered checkpoints, logs, Git, database, authentication and security controls.
For a technical team, that depth matters.
For this intentionally simple task, however, the experience felt heavier than necessary.
And publishing failed.
Three attempts.
More than 30 minutes.
No live result.
What worked
- real editable code;
- deep developer environment;
- checkpoints;
- logs;
- Git;
- database;
- authentication;
- security controls;
- strong platform for technical teams.
What held it back
- publish failure after multiple attempts;
- much more complexity than the task needed;
- reliability and last-mile completion were the weakest part of the observed test.
Replit creates the cleanest contrast in the entire benchmark:
the build took about six minutes; the publishing problem took more than five times longer.
Fast generation is not the same as fast completion.
Speed and publishing: the snapshot that changed the ranking
| Agent | Observed build / preview time | Publishing outcome |
|---|---|---|
| Base44 | ~2 min | Published |
| Lovable | ~5 min | Published |
| Replit Agent | ~6 min build | Publish failed |
| Emergent | ~7 min | Paid publish |
| Kimi Websites | ~10 min | Published after fix |
If we ranked only by generation speed, Base44 wins comfortably.
If we ranked by developer depth, Replit becomes much stronger.
If we ranked by guided discovery, Emergent leads.
If we ranked by balance across the full journey, Lovable performed best in this test.
That is exactly why the category cannot be reduced to one metric.
The side-by-side experience matrix
Strong / Medium / Weak reflects this controlled test, not every possible use case.
| Agent | Questions | Planning | UI | Visual edit | Publishing | Reliability | Cost friction | Dev depth |
|---|---|---|---|---|---|---|---|---|
| Emergent | Strong | Strong | Strong | Medium | Weak | Strong | Weak | Strong |
| Lovable | Strong | Strong | Strong | Strong | Strong | Strong | Medium | Medium |
| Base44 | Weak | Medium | Strong | Strong | Strong | Strong | Medium | Medium |
| Kimi | Medium | Medium | Medium | Weak | Strong | Weak | Medium | Weak |
| Replit | Strong | Strong | Medium | Medium | Weak | Weak | Medium | Strong |
The useful takeaway is not "Product X wins."
The useful takeaway is that no single product owned every part of the journey.
What all five already do well
This benchmark also showed how far the category has moved.
Every product was operating beyond simple code completion.
Across the group, we saw:
- prompt-to-UI generation;
- responsive-layout ambition;
- placeholder content generation;
- iterative chat editing;
- preview-first workflows;
- polished visual ambition;
- deployment ambition;
- reusable code or application structure.
That means fast visual generation is becoming the baseline.
It is no longer the whole differentiation.
The next battle is:
control, trust, cost and reliable completion.
What the simple portfolio test could not prove
This is where the handwritten FloNeo framework becomes important.
The benchmark directly tested the visible experience.
It did not prove:
Reliability at production level
- uptime guarantees;
- dependency-update policy;
- failover;
- redundancy.
Accuracy at business-logic level
- prompt precision;
- workflow precision;
- backend logic precision;
- output accuracy;
- expectation / customer-delight score.
Performance at real scale
- development speed beyond the small test;
- heavy concurrent usage;
- larger data volumes;
- long-running workflows.
Cost predictability
We saw cost friction, credit pressure and tool limits.
That is not the same as proving:
- predictable total cost;
- clear visibility of cost burning over a larger project.
Those are the questions a second-stage benchmark has to answer.
What this test changed for FloNeo
The answer is not to copy one competitor.
The test showed pieces worth combining.
From Emergent: guided discovery and auto-answer.
From Lovable: visible planning and design choice.
From Base44: speed and direct visual editing.
From Replit: developer control and code visibility.
And from the failures across the test: publishing, reliability and cost transparency cannot be treated as afterthoughts.
That maps cleanly to the direction FloNeo has been working toward:
Ask -> Plan -> Build -> Refine -> Deploy
The earlier FloNeo article, How FloNeo's 4-Layer Architecture Makes AI Prototyping Ultra-Affordable, explains the architectural side: compress context, structure the application, route models intelligently and patch only what changes.
This benchmark adds the product-experience side.
Part 2 of this series goes deeper into the biggest trap the test revealed:
a beautiful AI-generated UI can still be an unfinished product.
References and Source Notes
- FloNeo internal competitive review, August 2026. Hands-on benchmark of Emergent, Lovable, Base44, Kimi Websites and Replit Agent using the same prompt: "Create a portfolio." All timings, publishing outcomes, experience notes and matrix ratings in this article come from that controlled test.
- FloNeo internal evaluation notes, 8 August 2026. Evaluation framework covering prompt clarity, plan/spec creation, UI and usability, visual editing, developer depth, editing without token loss, user control, token visibility, publishing, reliability, accuracy, performance and cost model.
- FloNeo. How FloNeo's 4-Layer Architecture Makes AI Prototyping Ultra-Affordable.