The Nine Criteria for Evaluating a Volume Bot
Nine properties separate a tool you can hold to account from one you can only hope about, and each can be checked before money moves. This page turns them into a scoring sheet.
What this page settles
- The nine criteria are custody, venue coverage, pricing, control surface, failure reporting, observability, support behaviour, contract terms and operational maturity, and each can be checked before payment.
- Every criterion scores 2, 1 or 0 against evidence you collect yourself, which is why the sheet still works on a vendor nobody has written about.
- The total ranks a shortlist rather than rating a product, the weights belong to your situation, and a zero on custody or contract terms stops the process outright.
Why a scoring sheet beats a feature list
Nine criteria decide whether a volume tool is worth paying for: custody model, venue coverage, pricing structure, control surface, failure reporting, observability, support behaviour, contract terms and operational maturity. Score each one 2, 1 or 0 against evidence you can collect before payment. The total sorts a shortlist, and a zero on custody or contract terms ends the process by itself.
A feature list is written by the party selling the software, so it names what that vendor happens to be good at. Two products with identical bullet points can differ on who holds the keys and on whether failed transactions are billed to you. Feature lists have no vocabulary for an absence, which the criteria section exists to expose.
This is a filter, not a rating. Nothing here produces a score you could publish beside a product name. The sheet shortens a list to the two or three candidates worth a small paid test, and makes each elimination one you can still explain a week later.
How to use the sheet before you pay
Scoring is reliable only when evidence is gathered in one pass and judged in another. Mixing the steps lets the first product you open set the standard for the rest. Collect first, score second, and finish one criterion across all candidates before the next. The questions to put to a vendor are the interview half of the job.
- Write the requirement first. Put down what the run must achieve before opening a vendor page, or the sheet becomes a description of whichever product you saw first.
- Collect the evidence in one pass. Save the pricing page, terms page, venue list, settings screen, any sample report and two support replies. A missing item is the score.
- Score cold, one criterion at a time. Give every candidate a C1 before anyone gets a C2, so one impression cannot carry all nine scores.
- Total and rank. Add the raw scores and apply weights if the situation calls for them. The ordering is a plan for work, not a verdict.
- Apply the veto conditions. Strike out any candidate scoring zero on C1 or C8, whatever the total says.
- Re-score after a small paid test. Buy the smallest job the vendor sells, then score C5, C6 and C7 from what the run produced.
- Written price expression, network costs included
- Terms page readable before payment
- Venues named individually, not by category
- Settings screen with its limits
- Sample run report or export
- Two support replies, with times
- A dated change log or status history
C1 to C3: the structure of the deal
The first three criteria describe the shape of the arrangement rather than the quality of the software. They are hardest to change after payment and most often left vague on purpose. Score them from documents, because a structural answer that exists only inside a chat window can be withdrawn without trace.
C1 Custody model
Custody asks one question: can the tool move your funds without you signing for it. A non-custodial design has you sign from a wallet you control. A custodial design has you deposit into an address the operator holds, which is legitimate but a different risk, since withdrawal depends on that operator cooperating.
Check it by asking who signs each transaction and by watching what onboarding requests; the mechanics are worked through in the piece on custodial and non-custodial volume tools. One answer ends the conversation. If a seed phrase or private key is requested, for any reason, the score is zero, and the loss is irreversible once a key leaves your control.
C2 Venue coverage
Coverage means the venues the engine actually routes through, not the logos on a home page. A tool can display eight names and send everything to one of them. Coverage also has a maintenance dimension: venues change, and an engine that has added nothing since launch is trading against a market that moved on.
Verify it by asking for the list as individual names, then resolving signatures from a small run on a block explorer to see which programs were touched. Refusal to name venues settles this criterion. A phrase such as "all major decentralised exchanges" is not a list, and an operator who will not commit before payment will not commit later.
C3 Pricing structure
Pricing is scored on completeness, not cheapness. The full expression covers the headline model, everything that headline excludes, the network costs you pay regardless, and any minimum on a first run. Two vendors quoting the same percentage produce different totals once failed transactions enter the calculation, which is why a bare percentage never earns a 2.
Ask for a worked total on a run of your intended size, in writing; how each structure behaves as a run grows belongs to the comparison of pricing models. What ends it is a price that appears only after payment, or a message contradicting the public page. A price that cannot survive being written down is an opening position.
C4 to C6: what you can steer and what you can see
Criteria four to six describe your relationship with the running job: whether you can shape it, whether you learn when part of it fails, and whether you can reconstruct events from something other than the operator's summary. Products often score well on the first three and poorly here, because visibility costs engineering time.
C4 Control surface
The control surface is the set of parameters you may set: wallet count, order size and spread, timing between orders, venue selection, stop conditions and a hard budget ceiling. Its absence is sometimes presented as simplicity. A start button is simple in the way a locked door is quiet; the parameters still exist, chosen for you.
Inspect the settings screen before paying, or the configuration file where the tool is self-hosted, and note which limits are enforced. A budget ceiling and a working kill switch matter more than the number of tuning knobs, since they turn a bad run into a small loss. The disqualifying reply is that parameters appear after purchase.
C5 Failure reporting
Transactions fail in every run, and the base fee of 5,000 lamports per signature is charged whether a transaction succeeds or fails. Failure reporting therefore holds two questions: does the tool show the failures, and does it bill you for them. A report listing only completed orders is a summary with the expensive half removed.
Request a sample report before payment and read a real one afterwards. What you want is a per-transaction record carrying a status and a signature, since a headline success rate reconciles against nothing. Conversation ends when an operator cannot say whether failures are billed: that is not a hard question, and the silence is the finding.
C6 Observability
Observability covers what you can watch while the job runs and what you can take away once it stops. The strong version is a live view with per-order status plus an export containing transaction signatures. The weak version is a progress bar and a completion message, which confirms the software finished without establishing what it did.
Much of this can be scored from the interface without an account: open any hosted product, including a professional Solana volume bot, and check whether the run view exposes signatures. Interface model sets the ceiling, a difference covered in the piece on Telegram bots against web consoles. One request settles it: send me the signatures.
C7 to C9: the operator behind the software
The last three criteria concern the counterparty rather than the code. They matter most in the situation you are trying to avoid: something has gone wrong, money has moved, and the only thing between you and a total loss is how this operator behaves under pressure. All three can be sampled before payment.
C7 Support behaviour
Support is measured on specificity and on the gap between pre-sale and post-sale response. Send two questions that require a real answer, such as how failed transactions are billed and what the refund clause covers, then note how long each reply takes and whether it addresses the question actually asked.
Record the channel as well as the content, because support conducted in a chat either party can delete leaves nothing to point at later. Ask which channel is the channel of record. Two deflections on one question end this criterion: a single vague reply is a busy day, a second is a policy.
C8 Contract terms
Terms are the only part of an arrangement that survives a disagreement intact. What matters is not the presence of a terms page but whether four things appear inside it: refund conditions, cancellation, data retention, and a description of what happens when a run fails or is stopped part way through.
Read the page before payment and search it for the failure case. Where terms describe delivery in general language and never mention partial execution, the likeliest bad outcome has no agreed treatment at all. Nothing further needs saying when there are no terms: every disputed question is then decided by whoever holds the funds.
C9 Operational maturity
Maturity is evidence that the operation existed before you found it and will plausibly exist after you pay. Useful signals include a dated change log, a status history covering bad days as well as good ones, disclosed incidents, and documentation revised rather than written once. None of this requires a review to verify.
Judge the record by whether it admits to anything, since a change log listing only features reads as marketing. This is the criterion where a young product should score low without being condemned, because the right response is a smaller first test. An absence of any dated artefact closes it: you are then the operational history.
The nine-point scoring sheet
Here is the sheet in full. Copy it into a spreadsheet with one column per candidate and fill it in from the evidence pack, not from memory. A 2 requires something documented, a 1 covers an answer that is partial or verbal, and a 0 covers absence, refusal or contradiction.
| Code | Criterion | Evidence that satisfies it | Score 2 (clear) | Score 1 (partial) | Score 0 (fail) |
|---|---|---|---|---|---|
| C1 | Custody model | Written answer naming who signs, plus onboarding | You sign from your own wallet; no key requested | Custodial, withdrawal path documented | Custody unclear, or a key requested |
| C2 | Venue coverage | Venues named individually, plus resolvable signatures | Named list, demonstrated, with an additions process | Named list, claimed only, no evidence | A category, and no list on request |
| C3 | Pricing structure | Full price in writing, network costs included | Every line named, worked total for your run | Headline clear, extras unquantified | Price only after payment |
| C4 | Control surface | Settings screen showing parameter limits | Count, size, spacing, venue, stop condition, budget cap | Parameters exposed, no budget ceiling | Start button only |
| C5 | Failure reporting | Sample report separating confirmed from failed | Per-transaction status with signatures, billing in writing | Failures aggregated into a rate | Completed orders only |
| C6 | Observability | Live run view and an export containing signatures | Signatures during and after, readable export | Summary figures, signatures on request | No signatures at any point |
| C7 | Support behaviour | Two pre-sale questions and the replies | Specific answers, named channel, stated response window | General answers, or a deletable chat | Deflected twice, or answers change |
| C8 | Contract terms | Terms page readable before payment | Refund, cancellation, retention, failed-run handling written | Terms silent on failed runs | No terms, or no obligation stated |
| C9 | Operational maturity | Dated change log or status history | Dated log plus a history recording bad days | Features listed, no incident or fix | No dated artefact anywhere |
Two habits keep the sheet honest. Score missing evidence as a zero rather than leaving a cell blank, because blanks vanish from a total and absences are what you are trying to detect. Write a line of justification beside each score, since a number you cannot defend is an impression wearing a numeral.
Totalling the sheet and setting your own weights
Raw totals run from zero to eighteen. Read the result as three bands rather than a league table: a candidate in the low teens or above is worth a paid test, one around the middle needs specific gaps closed first, and anything lower has failed enough criteria that the remaining questions are academic.
Weighting is optional, and the weights belong to your situation rather than to this desk. If you use them, choose them before scoring begins, because weights selected afterwards are a way of confirming the answer you already preferred.
weighted total = 3 x (C1 + C2 + C3) + 2 x (C4 + C5 + C6) + 1 x (C7 + C8 + C9) One possible weighting, shown to illustrate the shape rather than to recommend a set. The multipliers are arbitrary; pick your own and hold them fixed. Illustrative
The figures below are invented to show the arithmetic and describe no real product. Suppose a candidate scores 2 on C1, 1 on C2, 2 on C3, 2 on C4, 0 on C5, 1 on C6, then 2, 2 and 1 across the last three.
- Raw total: 2 + 1 + 2 + 2 + 0 + 1 + 2 + 2 + 1 = 13 out of 18.
- Structural set at weight 3: (2 + 1 + 2) x 3 = 15.
- Operational set at weight 2: (2 + 0 + 1) x 2 = 6.
- Counterparty set at weight 1: (2 + 2 + 1) x 1 = 5.
- Weighted total: 15 + 6 + 5 = 26 out of 36.
Both totals look respectable and both hide the same thing. The zero sits on C5, so this candidate does not report failed transactions. Neither veto is triggered, so it survives, but the next step is one question about failure billing and a re-score.
Buyers in different positions should not run identical weights. The table below offers a starting point for three situations, expressed as multipliers on the three groups rather than on individual criteria, which keeps the arithmetic small. Adjust the numbers to match your own exposure.
| Situation | C1 to C3 structural | C4 to C6 operational | C7 to C9 counterparty | What usually decides it |
|---|---|---|---|---|
| First small run, own funds | x3 | x1 | x2 | Whether a bad purchase can be contained |
| Recurring runs, larger budget | x2 | x3 | x1 | Whether each run can be read and reconciled |
| Running for someone else | x3 | x2 | x3 | Whether you can show a third party what happened |
Keep the weights out of the scoring session. Score every candidate on the raw sheet first and apply multipliers afterwards, as a separate step. Weighting while you score lets the scheme bleed into your judgement of the evidence.
The two criteria that are veto conditions
Custody and contract terms behave differently from the other seven. A zero on either removes the candidate whatever the total, because both describe conditions under which the rest of the sheet stops mattering. Everything else is a trade-off: thin coverage, a narrow control surface or a young operation can be worked around by running smaller.
Vetoes also protect the sheet from itself. Any additive score can be gamed by an operator who looks strong on six visible criteria and stays silent on the two expensive ones, and a numerical filter would rank that candidate first. Pulling the vetoes out of the arithmetic stops a high total from buying its way past a structural failure.
What the sheet cannot tell you
The sheet measures accountability rather than outcome. A candidate scoring well across all nine criteria is one you can hold to account: you know who signs, what you pay, what you can configure and what you are owed when a run fails. None of that forecasts what a run does to a chart.
Time is the other limit. Every score is a snapshot of an operator on the day you looked, and the criteria most likely to move are support behaviour, observability and maturity. Re-score those three after your first paid run, and take the whole sheet again if you return after a gap of months.
Frequently asked questions
What are the nine volume bot evaluation criteria?
They are custody model, venue coverage, pricing structure, control surface, failure reporting, observability, support behaviour, contract terms and operational maturity. Each names a property you can check before payment rather than a feature a vendor can claim about itself. The set is deliberately short, so one person can score four or five candidates in an afternoon without the standard drifting between the first and the last.
How is each criterion scored?
Award 2 where documented evidence satisfies the criterion, 1 where the answer exists but is partial or verbal only, and 0 where the evidence is absent, refused or contradicted. Nine criteria give a raw range of zero to eighteen. Read the result as a band rather than a rank, because a gap of one or two points between two candidates sits within the noise of your own judgement.
Which criteria should stop a purchase outright?
Custody and contract terms. A zero on custody means funds can move under conditions nobody has written down, and the loss is irreversible once a key is out of your hands. A zero on contract terms means there is no agreed description of what happens when a run fails. Neither is offset by wide venue coverage, a pleasant interface or an attractive headline price.
Can this sheet be used on a vendor with no reviews at all?
That is the case it is built for. Every criterion is scored from material the vendor hands you directly or from a run you pay for yourself, so an unknown operator and a widely discussed one go through the same procedure. Absence of public discussion touches operational maturity only, and that criterion is worth two points out of eighteen.
Should the nine criteria be weighted equally?
Not always. A first small run leans on custody, pricing and contract terms, because the realistic downside is a bad purchase. A recurring operation leans on control surface, failure reporting and observability, because the downside is a result nobody can read. Choose weights before scoring, write them down, and do not adjust them once the totals are visible.
What does a high total actually predict?
It predicts accountability, not profit. A high scorer is one you can hold to account: you know who signs, what you pay, what you can configure, what a failure looks like and what you are owed when one happens. The outcome of a run depends on the token, the venue, the depth available and decisions that stay yours throughout.
How often should a vendor be re-scored?
Re-score support behaviour, failure reporting and observability immediately after the first paid run, because pre-sale answers and post-payment behaviour are separate measurements. Score the whole sheet again if you return to an operator after a gap of months. Staff, pricing and terms change quietly, and the operation you evaluated may not be the one you are now paying.
Filed under Criteria by The Volume Bot Review Desk. Every figure on this page is either a protocol fact or arithmetic explicitly labelled illustrative. How we handle numbers is set out in the editorial policy.