Sales

Free lead scoring model builder with computed weights

By Charles Summers · Updated · Free, no signup

Short answer

This builds a point-based lead scoring model where the weights are computed rather than asserted: your ICP attributes and observable signals are classified, given a raw priority, then normalised so fit sums to 50 points and behaviour sums to 50 separately. The threshold is derived from your actual sales capacity, reps multiplied by leads per rep per week, expressed as a share of lead volume and read off a right-skewed score curve, so qualified volume matches what your team can work. You also get negative scoring large enough to outvote engagement, a decay rule with a stated half-life, and a four-box routing matrix that keeps fit and behaviour apart.

Use the lead scoring model builder

What does this tool actually do?

This builds a point-based lead scoring model where the weights are computed rather than asserted: your ICP attributes and observable signals are classified, given a raw priority, then normalised so fit sums to 50 points and behaviour sums to 50 separately. The threshold is derived from your actual sales capacity, reps multiplied by leads per rep per week, expressed as a share of lead volume and read off a right-skewed score curve, so qualified volume matches what your team can work.

It runs entirely in your browser. Nothing you type is sent to a server, no account is required, and there is no usage limit, because there is no cost per run to control.

One blended score collapses two things that require opposite responses

Add fit and behaviour together and a score of sixty has at least two meanings. It can be a company that matches your ideal profile exactly and has read one blog post, or a company you cannot sell to at all which has spent an hour on your pricing page. Those are not similar leads with similar handling. The first is a targeting success and a timing problem, and it belongs in a nurture track with an outbound touch when a trigger appears. The second is a targeting failure and belongs anywhere except a rep's calendar, because a rep working it will lose the time and then, worse, stop trusting the score for the next hundred leads.

The reason addition fails here is that it assumes substitutability: that enough behaviour can compensate for the wrong company. It cannot, because fit is a constraint and behaviour is a state. Constraints do not trade off against states. A student researching your category for a dissertation can generate more engagement points in an afternoon than a genuine buyer generates in a month, and a blended model will rank them accordingly, which is precisely the outcome that convinces sales floors that lead scoring does not work.

Keeping the two axes separate costs nothing and produces four boxes, each with an owner. High fit and high behaviour goes to a rep now, and speed matters more than polish. High fit and low behaviour stays with marketing and outbound: the account is right, so the job is to be present when something changes. Low fit and high behaviour goes to self-serve, community or documentation, and if that box is unexpectedly large you have learned something valuable and expensive, which is that your content is attracting an audience you cannot sell to. Low fit and low behaviour is suppressed, and suppression is a real action rather than an absence of one.

This also gives you a diagnostic no blended model can produce. Track conversion rate by box, not overall. If the top box converts no better than the second, your behaviour signals carry no information and are decoration. If the second converts nearly as well as the first, your fit criteria are doing all the work and you could route on firmographics alone, which would be simpler and faster.

The threshold is an arithmetic consequence of capacity

Most qualification thresholds are set in a meeting and are therefore a preference. The honest version is arithmetic with two terms. Capacity is the number of reps multiplied by the number of leads one rep can genuinely work in a week, multiplied by about 4.33 weeks in a month. Volume is what arrives. The share you can qualify is the first divided by the second, and the threshold is whatever score produces that share. Nothing about lead quality enters the calculation, which is the point: a threshold is a rationing decision and rationing is set by the size of the queue and the size of the team.

The calibration procedure takes ten minutes and beats any modelled curve, including the one this tool uses to give you a starting number. Score last month's leads with your weights, sort descending, count down to the row equal to your monthly capacity, and read the score on that row. That is your threshold, empirically, for your actual distribution. Do it again whenever headcount changes or volume moves by around a fifth, because both terms drift and a threshold from last year is rationing against a team that no longer exists.

The two failure modes are symmetrical and both destroy the model's credibility. Set it too low and reps receive more than they can work, so they cherry-pick, which means the score is no longer determining what gets worked and the whole apparatus becomes decorative. Set it too high and reps starve while marketing reports a healthy qualified count, and the argument that follows is about definitions rather than about anything real. Watch the acceptance rate as the check: if reps are rejecting most of what clears the bar, the bar encodes a marketing preference rather than a qualification.

One case deserves naming because the arithmetic gives an unusual answer. If capacity exceeds the entire in-profile lead volume, no threshold helps you. Scoring can only ration, and there is nothing to ration. The output should be to route everything that matches the profile, put the surplus capacity into outbound against the same profile, and stop treating a demand problem as a routing problem.

Weights you can defend, and the signal that ruins a model

Assigned-by-committee weights have a signature: round numbers, everything worth ten points, and a total that happens to reach a hundred. The defensible method is to compare conversion rates. For each attribute, calculate the rate at which leads carrying it become customers and the rate for leads without it, then set the weight in proportion to the lift. An attribute that doubles conversion should carry roughly twice the weight of one that adds a fifth. This is a crude approximation of what a logistic regression would produce and it is close enough to beat any workshop.

It needs data to be real. Below roughly thirty conversions carrying an attribute, the observed lift is noise and you will be encoding an accident of last quarter. Until you have that, normalise a priority ordering instead, which is what this tool does, and be explicit that it is a starting hypothesis rather than a measurement. A stated hypothesis gets tested. A number presented as a finding does not.

Then there is the trap that quietly ruins otherwise careful models: scoring the conversion event itself. "Requested a demo" is not a predictor of requesting a demo, it is the thing you are predicting, and a model built on it will look magnificent in validation and add nothing in operation, because those leads were going to a rep regardless of the score. Route conversion events directly and immediately, and keep them out of the model. The signals worth weighting are the ones observable before the buyer identifies themselves as ready.

The same rule applies in a subtler form to any attribute you cannot see at scoring time. Enrichment that arrives two days later is not available at routing, so a weight resting on it delays every decision. If a field is only populated after a sales conversation, it belongs in opportunity qualification, not in lead scoring. The practical test is whether the attribute is present within a minute of form submission for at least four leads in five.

Negative scoring must be able to outvote engagement, and scores must decay

A model that only adds points has a predictable failure. Every newsletter subscriber eventually becomes qualified through accumulation, and reps are handed leads whose most recent activity is three months old. The signature is easy to spot: sort your qualified leads by most recent activity date and look at the tail. If a meaningful number have not done anything for a quarter, your score is measuring history rather than intent.

Decay fixes it and the only decision is the half-life. Set it to the length of the buyer's active research window, which you can measure from your own data as the median time between first touch and demo request. Thirty days is a reasonable default for most business software and too long for anything transactional. A thirty-day half-life is a daily multiplier of about 0.977, which is the fiftieth root of a half, and it means a lead sitting at forty behaviour points falls to twenty after a month and ten after two, dropping out of routing without anyone having to intervene. Fit does not decay in the same way, because headcount and industry change slowly, though it should be re-verified quarterly since roughly a fifth of contact data goes stale in a year.

Negative scoring is the other half, and the common mistake is making the negatives too small to matter. If your worst disqualifier subtracts ten points while a maximum behaviour score is fifty, then a competitor doing thorough research still clears the bar, which is exactly who you least want in a rep's queue. Soft negatives should exceed your single largest behaviour signal so that no one action can cancel one.

Absolute disqualifiers should not be points at all. Competitor domains, existing customers, out-of-territory records, job applicants and free email addresses on a business product are gates: they remove the record from routing regardless of every other value in the model. The reason is arithmetic rather than philosophical. Any point value can in principle be outvoted by a sufficiently active lead, so if the answer must always be no, the mechanism has to sit outside the addition. Keep the count of gates small and reviewed, because a gate applied by mistake is invisible: nobody ever complains about the lead they never received.

Numbers worth knowing

MetricTypicalWhat it means
Conversions needed before a weight is a measurementabout 30 carrying the attributeBelow that, observed lift is noise from last quarter. Normalise a stated priority ordering instead and label it a hypothesis, which at least invites someone to test it.
Sales acceptance of scored leadsif under about half, the bar is wrongA rejection rate above half means the threshold encodes a marketing preference. Read it alongside cherry-picking: reps quietly skipping leads is the same signal in a different form.
Behaviour half-life30 days is a workable defaultSet it to your median time from first touch to demo request. Thirty days gives a daily multiplier of about 0.977; anything transactional needs a much shorter one.
Contact data decayroughly a fifth stale per yearPeople change jobs and companies restructure. Fit scores do not decay day to day but they do rot, which is why the model needs a quarterly re-verification pass rather than a one-off enrichment.
Weeks between threshold recalculationsmonthly, or on a 20% volume moveBoth terms drift. New headcount changes capacity immediately, and a campaign that doubles volume makes yesterday's threshold pass twice as many leads as the team can work.

Mistakes that quietly cost you results

Adding fit and behaviour into one number
A score of sixty then means either a perfect-fit account that read one page or an unsellable one that toured your pricing. Those need opposite handling, so keep the axes separate and route on the pair.
Scoring the conversion event itself
Weighting "requested a demo" predicts the thing you already know, so the model validates beautifully and changes nothing. Route conversion events straight to a rep and build the model from signals visible before the buyer declares themselves.
Setting the threshold at a round number like 75
It has no relationship to how many leads your team can work. Sort last month's scored leads descending, count down to the row equal to monthly capacity, and read the score there. That number is defensible because it came from your own distribution.
Using small negative points for absolute disqualifiers
Minus ten against a fifty-point behaviour ceiling means an active competitor still clears the bar. Absolute exclusions belong outside the addition as gates, and soft negatives should exceed your largest single behaviour signal.
Letting scores accumulate without decay
Every long-term subscriber eventually qualifies, and reps get leads whose last activity was in the previous quarter. Apply a half-life equal to your research window so stale intent falls out of routing without anyone having to intervene.

What does the output look like?

This is the exact output the tool produces from the example inputs. It is generated by the same code that runs when you click the button, so what you see here is what you get.

LEAD SCORING MODEL Fit is scored out of 50 and behaviour out of 50, separately, and they are never added into a single number for routing. FIT SCORE (firmographics, out of 50) 7 pts US and UK Geography: reliably observable but weakly predictive on its own; usually better as a gate than as points 11 pts 50-500 employees Company size: strong proxy for budget and process complexity, and available from enrichment at form fill 12 pts ecommerce retailers Industry or vertical: usually the highest-lift attribute, because it determines whether the problem exists at all rather than how big it is 12 pts using Shopify Plus Technology in use: the strongest routinely observable fit signal, because it evidences the workflow rather than describing the company 8 pts Head of Ecommerce or Ops Role or seniority: predicts authority but is self-reported on forms and therefore the least trustworthy field you collect Weights are normalised from a priority ordering, not measured. Once about 30 conversions carry an attribute, replace its weight with the measured lift: conversion rate of leads with it divided by conversion rate of leads without it. BEHAVIOUR SCORE (observable signals, out of 50) 15 pts pricing page visit High intent: commercial-page behaviour, the closest thing to a stated intention that you can observe passively 16 pts demo request High intent: this is a conversion event, not a predictor of one 8 pts return visit within 7 days Evaluation: evidence of an active evaluation rather than of casual interest, and it precedes the commercial pages 3 pts opened onboarding email Awareness: weak and abundant; useful for ordering a nurture list, close to worthless for routing 8 pts downloaded integration guide Evaluation: evidence of an active evaluation rather than of casual interest, and it precedes the commercial pages REMOVE THESE FROM THE MODEL: "demo request". They are conversion events rather than predictors of conversion. Scoring them makes the model look accurate in validation and change nothing in operation, because those leads were going to a rep whatever the score said. Route them directly, within minutes, and rebuild the weights from the signals that appear before a buyer declares themselves. CAPACITY AND THRESHOLD Capacity 3 reps x 18 leads per week x 4.33 weeks = 234 leads a month Lead volume 1,400 a month Qualifiable share 16.7% (capacity divided by volume) THRESHOLD: 55 of 100, which is the score that passes roughly the top 16.7% of a right-skewed inbound distribution. Split as a fit gate of 28 of 50 and a behaviour gate of 27 of 50, which must both be cleared. Both gates together equal the threshold, so nothing is double counted. The curve above is a model of a typical inbound distribution, not your distribution. Replace it in ten minutes: score last month's leads, sort descending, count down to row 234, and read the score on that row. That number is the defensible one because it came from your own data. ROUTING MATRIX (monthly volumes, assuming 30% of inbound matches the ICP) A High fit, high behaviour 234 to a rep now, 18.0 per rep per week against a stated capacity of 18 B High fit, low behaviour 186 marketing and outbound keep these: right account, wrong moment. Trigger on a job change, a funding round or a stack change rather than on a date C Low fit, high behaviour 546 self-serve, docs or community. Never a rep. If this box is larger than box A, your content is earning an audience you cannot sell to, which is a specific and fixable problem D Low fit, low behaviour 434 suppress, and treat suppression as an action rather than as an absence of one Boxes sum to your 1,400 leads. Behaviour is assumed independent of fit, which overstates box C, since in practice the two correlate positively. Treat C as an upper bound until you can count the real one. Box A is sized to your capacity by construction. If real box A volume comes in above it, raise the behaviour gate rather than the fit gate: you want fewer, later-stage leads from the same accounts, not a narrower definition of the accounts. NEGATIVE SCORING Soft negatives, applied as points: -21 free email address on a business product, or an unfillable company field -21 career, jobs or press pages visited in the same session as your product pages -10 unsubscribed from email but still browsing, which is a real but weaker negative Each soft negative is larger than your biggest single behaviour signal (16 pts), which is deliberate: if one strong action can cancel a disqualifier, the disqualifier does nothing. Hard gates, applied outside the addition: competitor domains, existing customers, current open opportunities, job applicants, out-of-territory records These are gates rather than points because any point value can be outvoted by a sufficiently active lead. Keep the gate list short and reviewed, since a wrongly applied gate is invisible: nobody complains about the lead they never received. DECAY RULE Behaviour points decay with a 30-day half-life, a daily multiplier of about 0.977. Set the half-life to your own median gap between first touch and demo request rather than to 30 days. A lead at the behaviour gate of 27 falls to 14 after 30 days of silence and 7 after 60, dropping out of routing without anyone deciding to remove it. Fit points do not decay day to day, but re-verify them quarterly: contact and firmographic data goes stale at roughly a fifth a year, so an unmaintained fit score slowly becomes a record of who these companies used to be.

Frequently asked questions

How are the weights calculated rather than guessed?

Each ICP attribute and each signal is classified, given a raw priority based on how much it typically separates buyers from non-buyers and how reliably it can be observed at routing time, and then normalised so that fit sums to exactly fifty points and behaviour sums to exactly fifty. That makes the weights relative to each other rather than invented, and it keeps the two halves comparable. It is still a hypothesis until you have data: once around thirty conversions carry an attribute, replace the priority with the measured conversion lift for leads that have it versus leads that do not.

Why is the threshold based on sales capacity?

Because a threshold is a rationing decision, not a statement about quality. Capacity is reps multiplied by leads per rep per week multiplied by about 4.33 weeks a month; the share you can qualify is that divided by lead volume, and the threshold is whatever score produces that share. Set it lower and reps cherry-pick, which means the score no longer decides what gets worked. Set it higher and reps starve while the qualified count still looks healthy on a dashboard.

Why keep fit and behaviour scores separate?

Because the same total means opposite things. A right-fit company with almost no activity is a timing problem that belongs in nurture with an outbound trigger. A wrong-fit company with heavy activity is a targeting problem that belongs in self-serve, and giving it to a rep costs the time and then costs you their trust in the score. Adding the two assumes engagement can compensate for being the wrong company, which it cannot, because fit is a constraint and behaviour is a state.

What should the decay rule be?

A half-life equal to your buyer's active research window, measurable as the median gap between first touch and demo request. Thirty days is a sensible default for business software and works out at a daily multiplier of about 0.977. Without decay, every long-term newsletter subscriber eventually accumulates enough points to qualify, and the giveaway is qualified leads whose most recent activity was a quarter ago.

How large should negative scores be?

Soft negatives need to exceed your single largest behaviour signal, otherwise one strong action cancels a genuine disqualifier and an active competitor lands in a rep's queue. Absolute exclusions such as competitor domains, current customers, job applicants and out-of-territory records should not be points at all, since any point value can be outvoted by enough engagement. Make those gates that remove the record from routing regardless of its score.

Related free tools

Some links on this site are affiliate links, which means Hacking Demand may earn a commission if you buy through them at no extra cost to you. This does not influence which tools are listed. The tools on this page are free and have no affiliate relationship of any kind.