Teacher Evaluation Software Problems: The 5 That Cost Districts Real Money, and How to Avoid Them
The most expensive failure mode is a rating that is correct on the merits and indefensible on the record. A principal logged eleven of thirteen required unannounced visits, entered one of them in April for a walkthrough he says happened in November, and the association asks a single question: show me the dates. The district settles, the non-renewal is withdrawn, and every principal in the building learns that the process is theatre. The cost is the settlement, the arbitration exposure, and the two to three weeks of HR and principal time that vanish every spring reconstructing a paper trail that should have been a byproduct of doing the work.
Why does a teacher evaluation project become a forms project?
Nearly every district that tells us their evaluation software failed describes the same scope failure. The requirement list was written as forms: an observation form, a pre-conference form, a summative form, a routing rule for signatures. That is a well understood build, it demos beautifully in June, and it does not touch the thing that actually loses grievances.
What loses grievances is the timeline. How many announced and unannounced observations by tenure status. How many school days, not calendar days, may pass between a visit and written feedback. How much notice a pre-conference requires, and who is permitted to observe. A forms project records what you did. It does not stop a principal doing it late, and it does not tell him on the third of March that he has eleven school days left and two visits outstanding.
Scope the first release around observation capture, evidence, and a contract-aware timeline engine with principal caseload dashboards that surface days remaining rather than work completed. That is $70,000 to $140,000 over 12 to 18 weeks in our delivery experience, and it should be in principals' hands before the first observation window, not after it. Composites, student learning objectives, improvement plans, professional development hours and state reporting are the second phase, and they are considerably easier to fund once spring has been quiet.
What goes wrong when prior evaluation records are migrated?
Districts almost always ask for the last three to five years of evaluations to come across, and the request is reasonable, because improvement plans and non-renewal recommendations reference them. The trouble is that historic records were produced under different rules. The rubric was revised in a settlement two years ago. The growth weighting changed when the state amended its model. A domain was renumbered. Load those records into a current structure and every historic rating silently reads as if it had been produced under today's rules, which is precisely the claim you cannot afford to make in a hearing.
The second problem is staff identity across time. Teachers change buildings, assignments and bargaining units. They leave and return. Payroll identifiers get reissued after a system upgrade. If the migration keys on a current staff record, an evaluation from a teacher's previous assignment attaches to the wrong caseload or disappears from their file entirely, and nobody notices until the file is requested.
What works: migrate historic evaluations as documents with their metadata, dated and attributed, rather than as scored records in the new model. Store the rubric version and formula version that applied at the time alongside them. Reserve structured, recomputable records for years that begin inside the new system. Then reconcile staff identity against your human resources system by employee identifier with a manual review queue for the ambiguous cases, and expect that queue to be longer than anyone predicted, particularly for long-serving staff.
Why do the human resources and student information system links break after launch?
Rostering integration looks like a solved problem in August and becomes the top support ticket in October. Staff arrive as a nightly file from the human resources system and assignments arrive from the student information system, and the two disagree in ways that are invisible until they matter. A long-term substitute appears in the student information system with a class assignment and does not exist in human resources. A teacher splits a schedule across two buildings and only one principal sees them. A co-taught section attributes both teachers to the same roster, so student growth attribution doubles.
The other reliable breakage is the mid-year change. A teacher transfers in November. The evaluation already has three completed observations under the previous evaluator in a different building, possibly under a different bargaining unit if they moved from a counselling role. Systems that model an evaluator as a field on a staff record cannot represent that cleanly, so somebody fixes it by editing the record, and the audit trail you built the system for is gone.
The fix is to model the evaluation cycle as its own object with an evaluator assignment history, a building history and a unit at each point in time, so a transfer is a new segment rather than an overwrite. Reconcile human resources and student information system feeds daily with an exception report that names the person, and give HR an owner for that report. Any staff member appearing in one system and not the other is usually a person nobody is evaluating.
What happens when composites and personnel confidentiality are not covered?
Two gaps sit outside the observation workflow and both bite late. The first is recomputation. A summative rating combines weighted observation domains, a state produced growth measure for tested subjects, and locally written student learning objectives for everyone else, with minimum data thresholds, rounding rules and cut scores. If the system stores only the output, an appeal against a rating from three years ago cannot be answered, because nobody can reproduce the number. Version the formula by school year and bargaining unit, store the inputs rather than the result, and expose a recompute function that prints the derivation term by term.
The second gap is the access model. Evaluation data is an employee personnel record, not a student education record, so the rules that govern it are your state's personnel file and public records statutes plus the confidentiality terms in your bargaining agreement. A principal sees their caseload. Human resources sees the district. Whether a superintendent may see a draft before it is finalised is often a contract question, and getting it wrong is itself a grievance. Any developer whose access discussion begins and ends with FERPA has misread the problem, although FERPA does apply to student data appearing inside a growth measure.
Build both properly in the first phase that touches ratings. Retrofitting an access model into a system already holding three years of personnel records is expensive and visible.
Should you build custom or configure what you already own?
If you have one bargaining unit, one rubric, fewer than roughly 600 certificated staff and a contract that has been stable for several cycles, do not build. Frontline Professional Growth or Standard for Success will serve you, and a custom build would be an expensive way to own a maintenance burden. Standard for Success in particular is genuinely good at fast mobile capture with evidence tagged to indicators, which is the thing principals feel daily.
Before you commission anything, spend a week with the product you already licence and a person from your association. Most districts have never configured their existing tool against the actual contract, because the person who set it up was in technology rather than human resources. Map every observation count, notice period and feedback window in the agreement to a setting in the product, and write down what has no home. If almost everything maps, your problem is training and process, not software, and a build will not fix it. If the list of homeless clauses is long, you now have a specification that cost you nothing.
Build when two or more of these are true. You run three or more distinct rubrics or bargaining units, because counsellors, psychologists and speech pathologists usually need their own process rather than a bent copy of the teacher one. You carry more than roughly 1,500 certificated staff, at which point spring reconstruction becomes a staffing problem. You have lost or settled a grievance on process rather than substance. Your state mandates a composite your vendor cannot configure without a change request. Or you are a charter network, state agency or regional service centre running evaluation for many districts, where multi-tenancy with genuinely different rules rules out most packaged options.
How do hidden costs get into the quote?
Five items produce most of the overrun in district evaluation builds.
- Bargaining units counted as configuration. Each unit is a separate process model with its own rubric, counts, notice rules and composite, not a copy with different labels. Three units is roughly three times the process work, and quotes routinely price it as one.
- The state reporting file. Usually a fixed width layout with validation rules and a submission window. It takes real weeks, it cannot be shortened, and it is often scoped as an export.
- Offline mobile. More buildings have poor wireless coverage than anyone admits, and reliable offline capture with synchronisation is genuine mobile engineering rather than a setting.
- Single sign-on and human resources integration, particularly against an older on-premise system where the district's own vendor controls the timeline.
- Contract interpretation. Human resources, the association and building principals frequently describe the same clause three different ways, and someone has to convene them and produce one authoritative reading. Districts that do this in week one move noticeably faster than those that discover the disagreement in user acceptance testing.
What separates a build that works from one that fails here?
Ask a prospective developer to model a school day calendar in front of you. If they reach for calendar days, or cannot immediately name the problem of a snow day falling inside a five day feedback window, they have not built this and your timeline engine will be wrong in exactly the cases that reach arbitration.
Ask what happens when an evaluator wants to change a word after the teacher has acknowledged an observation. The right answer is an amendment record with the original preserved and the author named, not an edit. That single design decision is what turns a hearing from a two day reconstruction into a printed timeline, and it changes principal behaviour once everybody understands that the record is the record.
Ask how a rubric is versioned mid-year when a settlement lands in October and applies retroactively. Process templates need to be effective dated so observations completed under the old rules keep them. A product that stores a single current configuration cannot answer that question honestly.
Then start with the classroom teacher unit and one school year, add units in year two, and settle ownership in writing before kickoff. You should own the repository, the cloud accounts and the evaluation records themselves. At Digital Heroes the client owns everything from the first commit, and these records may be requested years after a teacher has left the district.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- Organizations that scaled intelligent automation report an average cost reduction of 32% (up from 24% in 2020), and respondents expect an average 31% cost reduction over the next three years. Source: Deloitte (2022) →
- An EY survey found one in five U.S. payrolls contains errors, each costing an average of $291 to remediate, with a typical 1,000-employee organization spending roughly 29 workweeks per year fixing common payroll errors. Source: EY (Ernst & Young) (2022) →
- McKinsey found that currently demonstrated technologies can fully automate about 42% of finance activities and mostly automate a further 19%, indicating roughly 60% of finance work is technically automatable. Source: McKinsey & Company (2018) →
- In an October 2025 survey of 530 small-business employers (conducted by TechnoMetrica, October 3-9, 2025), 88% reported using AI tools and 73% said those tools had been important to their competitiveness and growth over the past year, with 60% citing efficiency and productivity as the primary motivation for adoption (42% cited improving customer service). Source: Small Business & Entrepreneurship Council (SBE Council) (2025) →
Olivia is a senior product designer working on the software side of Digital Heroes: dashboards, admin tools, internal systems and the screens people use all day rather than once. She writes about designing for repeat use, where speed and clarity matter more than a striking first impression.
View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.
Frequently asked questions
Why do districts lose grievances even when the evaluation software worked?
Should we migrate five years of prior evaluations into the new system?
What happens when a teacher transfers buildings mid-year?
How do we answer an appeal against a rating from three years ago?
Is FERPA the right framework for evaluation record access?
What should we configure in Frontline or Standard for Success before considering a build?
Why does adding a second or third bargaining unit cost so much?
What happens when a settlement lands in October and applies retroactively?
Can we keep using BambooHR while the custom system is being built?
How long does it take to build a custom HR system?
How many SaaS seats do we need before building custom becomes cheaper?
Who owns the code if an agency builds our HR software?
Is Workday realistic for a company under 500 employees?
How many developers does it take to build an HR platform?
How do I calculate whether custom software will pay for itself?
Who owns the code when an agency builds my software?
What tech stack should custom HR software use?
What should I prepare before contacting an agency about HR software?
Who can build a custom HR software system?
Digital Heroes builds custom HR software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.
Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.
What makes Digital Heroes different from other HR software companies?
Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.
Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.
How can I check Digital Heroes is legitimate before getting in touch?
Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.
Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.