Common Gherkin Anti-Patterns
Gherkin anti-patterns are practices that break BDD principles, reduce readability, and create brittle, high-maintenance automation. Gherkin is meant to describe business behavior in a shared language that product owners, business analysts, developers, testers, and automation engineers can all understand. When scenarios become technical, procedural, vague, oversized, or inconsistent, the feature files stop serving that purpose. They become test scripts with Given, When, and Then keywords.
Recognizing common anti-patterns is essential for scalable BDD. A small project may survive with inconsistent wording and UI-heavy steps for a while, but a growing product cannot. As the number of feature files increases, every weak habit becomes expensive. Duplicate step definitions multiply. Scenarios become harder to review. Reports become harder to interpret. Automation becomes tightly coupled to the UI. The cost is not only technical; the team also loses the collaboration value that BDD is supposed to provide.
Anti-patterns also tend to spread. If one feature file uses click-by-click wording, new contributors often copy that style. If one scenario uses vague assertions, similar vague assertions appear elsewhere. If one team overuses Scenario Outline, another team may repeat the pattern because it appears efficient. This is why Gherkin quality needs active review. Feature files should be treated with the same seriousness as production code. They deserve naming standards, refactoring, review, and shared ownership.
The purpose of this article is not to make Gherkin rigid or academic. The goal is practical. Good Gherkin should help teams agree on behavior before coding, automate that behavior after implementation, and report meaningful results during execution. Anti-patterns weaken one or more of those goals. The better a team becomes at spotting them, the easier it becomes to keep BDD useful over time.
1. UI-Driven Procedural Scenarios
The most common Gherkin anti-pattern is writing scenarios as screen-level procedures. These scenarios describe clicks, fields, buttons, page navigation, browser actions, and element interactions. They look detailed, but they describe how automation interacts with the UI rather than what business behavior the system provides.
Consider this anti-pattern:
Given I open the browser
When I click the Login button
Then I enter username and password
This scenario is tied to UI details. It is fragile because the wording depends on a particular interface design. If the login button changes, if the page is redesigned, or if the login flow moves to single sign-on, the scenario becomes outdated even if the business behavior remains valid. It is also not business-readable because it talks about browser mechanics instead of authentication behavior.
A behavior-focused version is clearer:
Given the user has valid credentials
When the user logs in
Then the user should be authenticated
The improved version describes the business intent. The implementation may still open a browser, enter credentials, and click a button, but those details belong in step definitions, page objects, or action classes. The Gherkin should express behavior. This separation makes scenarios more stable and more meaningful as living documentation.
How to Detect Procedural Gherkin
If the scenario reads like a manual test case or a Selenium script, it is probably procedural. Words like click, type, scroll, hover, browser, textbox, button, dropdown, and page are review signals. They do not always make a step wrong, but they should make the team ask whether the step is describing business intent or UI mechanics.
A useful review question is: would this scenario still make sense if the UI changed completely? If the answer is no, the scenario is probably too procedural. A login capability may still exist even if the form becomes a single sign-on redirect. An order submission may still exist even if the checkout button becomes a swipe action in a mobile app. Good Gherkin should stay meaningful across those implementation changes.
Another useful question is: could this behavior be executed through an API or a mobile app without changing the scenario wording? If the answer is yes, the wording is probably behavior-focused. If the wording only makes sense in one browser screen, it is likely UI-driven.
2. Technical Assertions in Gherkin
Another common anti-pattern is placing technical assertions directly in Gherkin. Examples include HTTP status codes, database table checks, JSON paths, framework terms, environment names, or internal service details. These checks may be valid from a testing perspective, but they often do not belong in the business-facing scenario text.
For example:
Then HTTP status code should be 200
This step may be useful in an API test, but it is not a strong BDD acceptance statement unless the audience is explicitly technical and the status code is the actual requirement. Most business readers do not think in HTTP response codes. They think in outcomes such as an order being confirmed, an account being created, or an invalid request being rejected.
A better step is:
Then the order should be confirmed
The step definition can still verify that the HTTP status code is 200, that the response body contains the correct order number, and that the database has the expected state. Those technical checks support the behavior, but the feature file does not need to expose them. This keeps Gherkin readable and prevents implementation leakage.
Where Technical Checks Belong
Technical checks belong in automation code, helper methods, API clients, database utilities, or assertion layers. The Gherkin step should describe the business result. This does not reduce test rigor. It improves communication while still allowing strong validation underneath.
There are cases where technical language is acceptable, especially in a feature file intended for a technical API team. Even then, the team should be deliberate. If the business requirement is truly that an API returns a 401 for unauthorized access, the status code may be part of the expected behavior. But if the real requirement is that unauthorized users cannot access account details, the Gherkin should usually express the access rule and let the step definition validate the technical response.
The safest principle is to make the feature file readable to its intended audience. If the audience includes business stakeholders, avoid implementation terms. If the audience is purely technical, use technical terms only when they describe the behavior being specified, not merely the mechanism used to verify it.
3. Overusing Background
Background is useful when several scenarios share the same simple context. However, overusing Background is a common anti-pattern. Large Background blocks hide important setup and make scenarios harder to understand. Readers must constantly combine the Background with each scenario to know what is really being tested.
Consider this example:
Background:
Given user exists
And user is logged in
And user has cart
And payment method exists
This Background may look efficient, but it hides several assumptions. Does every scenario really require a logged-in user? Does every scenario need a cart? Does every scenario depend on an existing payment method? If some scenarios rely on those conditions and others do not, the Background creates confusion. It also makes failures harder to reason about because setup problems may appear unrelated to the scenario being read.
The fix is to keep Background limited to true, minimal common context. Critical assumptions should appear inside the scenario where they matter. If payment method eligibility is important to a payment scenario, include it in that scenario. Do not hide it in a large Background block where readers may miss it.
Good Background Usage
A good Background is short, stable, and genuinely common. It should establish broad context, not hide business rules. If a Background contains more than a few lines or includes details that affect only some scenarios, it is a candidate for refactoring.
Background should also avoid procedural setup. A Background that logs in through the UI, creates data through several screens, and navigates to a page is usually doing too much. It may be better to create the required state through fixtures, APIs, or direct setup methods. The scenario can then describe the business context in a simple way, such as "Given the customer has an active account."
When reviewing Background, ask whether every scenario in the feature truly depends on every line. If not, move the condition into the specific scenario that needs it. This makes each scenario more self-explanatory and reduces accidental dependencies between scenarios.
4. Mega Scenarios That Are Too Coarse
Mega scenarios occur when one scenario validates multiple behaviors. A scenario titled "User completes purchase flow" may include login, cart updates, shipping, discounts, payment, confirmation, email notification, and order history. It may appear valuable because it covers a realistic journey, but it usually has too many reasons to fail.
Scenario: User completes purchase flow
The problem with a mega scenario is poor diagnostics. If it fails, the report tells the team that the purchase flow failed, but it does not immediately identify which behavior is broken. The failure could be caused by authentication, inventory, payment, address validation, or notification. This makes triage slower and reduces confidence in reports.
The fix is to split broad flows into focused scenarios. For example, create separate scenarios for successful order placement, payment failure handling, invalid coupon rejection, and confirmation email delivery. Each scenario should validate one outcome with one clear reason to fail.
When a Broad Flow Is Acceptable
A small number of broad end-to-end scenarios may be acceptable for smoke testing or release readiness. The anti-pattern is not having any end-to-end coverage. The anti-pattern is using mega scenarios as the main way to validate every rule. Detailed behavior should be covered by focused acceptance scenarios, API tests, integration tests, or unit tests where appropriate.
A useful way to refactor a mega scenario is to list every Then outcome and every major business action. Each separate outcome may become its own scenario. Each major action may become a Given state in another scenario if it is only setup. For example, instead of testing cart creation, payment, and confirmation in one scenario, you can start with "Given the user has eligible items in the cart" and focus the scenario on payment confirmation.
This refactoring improves reports. Instead of one failure saying "purchase flow failed", the report can say "payment is declined for an expired card" or "order is confirmed after successful payment." Those names are more useful during triage and release decisions.
5. Over-Splitting Scenarios That Are Too Fine
The opposite anti-pattern is over-splitting. This happens when teams create one scenario for each tiny UI step. The feature file becomes full of scenarios such as "Click login", "Enter username", and "Enter password." These scenarios are small, but they are not meaningful business behavior.
Scenario: Click login
Scenario: Enter username
Scenario: Enter password
Over-splitting creates low-value scenarios and noisy reports. A business stakeholder does not usually care that a user can click a login button in isolation. The useful behavior is that a valid user can log in successfully or that an invalid user is denied access. Over-split scenarios also make automation tightly coupled to UI design.
The fix is to combine related UI actions into one behavior-focused scenario. Instead of separate scenarios for each login field interaction, write a scenario such as "Successful login with valid credentials." The step definitions can perform the UI actions behind the scenes.
Over-splitting sometimes happens because teams want more detailed reporting. But more scenarios do not automatically mean better reporting. A report with many low-value scenarios can hide the important signals. Good reporting comes from meaningful scenario boundaries, not from turning every interaction into a separate scenario.
The right question is not "Can this be tested separately?" It is "Does this represent a behavior worth documenting separately?" If the answer is no, keep the detail in the automation layer rather than the feature file.
6. Multiple When Actions
A scenario should usually have one primary When action. The When step represents the action that triggers the behavior under test. If a scenario contains several main actions, it likely combines multiple behaviors and has unclear intent.
When the user logs in
And the user updates profile
Logging in and updating a profile are separate behaviors. They may occur in sequence, but they do not necessarily belong in the same scenario. If the scenario fails, the team may not know whether authentication failed, profile update failed, or setup was incorrect. This weakens the value of the report.
The fix is to keep one primary When per scenario. If another action is only setup, move it into a Given step expressed as state. For example, "Given the user is logged in" can set up authentication, while "When the user updates the profile" remains the main behavior. If both actions are important behaviors, split them into separate scenarios.
One Main Action, One Main Outcome
This principle helps scenarios stay focused. A scenario with one main action and one main outcome is easier to read, easier to automate, and easier to debug. It also maps more naturally to acceptance criteria.
There can be supporting And steps after When, but they should not introduce a new behavior. For example, "When the user updates the shipping address and saves the changes" may be acceptable if saving is part of the same action. But "When the user updates the shipping address and places an order" is likely two behaviors. The difference is whether the second action is part of completing the same behavior or starts a new one.
7. Vague or Generic Steps
Vague steps are another serious anti-pattern. A step such as "Then the system should work correctly" says nothing concrete. It cannot be verified meaningfully, and it does not help the team understand acceptance criteria.
Then the system should work correctly
This step is weak because "work correctly" can mean many things. Does the user receive an email? Does the record appear in a report? Is the payment approved? Is access granted? A vague step hides the expected outcome and forces the reader to infer meaning from context or automation code.
A better step is specific and observable:
Then the user should receive a confirmation email
This wording gives the scenario a verifiable result. It also communicates business value. Good Gherkin should make expectations explicit enough that stakeholders can review and agree on them before development begins.
Vague wording often appears when requirements are unclear. Instead of hiding uncertainty behind generic language, use the scenario discussion to clarify the expected behavior. Ask what the user should see, what record should be created, what message should be sent, what access should be granted, or what rule should be enforced. The answer becomes the Then step.
Clear Then steps are especially important because they define success. A weak Then step leads to weak assertions. A strong Then step encourages strong validation in the automation code.
8. Step Explosion and Inconsistent Vocabulary
Step explosion happens when the same intent is expressed in many different ways. For example, one feature may say "user logs in", another may say "user signs in", and another may say "user authenticates." These phrases may mean the same thing, but automation tools may treat them as different steps unless they are carefully managed.
- user logs in
- user signs in
- user authenticates
The result is duplicate step definitions, ambiguous matches, and inconsistent documentation. Teams waste time maintaining multiple phrases for the same behavior. New contributors are unsure which phrase to use. Reports become less consistent.
The fix is to standardize domain vocabulary. Choose the phrase the product uses and reuse it consistently. If the product language is "sign in", use that everywhere. If the domain says "authenticate", use that consistently. A shared glossary can help larger teams keep wording aligned.
Why Vocabulary Discipline Matters
BDD depends on shared language. Inconsistent wording weakens that shared language. Standardized vocabulary improves readability, reduces duplicate automation, and makes feature files easier to search and maintain.
A shared vocabulary does not need to be complicated. It can be a simple page in the project wiki or a short glossary inside the repository. The important part is that the team agrees on terms. For example, decide whether the product says customer, user, member, applicant, or account holder. Decide whether the action is sign in, log in, or authenticate. These small decisions prevent step explosion later.
Automation engineers should also review new step definitions before adding them. If an existing phrase already represents the same intent, reuse it. This keeps the step library smaller and more predictable.
9. Over-Parameterized Steps
Parameterization is useful when it represents data variation, but it becomes an anti-pattern when it hides intent. A step such as the following may appear reusable, but it provides poor documentation:
When the user performs "<action>"
This step is too generic. It tells the reader that something happens, but the actual meaning is hidden in the example data or automation implementation. Over-parameterized steps often become dumping grounds for unrelated behavior. They reduce readability and make reports less meaningful.
The better principle is to parameterize data, not intent. For example, "When the user places the order" should remain a meaningful action. Data such as payment method, country, coupon code, or account type can be parameterized when the behavior is otherwise the same.
When the user places the order
Good parameterization keeps the scenario expressive. Bad parameterization makes the scenario abstract and empty.
Scenario Outline tables should also be reviewed for meaning. If an examples table contains an action column, an expected result column, an error type column, and several unrelated flags, the outline may be hiding multiple behaviors. Good outlines have a clear pattern: the same behavior runs with different data. If the behavior changes row by row, split the outline.
10. Treating Gherkin as Test Case IDs
Some teams write scenario names like manual test case identifiers. For example:
Scenario: TC_001_Login_Test
This is an anti-pattern because it provides no business value. A reader cannot understand the behavior from the scenario name. The identifier may be useful in a test management tool, but it should not replace a meaningful scenario title.
A better scenario name describes the behavior:
Scenario: Successful login with valid credentials
This title is readable, searchable, and useful in reports. If traceability is required, a tag can reference the external test case or requirement ID. The scenario name should remain human-readable.
For example, use a tag such as @REQ-123 or @TC-001 if the team needs traceability. The tag can satisfy reporting and mapping needs while the scenario title remains descriptive. This preserves both governance and readability.
11. Writing Scenarios After Development
BDD loses much of its value when scenarios are written only after development is complete. In that case, scenarios often mirror what was already built instead of shaping what should be built. The team misses the early conversation where examples can clarify ambiguous requirements.
When scenarios are written after development, BDD becomes documentation only. Documentation is useful, but it is not the full value of BDD. The best value comes when scenarios are discussed before coding and used as acceptance criteria. This allows product owners, business analysts, developers, and testers to agree on expected behavior early.
The fix is to write or at least outline scenarios during refinement or before implementation. They do not need to be perfect at first, but they should guide the conversation. As understanding improves, the scenarios can be refined and automated.
This early use of examples is one of the biggest differences between BDD and simple test automation. BDD is not just a tool choice. It is a collaboration practice. Writing scenarios before development helps reveal missing rules, unclear acceptance criteria, conflicting assumptions, and edge cases that might otherwise be discovered late.
12. Using Comments to Explain Poor Gherkin
Comments are sometimes used to compensate for unclear steps. For example:
# This means user is logged in
Given userSession=true
This is a sign that the Gherkin step is not self-explanatory. If a comment is required to explain what the step means, the step should usually be rewritten. Feature files should be readable without decoding technical variables or hidden meanings.
A better version is simple:
Given the user is logged in
Comments are not forbidden. They can be useful for temporary notes or clarifying unusual domain context. However, they should not be used to hide poor wording. Good Gherkin should communicate clearly through the scenario itself.
If a comment explains a business term, consider whether that term belongs in a glossary or a clearer step. If a comment explains a technical variable, rewrite the step in domain language. Comments should support readability, not rescue unclear scenarios.
13. Overusing Scenario Outline
Scenario Outline is powerful when the same behavior is tested with different data. It becomes an anti-pattern when teams force too many variations, columns, and outcomes into one outline. Large outlines are hard to read and often mix different behaviors.
An outline with many columns may look efficient, but it can hide intent. If each row represents a different rule or a different expected outcome, the outline may be doing too much. Readers must scan across the table to understand which behavior is being tested. Reports may show example failures without clear business meaning.
The fix is to use Scenario Outline only when the behavior is identical and the data changes. If outcomes differ significantly, split the scenario. If columns become too many, consider whether separate scenarios would communicate better. Readability is more important than reducing the number of scenario blocks.
Good Scenario Outline Usage
A good outline might test invalid password rules where the structure is identical for each row. A poor outline tries to combine login success, login failure, locked account, expired password, and two-factor authentication into one table. These are related areas, but they are not identical behavior.
A practical guideline is to scan the Then column or expected result column. If every row has the same type of expected result, the outline is probably reasonable. If every row tells a different story, the outline is probably mixing behaviors. Split it into scenario groups that read naturally.
14. Automation-First Thinking
Automation-first thinking happens when Gherkin is written to satisfy automation convenience rather than business understanding. Symptoms include UI language, technical checks, long procedural flows, over-generic steps, and scenario names that mirror implementation details. The feature file becomes optimized for step reuse rather than communication.
This anti-pattern is understandable because automation engineers often maintain the Cucumber framework. They naturally think about locators, page objects, APIs, fixtures, and assertions. That knowledge is valuable, but it belongs in the implementation layer. The Gherkin layer should be written for shared understanding first.
The fix is to write for the business first and implement automation second. During review, ask whether a product owner or business analyst can understand the scenario without reading code. If not, the scenario probably needs to be raised to a behavior level.
This does not mean automation concerns are ignored. Automation feasibility matters. But the order of thinking matters. First clarify the behavior. Then decide the best automation design. If the team starts with automation convenience, the feature file often becomes procedural. If the team starts with behavior, the implementation can still be efficient while the scenario remains readable.
Quick Anti-Pattern Checklist
Before finalizing a scenario, use a short checklist. Does it mention UI elements? Can the business understand it? Does it validate one outcome? Would it survive a UI rewrite? Does it avoid technical implementation language? Does it use the team's standard domain vocabulary? If the answer exposes weakness, refactor before automation hardens the problem.
- Does it mention UI elements unnecessarily?
- Can business understand it without code?
- Does it validate one clear outcome?
- Would it survive a UI rewrite?
- Does it avoid technical assertions in the feature file?
- Does it reuse standard vocabulary?
Interview-Ready Summary
In interviews, a strong answer should explain that Gherkin anti-patterns break the value of BDD by making scenarios unclear, brittle, technical, or hard to maintain. Common issues include UI-driven steps, technical assertions, overused Background blocks, mega scenarios, over-splitting, multiple When actions, vague steps, inconsistent vocabulary, over-parameterization, and automation-first thinking.
The solution is to focus on behavior, clarity, and intent. Good Gherkin should describe business truth, not test scripts. It should use consistent domain vocabulary, validate one outcome per scenario, and hide technical details inside step definitions and automation layers.
- Gherkin anti-patterns break BDD value.
- Common issues include UI-driven steps, mega scenarios, and technical checks.
- Focus on behavior, clarity, and intent.
- Standardize vocabulary and keep scenarios focused.
- Good Gherkin reads like business truth, not automation scripts.
Golden Rule
The golden rule is simple: if Gherkin optimizes automation over understanding, BDD has already failed. Automation matters, but it should support the shared behavior language, not dominate it. A scenario that is easy for automation but meaningless to the business is not good BDD.
Good Gherkin creates alignment before code is written, supports maintainable automation after code is built, and remains useful as documentation after the sprint is complete. Avoiding anti-patterns protects all three purposes. It keeps feature files readable, automation stable, and reports meaningful.