Decoding by Run-Length - The Elementary Unit: How the Decoder Turns Pulse Widths into Modules and Modules into Characters |
Subtitle: A Deep Dive into Run-Length Decoding, Module Normalization, and the Symbology Lookup Table - with Real-World Designs from Symbol, Zebra, Honeywell, Datalogic, Microchip, and NXP |

|
Opening Summary |
The edge counter has done its job. It has captured the sequence of pulse widths from the digitised waveform and stored them in a buffer. Now the decoder must make sense of these numbers. The first step is to convert the raw pulse widths into a standardized format. This is done by run-length decoding, which transforms the sequence of alternating bars and spaces into a series of element widths measured in units of the module - the narrowest bar or space. This normalized sequence is the elementary unit of the barcode, and it is the foundation for all subsequent decoding. |
This article is dedicated to run-length decoding - the process that turns the raw pulse widths into module-width units. We will explore the concept of the module, the estimation of the module width, and the normalization of the pulse widths. We will examine the different techniques for handling scanning speed variations and print quality variations. We will look at how major companies have implemented run-length decoding in their products. We will see how Symbol (now Zebra) used a simple but effective method in the LS2208, based on finding the shortest pulse. We will explore Honeywell's use of a histogram-based method to estimate the module width in their imagers. We will examine Datalogic's adaptive method that tracks speed changes during the scan. We will also look at reference designs from Microchip, NXP, and STMicroelectronics, which include complete run-length decoding examples. |
By the end of this journey, you will understand that run-length decoding is the first and most crucial step in the decoder's algorithm. It is the bridge between the raw timing measurements and the symbolic representation of the barcode. |

|
Full Article |
Section 1: Run-Length Decoding - The First Step |
Run-length decoding is the process of converting a sequence of measured pulse widths into a sequence of element widths, measured in units of the module. The module is the width of the narrowest bar or space in the barcode. All other bar and space widths are integer multiples of the module: 2x, 3x, or 4x. |
The decoder receives a sequence of pulse widths from the edge counter. These pulse widths are measured in microseconds or timer counts. The decoder must estimate the module width from this sequence. It then divides each pulse width by the module width to obtain the normalized element width. The normalized element widths are integers (1, 2, 3, or 4), which represent the widths of the bars and spaces in module units. |
Run-length decoding is the first step in the decoder's algorithm. It transforms the raw timing data into a standardized format that can be processed by the symbology decoder. |

|
Section 2: The Module - The Fundamental Unit |
The module is the fundamental unit of the barcode. It is the width of the narrowest bar or space. All other bar and space widths are integer multiples of the module. The module is the ruler that measures all the other elements. |
The module width is not a fixed value. It depends on the print quality and the scanning speed. A faster scan produces a smaller module width. A slower scan produces a larger module width. The decoder must estimate the module width from the pulse widths. |

|
Section 3: Estimating the Module Width - The Shortest Pulse Method |
The simplest method for estimating the module width is the shortest pulse method. The decoder finds the shortest pulse in the entire sequence of pulse widths. This shortest pulse is assumed to be one module wide. The module width is set to the duration of this shortest pulse. |
The shortest pulse method is simple and effective. It works well when the barcode has at least one narrow element. The method is used in Symbol's LS2208 and in many other scanners. |
The shortest pulse method is vulnerable to noise. If a noise spike creates a very short pulse, the module width will be underestimated. To mitigate this, the decoder may use a median filter or a histogram-based method. |

|
Section 4: Symbol's LS2208 - The Shortest Pulse Method |
Symbol's LS2208 uses the shortest pulse method to estimate the module width. The decoder captures a complete sequence of pulse widths from the barcode. It then finds the smallest pulse width in the sequence. This smallest pulse width is used as the module width. |
The LS2208's module width estimation is robust enough for most hand-scanning applications. The scanner's designers have tuned the algorithm to handle the typical variations in scanning speed. |
Section 5: The Histogram-Based Method - A Robust Approach |
The histogram-based method is a more robust approach to estimating the module width. The decoder constructs a histogram of the pulse widths. The histogram is a bar chart that shows the number of pulses of each width. The histogram will have peaks at the module width, at twice the module width, and at three or four times the module width. |
The decoder finds the first peak in the histogram. This first peak corresponds to the module width. The histogram-based method is less vulnerable to noise than the shortest pulse method. The histogram averages the pulse widths, reducing the impact of outliers. |
Honeywell uses a histogram-based method in their imaging scanners. |

|
Section 6: Honeywell's Histogram-Based Module Estimation |
Honeywell's imagers use a histogram-based method to estimate the module width. The decoder constructs a histogram of the edge-to-edge distances (the pulse widths) from the captured image. The histogram is then analyzed to find the first peak, which is the module width. |
The histogram-based method is more computationally intensive than the shortest pulse method, but it is also more accurate and robust. Honeywell's use of the histogram-based method contributes to their scanners' excellent performance on low-quality barcodes. |
Section 7: The Running-Average Method - Adapting to Speed Changes |
The running-average method is a technique for adapting to changes in scanning speed during a scan. The decoder maintains a running average of the pulse widths. The running average is updated with each new pulse. The module width is estimated from the running average. |
The running-average method is useful when the scanning speed varies significantly during a scan. The decoder can adapt to the speed changes in real-time. |
Datalogic uses a running-average method in their industrial scanners. |

|
Section 8: Datalogic's Running-Average Module Estimation |
Datalogic's industrial scanners use a running-average method to estimate the module width. The decoder maintains a running average of the pulse widths. The running average is updated with each new pulse. The module width is estimated from the running average. |
The running-average method allows Datalogic's scanners to handle the rapid speed changes that can occur on conveyor belts. |
Section 9: Normalizing the Pulse Widths - Dividing by the Module |
Once the module width has been estimated, the decoder normalizes the pulse widths by dividing each pulse width by the module width. The result is a sequence of element widths measured in module units. The element widths are integers (1, 2, 3, or 4), although they may be fractional due to measurement errors and quantization. |
The normalization is the core of run-length decoding. It transforms the raw pulse widths into a standardized format. |

|
Section 10: The Quantization - Rounding to Integers |
The normalized element widths are often fractional. The decoder rounds them to the nearest integer. The rounding is done to accommodate measurement errors and quantization errors. |
The decoder uses a tolerance for the rounding. An element width of 0.8 to 1.2 is rounded to 1. An element width of 1.8 to 2.2 is rounded to 2. The tolerance is typically 20-25%. |
Section 11: The Tolerance - Accounting for Errors |
The tolerance accounts for errors in the pulse width measurement. The errors can be caused by noise, jitter, scanning speed variations, and print quality variations. The tolerance ensures that the decoder can handle these errors. |
The tolerance is a critical parameter. A larger tolerance makes the decoder more robust to errors but can also cause misclassifications. A smaller tolerance makes the decoder less robust but more accurate. |

|
Section 12: The Element Sequence - The Barcode's Pattern |
The normalized and rounded element widths form a sequence of integers. This sequence is the barcode's pattern. The pattern consists of alternating bars and spaces. The values in the sequence represent the widths of the bars and spaces in module units. |
The sequence is the input to the symbology decoder. The symbology decoder interprets the pattern and turns it into characters. |
Section 13: The Symbology - The Barcode's Grammar |
The symbology is the barcode's grammar. It defines the rules for encoding characters into bars and spaces. Different symbologies have different rules. Code 39, UPC, Code 128, and EAN are all different symbologies. |
The symbology decoder must know the symbology to decode the barcode. The symbology is usually determined by the barcode's start and stop characters. |

|
Section 14: Code 39 - A Variable-Length Symbology |
Code 39 is a variable-length symbology. It can encode any number of characters. Each character is represented by a 9-element pattern (5 bars and 4 spaces). The pattern has 3 wide elements and 6 narrow elements. |
The Code 39 decoder must identify the start and stop characters, which are usually an asterisk (*). The decoder then groups the elements into 9-element patterns and looks up the corresponding character in a lookup table. |
Section 15: UPC - A Fixed-Length Symbology |
UPC (Universal Product Code) is a fixed-length symbology. It is used in retail. The UPC code is 12 digits long. The first 6 digits are the manufacturer code, and the next 5 digits are the product code. The last digit is a checksum. |
The UPC decoder must identify the left and right halves of the barcode and the center guard pattern. The decoder then decodes the digits using a lookup table. |

|
Section 16: Code 128 - A High-Density Symbology |
Code 128 is a high-density symbology. It can encode all 128 ASCII characters. It uses a 6-element pattern for each character. The pattern has 3 bars and 3 spaces. |
The Code 128 decoder must handle the start, stop, and checksum characters. Code 128 is a more complex symbology than Code 39 or UPC. |
Section 17: The Start and Stop Characters - The Barcode's Boundaries |
The start and stop characters are special patterns that mark the beginning and end of the barcode. They are essential for the decoder. The decoder uses the start and stop characters to locate the barcode and to determine the symbology. |
The start and stop characters are unique patterns that are not used for data. They are always at the beginning and end of the barcode. |

|
Section 18: The Quiet Zone - The White Margin |
The quiet zone is a white margin that surrounds the barcode. The quiet zone is typically at least 10 times the module width. The decoder uses the quiet zone to detect the barcode's presence and to reset its timing. |
The quiet zone is not part of the element sequence. It is used only for detection and synchronization. |
Section 19: The Element Sequence and the Distortion |
The element sequence can be distorted. The distortion can be caused by print quality variations, scanning speed variations, or noise. The decoder must be robust to distortion. |
The decoder uses tolerances and error correction to handle distortion. |

|
Section 20: The Element Sequence and the Noise |
The element sequence can be affected by noise. The noise can cause errors in the pulse width measurement. The noise can also cause false edges. |
The decoder's tolerance helps to mitigate the effects of noise. |
Section 21: The Element Sequence and the Jitter |
The element sequence can be affected by jitter. The jitter is the uncertainty in the edge timing. The jitter causes errors in the pulse width measurement. |
The decoder's tolerance helps to mitigate the effects of jitter. |

|
Section 22: The Element Sequence and the Quantization Error |
The element sequence can be affected by the quantization error. The quantization error is caused by the timer's finite resolution. The quantization error causes the pulse widths to be measured to the nearest timer count. |
The decoder's tolerance helps to mitigate the effects of the quantization error. |
Section 23: The Element Sequence and the Symbology Decoder |
The element sequence is the input to the symbology decoder. The symbology decoder interprets the sequence and turns it into characters. The symbology decoder is the heart of the decoder. |

|
Section 24: The Symbology Decoder - A Pattern Matcher |
The symbology decoder is a pattern matcher. It compares the element sequence to the patterns in a lookup table. The lookup table is specific to the symbology. |
The symbology decoder finds the matching pattern and outputs the corresponding character. |
Section 25: The Lookup Table - A Memory of Patterns |
The lookup table is a memory that stores the patterns for each character. The lookup table is specific to the symbology. The lookup table is stored in the microcontroller's program memory. |
The lookup table is a critical part of the decoder. |

|
Section 26: The Decoder's State Machine - A Sequential Processor |
The decoder is a state machine. It processes the element sequence sequentially. The state machine has states for the quiet zone, the start character, the data characters, the checksum, and the stop character. |
The state machine is implemented in the microcontroller's firmware. |
Section 27: The Decoder's Output - The Barcode's Data |
The decoder's output is the barcode's data. The data is the sequence of characters that were encoded in the barcode. The data is outputted as a string of ASCII characters. |

|
Section 28: The Checksum - Validating the Data |
The decoder calculates a checksum from the decoded data. The checksum is a mathematical function of the data. The checksum is compared to a checksum that is encoded in the barcode. If the two checksums match, the data is valid. If they do not match, the scan is rejected. |
The checksum is an essential error-checking mechanism. |
Section 29: The Run-Length Decoding in Microchip's Reference Design |
Microchip's reference design for a barcode scanner includes a complete run-length decoding example. The design uses a PIC microcontroller. The firmware includes a module width estimation algorithm and a normalization routine. |
The Microchip reference design is a useful starting point for engineers developing barcode scanners. |

|
Section 30: The Run-Length Decoding in NXP's Reference Design |
NXP's reference design includes a run-length decoding example. The design uses an LPC microcontroller with a DMA engine. The firmware includes a histogram-based module width estimation algorithm. |
The NXP reference design demonstrates the use of a histogram-based method. |
Section 31: The Run-Length Decoding in STMicroelectronics' Reference Design |
STMicroelectronics' reference design includes a run-length decoding example. The design uses an STM32 microcontroller. The firmware includes a shortest-pulse module width estimation algorithm. |
The STMicroelectronics reference design is a useful starting point for engineers developing barcode scanners. |

|
Section 32: The Run-Length Decoding and the Scanning Speed |
The run-length decoding is affected by the scanning speed. A faster scan produces a smaller module width. A slower scan produces a larger module width. The decoder must estimate the module width from the pulse widths. |
The module width estimation algorithm must handle the variations in scanning speed. |
Section 33: The Run-Length Decoding and the Print Quality |
The run-length decoding is affected by the print quality. A poorly printed barcode has a larger variation in the element widths. The decoder must be robust to the variations in print quality. |
The histogram-based method is more robust to print quality variations than the shortest pulse method. |

|
Section 34: The Run-Length Decoding and the Noise |
The run-length decoding is affected by the noise. The noise can cause errors in the pulse width measurement. The decoder's tolerance helps to mitigate the effects of noise. |
Section 35: The Run-Length Decoding and the Jitter |
The run-length decoding is affected by the jitter. The jitter causes errors in the pulse width measurement. The decoder's tolerance helps to mitigate the effects of jitter. |

|
Section 36: The Run-Length Decoding - A Summary of Best Practices |
Based on our exploration, let us summarize the best practices for run-length decoding in a barcode scanner: |
1. Estimate the Module Width: Use a robust method to estimate the module width. The histogram-based method is preferred. |
2. Normalize the Pulse Widths: Divide each pulse width by the module width to obtain the element widths in module units. |
3. Round to Integers: Round the normalized element widths to the nearest integer, using a tolerance. |
4. Handle Variations: Use a tolerance to handle scanning speed variations, print quality variations, and noise. |
5. Use a Lookup Table: Use a lookup table to map the element patterns to characters. |
6. Test the Decoder: The run-length decoder must be tested with a variety of barcodes, under a variety of conditions, to ensure it is working correctly. |

|
Final Summary |
Run-length decoding is the first and most crucial step in the decoder's algorithm. It transforms the raw pulse widths into a standardized sequence of element widths measured in module units. The module is the fundamental unit of the barcode. The decoder estimates the module width from the pulse widths and then normalizes the pulse widths by dividing by the module width. |
We have seen how major companies have implemented run-length decoding in their products. Symbol's LS2208 uses the shortest pulse method. Honeywell uses a histogram-based method. Datalogic uses a running-average method. Microchip, NXP, and STMicroelectronics provide reference designs that include complete run-length decoding examples. |
Run-length decoding is the bridge between the raw timing measurements and the symbolic representation of the barcode. It is the key to unlocking the barcode's data. |