# Check number formatting

**URL:** <https://forum.xbench.net/t/check-number-formatting/332>\
**Category:** Technical Support\
**Created:** [February 4, 2018, 2:49pm UTC](https://forum.xbench.net/t/check-number-formatting/332 "2018-02-04T14:49:24Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![pcs](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/pcs/32/106_2.png) [@pcs](https://forum.xbench.net/u/pcs)\
**Post date:** [February 4, 2018, 2:49pm UTC](https://forum.xbench.net/t/check-number-formatting/332/1 "2018-02-04T14:49:25Z")

</div>

Good day, I am new on this forum but I have been using XBench for two years now for QA. I need help with RegEx to fix my TM.

**_That’s the story…_**

> In 2016, a long-standing client introduced a style guide for numbers/measures used in their manuals. Before then, they were using the AmE number ‘spelling’, i.e. 34,000 for 34 thousands, 12.15 for 12 units and 15 decimal points. In 2016 they decided to match some technical standard, for which the thousands separator should be a non-breaking space (i.e. 34 000) and the decimal separator should still be the point (12.15). They decided that also the translated manuals should stick to this rule, regardless of the local custom – for example, in my country we use the comma as the decimal separator, but I should stick to the point (no pun intended) for this client.
> 
> After 5 years, my TM has a mix of bad sources and bad targets due to these style changes. I would like to fix my TM so that anything I pre-translate is pre-translating according to the new style guide.

I am working with SDL Trados Studio and thus I have exported my TM from \*.sdltm to \*.tmx.  
I have loaded the \*.tmx in XBench.  
I need to check and edit those TUs that do not match the current style guide. I.e. I need to run a series of check on the target, according to these rules:

- If the number has 5 or more digits, the thousand separator has to be a non-breaking space (i.e. **34 000** , however if it has just 4 digits, no space should be used, i.e. **4000** )

- If the number has decimals, the decimal separator should be the point and not the comma (i.e. **12.15** is okay, _12,15_ is not)

- If the source contains **in.** the target should contain **in.** (with the point) as well.

- If the source contains **cu.ft.** the target should contain **cu.ft.** (with the points) as well.

- If the source contains **L** the target should contain **L** (single letter, capitalized) as well.

- Ranges should be indicated using a en dash, so **10-20** is okay, while _10-20_ is not okay. (_in this post they are displayed the same, but the en dash should look slightly longer than the regular dash_)

- If the source contains **red rose** the target should contain **rosa rossa**.

- If the target contains **km** , it should be _preceded_ by a non breaking space.

- If the target contains **ASTM** , it should be _followed_ by a non breaking space

Can you please help me translate this into RegEx? Thank you in advance!

---

<div class="post-metadata">

**Author:** ![pcondal](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/pcondal/32/305_2.png) [@pcondal](https://forum.xbench.net/u/pcondal)\
**Post date:** [February 5, 2018, 6:27pm UTC](https://forum.xbench.net/t/check-number-formatting/332/2 "2018-02-05T18:27:45Z")

</div>

The Xbench RegEx grammar is documented [here](https://docs.xbench.net/user-guide/regular-expressions/).

Xbench 3.0 also support the _.sdltm_ format as input format.

Here shows how you can approach them (or at least an approximation that could be a good starting point if I did not understand the exact requirements well).

> [@pcs](#):
>
> If the number has 5 or more digits, the thousand separator has to be a non-breaking space (i.e. 34 000, however if it has just 4 digits, no space should be used, i.e. 4000)

You can probably try to find if there are still entries of the wrong target thousands separator (I assume it was a comma in target):

- Target: `[0-9]{2,3},[0-9]{3}`
- Mode: `Regular Expressions`

For the 4 digit number, you can perhaps check if there is a 4 digit number with an old thousands separator:

- Target: `<[0-9]{1},[0-9]{3}>`
- Mode: `Regular Expressions`

> [@pcs](#):
>
> If the number has decimals, the decimal separator should be the point and not the comma (i.e. 12.15 is okay, 12,15 is not)

This one needs to rely on the source and use [PowerSearch](https://docs.xbench.net/user-guide/advanced-features/#powersearch-function). I’m assuming that source uses decimal point and that target also should use decimal point:

- Source: `"([0-9]+\.[0-9]+)=1"`
- Target: `-@1`
- Mode: `Regular Expressions`
- PowerSearch: `On`

> [@pcs](#):
>
> If the source contains **in.** the target should contain **in.** (with the point) as well.

- Source: `"<in\."`
- Target: `-"<in\."`
- Mode: `Regular Expressions`
- PowerSearch: `On`

> [@pcs](#):
>
> If the source contains **cu.ft.** the target should contain **cu.ft.** (with the points) as well.

- Source: `"<cu\.ft\."`
- Target: `-"<cu\.ft\."`
- Mode: `Regular Expressions`
- PowerSearch: `On`

> [@pcs](#):
>
> If the source contains **L** the target should contain **L** (single letter, capitalized) as well.

- Source: `"<L>"`
- Target: `-"<L>"`
- Mode: `Regular Expressions`
- PowerSearch: `On`
- Case-sensitive:: `On`

> [@pcs](#):
>
> Ranges should be indicated using a en dash, so 10-20 is okay, while 10-20 is not okay. (in this post they are displayed the same, but the en dash should look slightly longer than the regular dash)

- Target: `<[0-9]+-[0-9]+>` (use the bad dash here)
- Mode: `Regular Expressions`

> [@pcs](#):
>
> If the source contains **red rose** the target should contain **rosa rossa**.

- Source: `"red rose"`
- Target: `-"rosa rossa"`
- Mode: `Simple`
- PowerSearch: `On`

> [@pcs](#):
>
> If the target contains **km** , it should be preceded by a non breaking space.

- Target: `[^\x00a0]km>`
- Mode: `Regular Expressions`

> [@pcs](#):
>
> If the target contains **ASTM** , it should be followed by a non breaking space

- Target: `ASTM[^\x00a0]`
- Mode: `Regular Expressions`

---

<div class="post-metadata">

**Author:** ![pcs](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/pcs/32/106_2.png) [@pcs](https://forum.xbench.net/u/pcs)\
**Post date:** [February 6, 2018, 3:19am UTC](https://forum.xbench.net/t/check-number-formatting/332/3 "2018-02-06T03:19:38Z")

</div>

Hello @pcondal,

thank you **_so_** much for your help. You are a life saver! 🤩  
I will reference the Xbench RegEx grammar in the future. RegEx is not so straightforward and it looks a little intimidating, though well worth learning!

As per your suggestion, I have been searching directly the \ ***.sdltm** file, though to make edits, SDL Studio is open/used. Correct? I thought it was possible to batch/bulk edit in Xbench directly, but maybe it is just for other file formats.

Anyway, most of the strings worked 🤖, however these were not working:

> [@pcondal](#):
>
> Source: “\<cu.ft.”
> 
> Target: -“\<cu.ft.”
> 
> Mode: Regular Expressions
> 
> PowerSearch: On

This one is not working (the search fields turn red), however the previous one searching for the abbreviation of inches, works. 🙄

> [@pcondal](#):
>
> For the 4 digit number, you can perhaps check if there is a 4 digit number with an old thousands separator:
> 
> Target: \<[0-9]{1},[0-9]{3}\>
> 
> Mode: Regular Expressions

This one is not working (the search fields turn red).

Thank you again for your time and efforts! 🤭

---

<div class="post-metadata">

**Author:** ![pcondal](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/pcondal/32/305_2.png) [@pcondal](https://forum.xbench.net/u/pcondal)\
**Post date:** [February 6, 2018, 8:19am UTC](https://forum.xbench.net/t/check-number-formatting/332/4 "2018-02-06T08:19:24Z")

</div>

Xbench is a browser, not really an editor. We always try to find ways to call the home application of the format to ensure data integrity (slight changes for example in a home application update can easily corrupt data for the home application proprietary data).

However, I agree that it would be great that it was possible from Xbench to open the TM in Studio right at the segment.

I created [this idea](https://community.sdl.com/ideas/translation-productivity-ideas/i/trados-studio-ideas/add-translationunit-argument-to-command-line-argument-openfiletm-to-allow-better-integration-with-xbench) in the SDL Community site. If you vote it (and manage that other interested users vote it), and SDL eventually decides to implement it, the functionality of segment positioning will be eventually available in Xbench.

For the cu.ft. I notice you are missing the backslashes in front of the dot (a dot has an special meaning in Regex). In any case, could you provide specific source/target examples on where does it not work?

---

<div class="post-metadata">

**Author:** ![pcs](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/pcs/32/106_2.png) [@pcs](https://forum.xbench.net/u/pcs)\
**Post date:** [February 7, 2018, 12:28am UTC](https://forum.xbench.net/t/check-number-formatting/332/5 "2018-02-07T00:28:36Z")

</div>

> [@pcondal](#):
>
> For the cu.ft. I notice you are missing the backslashes in front of the dot (a dot has an special meaning in Regex). In any case, could you provide specific source/target examples on where does it not work?

I am still getting the same error. I believe it is a syntax error, but I cannot tell where the error lies.

 ![31](https://canada1.discourse-cdn.com/flex035/uploads/xbench/original/1X/2ae5068e20a1e14431e485d0a2603c3760911850.png)

---

<div class="post-metadata">

**Author:** ![pcondal](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/pcondal/32/305_2.png) [@pcondal](https://forum.xbench.net/u/pcondal)\
**Post date:** [February 7, 2018, 7:05am UTC](https://forum.xbench.net/t/check-number-formatting/332/6 "2018-02-07T07:05:15Z")

</div>

Red background in Xbench means “nothing found” in search.

Are you pressing **Ctrl+P** to search in PowerSearch mode?

Is there any segment that has “something cu.ft.” in source and does not have “something cu.ft.” in target?

---

<div class="post-metadata">

**Author:** ![Claudia\_Cappelletti](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/claudia_cappelletti/32/169_2.png) [@Claudia\_Cappelletti](https://forum.xbench.net/u/Claudia_Cappelletti)\
**Post date:** [February 7, 2018, 6:15pm UTC](https://forum.xbench.net/t/check-number-formatting/332/7 "2018-02-07T18:15:48Z")

</div>

In this case:

Source: “\<cu.ft.”

Target: -"\<cu.ft."

Mode: Regular Expressions

PowerSearch: On

I believe you need to “escape” the dots so as to tell xbench that this dot is not a regular expression. So you should insert a backslash right before each dot.

---

<div class="post-metadata">

**Author:** ![RSchiaffino](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/rschiaffino/32/15_2.png) [@RSchiaffino](https://forum.xbench.net/u/RSchiaffino)\
**Post date:** [February 7, 2018, 7:11pm UTC](https://forum.xbench.net/t/check-number-formatting/332/8 "2018-02-07T19:11:08Z")

</div>

"However, I agree that it would be great that it was possible from Xbench to open the TM in Studio right at the segment.

I created this idea3 in the SDL Community site. If you vote it (and manage that other interested users vote it), and SDL eventually decides to implement it, the functionality of segment positioning will be eventually available in Xbench."

As an alternative, while this is still not possible with sdltm memories, it should be possible to access the “offending segment” of a TM for editing by first exporting the TM to TMX format, then loading it in Xbench as “ongoing translation”.

---

<div class="post-metadata">

**Author:** ![pcs](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.xbench.net/pcs/32/106_2.png) [@pcs](https://forum.xbench.net/u/pcs)\
**Post date:** [February 9, 2018, 6:50pm UTC](https://forum.xbench.net/t/check-number-formatting/332/9 "2018-02-09T18:50:27Z")

</div>

You are right, by now I probably have solved all issues with this specific string. I will be running more QA checks later next week and I might post again! Thank you!
