Text

Text helpers for Brazilian names, company names and addresses.

  • Parity matrix

Capitalize

Capitalizes the first letter of each word the way a Brazilian name, company name or address is written.

For example, "jose da silva" becomes "Jose da Silva", "empresa ltda" becomes "Empresa LTDA" and "santana/rs" becomes "Santana/RS".

  • options.lowerCaseWords lists the words kept in lower case between two words. The default is the prepositions and the conjunction e: a, ao, aos, à, às, ante, após, até, com, da, das, de, do, dos, e, em, na, nas, no, nos, num, numa, o, para, pela, pelas, pelo, pelos, perante, por, sem, sob, sobre, plus the particles del, della, den, der, di, du, van and von. The articles are only a and o. Until 2.4.0 ao, aos, à, às, para, pela(s), pelo(s), sob, sobre, até, num, numa, ante, após and perante were capitalized.
  • options.upperCaseWords lists the words always in upper case. The default is the company designations (CIA, EIRELI, EPP, LTDA, ME, MEI, S.A, S.A., S.S., S/A, S/S, SCP), the document abbreviations (CEP, CNPJ, CPF, RG, UF) and the roman numerals. The match ignores case. S/A and S/S are matched across the slash even though a slash splits words ("casa de carnes s/a" becomes "Casa de Carnes S/A"). A list given replaces its default, so { upperCaseWords: [] } turns "empresa ltda" into "Empresa Ltda". A state code after a /, or ending the value after a spaced hyphen, a spaced en dash or a comma, stays upper case even with upperCaseWords given.
  • Words split at whitespace, -, /, apostrophes and adjoining punctuation. Runs of whitespace (spaces, tabs, newlines) collapse into one space, and the whitespace at the start and at the end is dropped.
  • A lower-case word keeps its capital when it is first, last or followed by punctuation. ME is upper-cased only as a designation (last word, or right before another company designation such as EPP or S/A), never when a hyphen or an apostrophe attaches it to the previous word ("diga-me" becomes "Diga-Me"). S.A typed without the final dot is a designation too ("empresa s.a" becomes "Empresa S.A"; until 2.4.0 it was "Empresa S.a"). SA without dots is left alone (the surname Sá).
  • The elided particle d is lower case wherever it appears, even as the first word, but only when an apostrophe and a word follow it ("santa bárbara d'oeste" becomes "Santa Bárbara d'Oeste"; "rua d" becomes "Rua D"). A single letter right after an apostrophe is the English possessive and stays lower case ("bob's" becomes "Bob's").
  • Roman numerals from II to XXXIX are upper-cased ("rua xxiv de maio" becomes "Rua XXIV de Maio"). VI is left out because it is also the verb form "vi", and numerals from XL on are left out because letters such as L, C and D spell ordinary words. 2.4.0 stopped at XXIII.
  • A state code (UF) is upper-cased after a / ("santana/rs" becomes "Santana/RS"). As the last word, it is also upper-cased after a spaced hyphen, a spaced en dash or a comma, the "Cidade - UF" form of the Correios addressing guide: "brasília - df" becomes "Brasília - DF". Anywhere else, or after an unspaced hyphen ("brasília-df"), the two letters are an ordinary word. 2.4.0 upper-cased a state code only after /.
  • Only a real state code counts: "santana/br" becomes "Santana/Br".
  • A first letter whose upper case is two letters (ß) keeps its case: "straße" becomes "Straße" and "ßa" stays "ßa". The letters after the first are lowered one by one, so "İSTANBUL" becomes "İstanbul".
  • A lowerCaseWords or upperCaseWords that is not an array falls back to its default, and a member that is not a string is ignored. A value that is not a string returns an empty string.
ParameterTypeRequired
valuestringyes
optionsCapitalizeOptionsno
options.lowerCaseWordsstring[]no
options.upperCaseWordsstring[]no
returnsstring

Capitalize the first letter of each word, the way a Brazilian name, company name or address is written, with no options needed.

  • Options (CapitalizeOptions): lowerCaseWords, words kept in lower case between two words, by default the prepositions and the conjunction e, such as de, da, do, ao, para, pelo, sobre, até (the articles are only a and o); upperCaseWords, words always in upper case, by default company designations and abbreviations such as LTDA, S.A., ME, CNPJ and roman numerals. A list replaces its default.
  • Words split at whitespace, -, /, apostrophes and adjoining punctuation; whitespace runs collapse into one space.
  • A lower-case word that is first, last or followed by punctuation is a designator and keeps its capital.
  • ME is upper-cased only as a designation (last word, or before another designation); S.A without the final dot is a designation too, while SA without dots is left alone (the surname Sá). A state code after a / is upper-cased even with upperCaseWords given, and so is one that ends the value after a spaced -, a spaced – or a , , the Correios' "Cidade – UF".
import { capitalize } from '@brazilian-utils/brazilian-utils';

capitalize('jose da silva'); // Jose da Silva
capitalize('JOSÉ DA SILVA'); // José da Silva
capitalize('empresa ltda'); // Empresa LTDA
capitalize('banco do brasil s.a.'); // Banco do Brasil S.A.
capitalize('casa de carnes s/a'); // Casa de Carnes S/A ("S/A" is matched across the slash)
capitalize('mogi-guaçu'); // Mogi-Guaçu ("-" starts a new word)
capitalize("santa bárbara d'oeste"); // Santa Bárbara d'Oeste ("'" starts a new word, "d" stays lower case)
capitalize("bob's"); // Bob's (a single letter after an apostrophe is the English possessive)
capitalize('rua a, 100'); // Rua A, 100 (a preposition followed by punctuation is a designator)
capitalize('fulano comércio me'); // Fulano Comércio ME ("ME" as the last word is the designation)
capitalize('não-me-toque'); // Não-Me-Toque (anywhere else "me" is an ordinary word)
capitalize('(empresa) ltda'); // (Empresa) LTDA
capitalize('luiz von schmidt'); // Luiz von Schmidt
capitalize('casa para todos'); // Casa para Todos (contracted prepositions such as "ao", "às", "pelo" and "sobre" stay lower case too)
capitalize('empresa s.a'); // Empresa S.A
capitalize('santana/rs'); // Santana/RS ("RS" is a state code right after a "/")
capitalize('porto alegre/rs'); // Porto Alegre/RS
capitalize('brasília - df'); // Brasília - DF (a state code as the last word after " - ", " – " or ", ")
capitalize('santana rs'); // Santana Rs (no "/", so "rs" is just a word)
capitalize('rua xv de novembro'); // Rua XV de Novembro (roman numeral, "de" stays lower case)
capitalize('joão paulo ii'); // João Paulo II
capitalize('rua xxiv de maio'); // Rua XXIV de Maio (roman numerals from II to XXXIX, except VI, the verb "vi")
capitalize('de'); // De (a preposition keeps its capital when it is the first word)
capitalize('empresa ltda', { upperCaseWords: [] }); // Empresa Ltda (the list given replaces the default one)
capitalize('josé Ama MARIA', { lowerCaseWords: ['ama'] }); // José ama Maria
capitalize('doc inválido', { upperCaseWords: ['DOC'] }); // DOC Inválido (case-insensitive match)
capitalize('  josé   maria  '); // José Maria (every run of whitespace, tabs and newlines included, collapses into one space)

Source: Manual de Redação da Presidência da República.

Code: brazilian-utils/javascript
Try it with JavaScript capitalize
The inputs start with the first shared case. Change one to see the new result.

Runs @brazilian-utils/brazilian-utils 2.5.0 in your browser.

Shared test cases (162) and the result in each library text.capitalize

Remove accents

Removes diacritical marks (accents, tildes, cedillas) from a string.

  • The function decomposes every character (Unicode NFD) and drops every combining mark (general category M), so accents from any script go.
  • Letters with no decomposition into a base letter and a mark stay as they are (ß, ø, æ). A value that is not a string returns an empty string.
ParameterTypeRequired
valuestringyes
returnsstring

Remove diacritical marks (accents, tildes, cedillas) from a string.

  • Every combining mark (Unicode general category M) is dropped, so accents from any script go.
import { removeAccents } from '@brazilian-utils/brazilian-utils';

removeAccents('São Paulo'); // 'Sao Paulo'
removeAccents('Piauí'); // 'Piaui'
removeAccents('Ceará'); // 'Ceara'
removeAccents('Açaí'); // 'Acai'
removeAccents(''); // ''
Code: brazilian-utils/javascript
Try it with JavaScript removeAccents
The inputs start with the first shared case. Change one to see the new result.

Runs @brazilian-utils/brazilian-utils 2.5.0 in your browser.

Shared test cases (11) and the result in each library text.removeAccents

Official sources

Last updated on

On this page