Compare commits

..
7 Commits
12 changed files with 243 additions and 353 deletions
+42 -10
View File
@@ -1,14 +1,44 @@
# Financials-Extension
Version 3.3.0 includes improved cookie handling and somewhat improved logic to deal with network issues.
## Overview
This is a Python based extension for LibreOffice Calc to make market data available in Calc
spreadsheets - currently supporting Yahoo's (FX, crypto, equities, indices, futures, options) and Financial Times'
(FX, equities, indices, futures) websites using old-fashioned web scraping.
Starting with version 3.1.0, we received a contribution to get crypto data directly from Coinbase
## Latest version vs Yahoo HTTPS fingerprinting
Latest version 3.8.0 was created to bypass Yahoo's recently adding crazy HTTPS fingerprinting
to their website. If you don't use Yahoo, no further changes are required.
If you do use Yahoo as a source, here is what I had to do to get Yahoo working again on my Ubuntu system:
- install system-wide Python module curl_cffi - on my system as root: `pip3 install curl_cffi --upgrade`
- download latest binary of [curl-impersonate](https://github.com/lwthiker/curl-impersonate/releases) e.g.
libcurl-impersonate-v0.6.1.x86_64-linux-gnu.tar.gz and untar it somewhere
Now some of these bits need to be loaded/initialised before running LibreOffice: I used the below (adjust your location
to libcurl-impersonate-chrome.so) to run LibreOffice Calc directly from command line:
```
LD_PRELOAD=/tmp/curl-impersonate/libcurl-impersonate-chrome.so CURL_IMPERSONATE=chrome101 /usr/lib/libreoffice/program/soffice.bin --calc
```
With this I can see the below in the output from `=GETREALTIME("SUPPORT")` and the examples.ods file from this repo
can load data for Yahoo again.
```
...
requests=curl_cffi_0.10.0
LD_PRELOAD=/tmp/curl-impersonate/libcurl-impersonate-chrome.so
CURL_IMPERSONATE=chrome101
curl_version="libcurl/8.1.1 BoringSSL zlib/1.2.11 brotli/1.0.9 nghttp2/1.56.0"
```
Similar things should be possible on Windows - let me know if [this](https://stackoverflow.com/questions/1178257/ld-preload-equivalent-for-windows-to-preload-shared-libraries)
is helpful and share your experience.
Background: for a normal Python script just installing curl_cffi is enough to bypass Yahoo's HTTPS fingerprinting.
Because LibreOffice is loading the stock curl library before executing the extension code directly, the above hack
is required. Unless someone tells me otherwise...
### Feedback requested:
@@ -31,7 +61,8 @@ Getting data should be as simple as having this in a cell:
Codes 21 and 90 stand for "last price" and "close" (see below), respectively.
Only Yahoo has historic data available.
There is a file **examples.ods** there too with usage examples and possible arguments to functions.
There is a file **examples.ods** in the Release area too with usage examples
and possible arguments to functions.
You have to check the respective websites to work out what symbol is the right one for you. Make sure today or the date
requested is a trading day (exchange is not closed). If a website doesn't have
@@ -129,18 +160,19 @@ On my system (Ubuntu) I installed packages: libreoffice-dev libreoffice-java-com
cd ~/tech/IdeaProjects/Financials-Extension/
python3 -m unittest discover src
\# Assuming curl-cffi is installed, LD_PRELOAD is not required here
CURL_IMPERSONATE=chrome101 python3 -m unittest discover src
\# This builds file **Financials-Extension.oxt**
./compile.sh
### Tested with:
- Windows 10 / LibreOffice Calc 7.1.2.2 / Python 3.8.8
- Ubuntu 22.04.1 / LibreOffice Calc 7.3.7.2 / Python 3.10.6
- MacOS 10.15.7 / LibreOffice Calc 7.2.0.4 / Python 3.8.10
- Ubuntu 22.04.5 / LibreOffice Calc 7.3.7.2 / Python 3.10.12
(Previous versions)
(Previously)
- Windows 10 / LibreOffice Calc 7.1.2.2 / Python 3.8.8
- MacOS 10.15.7 / LibreOffice Calc 7.2.0.4 / Python 3.8.10
- Debian 10.3 / LibreOffice Calc 6.1.5.2 / Python 3.7.3
- Ubuntu 20.04.5 / LibreOffice Calc 6.4.7.2 / Python 3.8.10
- Ubuntu 18.04.5 / LibreOffice Calc 6 / Python 3.6.9
+3 -4
View File
@@ -53,7 +53,6 @@ python3 "${PWD}"/src/generate_metainfo.py
cp -f "${PWD}"/src/financials.py "${PWD}"/build/
cp -f "${PWD}"/src/datacode.py "${PWD}"/build/
cp -f "${PWD}"/src/baseclient.py "${PWD}"/build/
cp -f "${PWD}"/src/jsonParser.py "${PWD}"/build/
cp -f "${PWD}"/src/naivehtmlparser.py "${PWD}"/build/
cp -f "${PWD}"/src/tz.py "${PWD}"/build/
cp -f "${PWD}"/src/financials_ft.py "${PWD}"/build/
@@ -64,11 +63,11 @@ cp -f "${PWD}"/src/financials_coinbase.py "${PWD}"/build/
TMPFILE=`mktemp`
wget "https://files.pythonhosted.org/packages/36/7a/87837f39d0296e723bb9b62bbb257d0355c7f6128853c78955f57342a56d/python_dateutil-2.8.2-py2.py3-none-any.whl" -O $TMPFILE
wget "https://files.pythonhosted.org/packages/ec/57/56b9bcc3c9c6a792fcbaf139543cee77261f3651ca9da0c93f5c1221264b/python_dateutil-2.9.0.post0-py2.py3-none-any.whl" -O $TMPFILE
unzip $TMPFILE dateutil/\* -d "${PWD}"/build/
rm $TMPFILE
wget "https://files.pythonhosted.org/packages/7f/99/ad6bd37e748257dd70d6f85d916cafe79c0b0f5e2e95b11f7fbc82bf3110/pytz-2023.3-py2.py3-none-any.whl" -O $TMPFILE
wget "https://files.pythonhosted.org/packages/81/c4/34e93fe5f5429d7570ec1fa436f1986fb1f00c3e0f43a589fe2bbcd22c3f/pytz-2025.2-py2.py3-none-any.whl" -O $TMPFILE
unzip $TMPFILE pytz/\* -d "${PWD}"/build/
rm $TMPFILE
@@ -77,7 +76,7 @@ unzip $TMPFILE pyparsing.py -d "${PWD}"/build/
rm $TMPFILE
# Windows LibreOffice 7.1 Python is missing this...
wget "https://files.pythonhosted.org/packages/d9/5a/e7c31adbe875f2abbb91bd84cf2dc52d792b5a01506781dbcf25c91daf11/six-1.16.0-py2.py3-none-any.whl" -O $TMPFILE
wget "https://files.pythonhosted.org/packages/b7/ce/149a00dd41f10bc29e5921b496af8b574d8413afcd5e30dfa0ed46c2cc5e/six-1.17.0-py2.py3-none-any.whl" -O $TMPFILE
unzip $TMPFILE six.py -d "${PWD}"/build/
rm $TMPFILE
BIN
View File
Binary file not shown.
+70 -143
View File
@@ -8,17 +8,11 @@
# version 3 of the License, or (at your option) any later version.
import codecs
import gzip
import logging
import os
import pathlib
import random
import select
import urllib.request
from http import cookiejar
from http.client import HTTPConnection, HTTPSConnection, HTTPException
from importlib import util
from datacode import Datacode
logger = logging.getLogger(__name__)
@@ -27,160 +21,92 @@ logger = logging.getLogger(__name__)
# logger.setLevel(logging.DEBUG)
class RedirectException(HTTPException):
def __init__(self, location):
self.location = location
curl_cffi_present = not util.find_spec("curl_cffi") is None
requests_present = not util.find_spec("requests") is None
if curl_cffi_present:
logger.debug("Importing curl_cffi...")
from curl_cffi import requests, __version__ as requests_version, __name__ as requests_name
elif requests_present:
logger.debug("Importing requests...")
import requests
requests_version = requests.__version__
requests_name = requests.__name__
else:
raise Exception("Neither curl_cffi nor requests found.")
# import requests
class HttpException(HTTPException):
def __init__(self, url, status):
class HttpException(Exception):
def __init__(self, url, response):
self.url = url
self.status = status
self.response = response
def __str__(self):
if self.response is None:
return f"url='{self.url}'"
if type(self.response) is str:
return f"url='{self.url}' status='{self.response}'"
if self.response.headers:
h = '\n'.join(sorted(self.response.headers.__str__().splitlines(), key=lambda l: l.lower()))
return f"url='{self.url}' status={self.response.status_code} reason='{self.response.reason}' headers={h}\n"
else:
return f"url='{self.url}' status={self.response.status_code} reason='{self.response.reason}'"
class BaseClient:
def __init__(self):
self.connections = {}
self.cookies = cookiejar.CookieJar()
self.last_url = None
self.redirect_count = 0 # will be set later
self.redirect_count = 0
self.basedir = os.path.join(str(pathlib.Path.home()), '.financials-extension')
os.makedirs(self.basedir, exist_ok=True)
user_agents = [
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/110.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/111.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/112.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/113.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/114.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/115.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/116.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/117.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/118.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/119.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/120.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/121.0'
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:133.0) Gecko/20100101 Firefox/133.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:134.0) Gecko/20100101 Firefox/134.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:135.0) Gecko/20100101 Firefox/135.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:136.0) Gecko/20100101 Firefox/136.0',
'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:137.0) Gecko/20100101 Firefox/137.0',
]
self.default_headers = {
'User-Agent': random.sample(user_agents, 1)[0],
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
'Accept-Encoding': 'gzip, deflate',
'Accept-Language': 'en-US,en;q=0.5',
'Connection': 'keep-alive',
'Cache-Control': 'max-age=0'
}
if curl_cffi_present:
self.session = requests.Session()
if logger.isEnabledFor(logging.DEBUG) and self.session.curl:
self.session.curl.debug()
else:
self.session = requests.Session()
self.session.headers.update({'User-Agent': random.sample(user_agents, 1)[0],
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
'Accept-Encoding': 'gzip, deflate',
'Accept-Language': 'en-US,en;q=0.5',
'Connection': 'keep-alive',
'Cache-Control': 'max-age=0',
})
self.response = None
self.session.max_redirects = 5
def request(self, method: str, url: str, data=None, headers={}, **kwargs):
_headers = self.default_headers.copy()
if headers:
for key, value in headers.items():
_headers[key] = value
if method == 'POST' and 'Content-Type' not in _headers:
_headers['Content-Type'] = 'application/x-www-form-urlencoded'
connection = None
scheme, _, host, path = url.split('/', 3)
if (scheme, host) in self.connections:
connection = self.connections.get((scheme, host))
if connection and select.select([connection.sock], [], [], 0)[0]:
connection.close()
connection = None
if not connection:
logger.debug('Creating connection --------------------------------------------------')
connection = HTTPConnection(host, **kwargs) if scheme == 'http:' else HTTPSConnection(host, **kwargs)
logger.debug('Creating request -----------------------------------------------------')
logger.debug("%s %s", method, url)
self.last_url = url
# generate and add cookie headers
request = urllib.request.Request(url)
self.cookies.add_cookie_header(request)
if request.get_header('Cookie'):
_headers['Cookie'] = request.get_header('Cookie')
for key, value in _headers.items():
logger.debug('Header: %s=%s', key, value)
# request
connection.request(method, '/' + path, data, _headers)
response = connection.getresponse()
logger.debug('Processing response --------------------------------------------------')
logger.debug('response.status=%s', response.status)
for key, value in response.getheaders():
logger.debug('Header: %s=%s', key, value)
self.cookies.extract_cookies(response, request)
self.connections[(scheme, host)] = connection
return response
def urlopen(self, url, redirect=True, data=None, headers={}, cookies=[], **kwargs):
if cookies:
for c in cookies:
self.cookies.set_cookie(c)
def urlopen(self, url, data=None):
self.last_url = None
self.response = self.request('POST' if data else 'GET', url, data, headers, **kwargs)
text = self.response.read()
resp = self.session.request('POST' if data else 'GET', url, data=data)
# Allow redirects - used by Yahoo for some cookie based consent
self.redirect_count = 5
if 400 <= resp.status_code < 500:
if resp.headers.get('X-Cache') == 'Error from cloudfront':
resp = self.session.request('POST' if data else 'GET', url, data=data)
# (for Yahoo) AWS CloudFront occasionally returns an incorrect, cached error responses
# try mitigating by re-requesting straight away
if 400 <= self.response.status < 500:
if self.response.getheader('X-Cache') == 'Error from cloudfront':
self.response = self.request('POST' if data else 'GET', url, data, headers, **kwargs)
text = self.response.read()
if resp.status_code >= 400:
logger.warning("url='%s' status=%s reason='%s' headers=%s", resp.url,
resp.status_code, resp.reason,
'\n'.join(sorted(resp.headers.__str__().splitlines(), key=lambda l: l.lower())))
raise HttpException(url, resp)
while 300 <= self.response.status < 400 and self.redirect_count >= 0:
self.redirect_count = len(resp.history)
self.last_url = resp.url
self.redirect_count -= 1
location = self.response.getheader('Location')
if location and redirect:
if location.startswith('/'):
scheme, _, host, path = url.split('/', 3)
location = '{}//{}{}'.format(scheme, host, location)
self.response = self.request('GET', location, None, headers, **kwargs)
text = self.response.read()
else:
raise RedirectException(location)
if self.response.status >= 400:
logger.warning("last_url='%s' status=%s headers=%s", self.last_url, self.response.status,
'\n'.join(sorted(self.response.headers.__str__().splitlines(), key=lambda l: l.lower())))
raise HttpException(url, self.response.status)
if self.response.getheader('Content-Encoding') == 'gzip':
text = gzip.decompress(text)
content_type = self.response.headers.get_content_charset()
if content_type is None:
content_type = 'utf-8'
text = codecs.decode(text, encoding=content_type, errors='ignore')
return text
return resp.text
def get_ticker(self):
@@ -389,10 +315,11 @@ class BaseClient:
return None
def version(self):
return requests_name + "_" + requests_version
def curl(self):
return curl_version
def close(self):
for connection in self.connections.values():
try:
connection.close()
except BaseException:
pass
self.connections = {}
self.session.close()
+14 -1
View File
@@ -240,7 +240,8 @@ class FinancialsImpl(unohelper.Base, Financials):
if e.tag.endswith('version'):
version = e.attrib['value']
s = 'ctx={}\nid(self)={}\nversion={}\nfile={}\ncwd={}\nhome={}\nuname={}\npid={}\nsys.executable={}\nsys.version={}\nsys.path={}\nlocale={}\ndefaultlocale={}\ndateutil={}\npytz={}\npyparsing={}\nsix={}'.format(
s = ('ctx={}\nid(self)={}\nversion={}\nfile={}\ncwd={}\nhome={}\nuname={}\npid={}\nsys.executable={}\nsys.version={}\nsys.path={}\n' +
'locale={}\ndefaultlocale={}\ndateutil={}\npytz={}\npyparsing={}\nsix={}\nrequests={}').format(
self.ctx,
id(self),
version,
@@ -258,8 +259,20 @@ class FinancialsImpl(unohelper.Base, Financials):
pytz.__version__,
pyparsing.__version__,
six.__version__,
self.ft.version()
)
ld_preload = os.environ.get('LD_PRELOAD')
if ld_preload:
s += f"\nLD_PRELOAD={ld_preload}"
curl_impersonate = os.environ.get('CURL_IMPERSONATE')
if curl_impersonate:
s += f"\nCURL_IMPERSONATE={curl_impersonate}"
if 'curl_cffi' in self.ft.version():
s += f"\ncurl_version=\"{self.ft.session.curl.version().decode()}\""
if datacode:
s = '{}\ntype(datacode)={}\nstr(datacode)={}'.format(
s,
+1 -3
View File
@@ -20,7 +20,6 @@ import json
import dateutil.parser
import pytz
import jsonParser
from baseclient import BaseClient, HttpException
from datacode import Datacode
@@ -35,7 +34,6 @@ class Coinbase(BaseClient):
self.crumb = None
self.realtime = {}
self.js = jsonParser.jsonObject
def getRealtime(self, ticker, datacode):
@@ -61,7 +59,7 @@ class Coinbase(BaseClient):
url = 'https://api.exchange.coinbase.com/products/{}/stats'.format(ticker)
try:
text = self.urlopen(url, redirect=True, data=None, headers=None)
text = self.urlopen(url)
except BaseException as e:
logger.exception("BaseException ticker=%s datacode=%s last_url=%s redirect_count=%s", ticker, datacode, self.last_url, self.redirect_count)
del self.realtime[ticker]
+1 -3
View File
@@ -16,7 +16,6 @@ import urllib.parse
import dateutil.parser
import jsonParser
from baseclient import BaseClient
from datacode import Datacode
from tz import whois_timezone_info
@@ -48,7 +47,6 @@ class FT(BaseClient):
self.crumb = None
self.realtime = {}
self.historicdata = {}
self.js = jsonParser.jsonObject
def getRealtime(self, ticker: str, datacode: int):
@@ -78,7 +76,7 @@ class FT(BaseClient):
url = f'https://markets.ft.com/data/{asset_class}/tearsheet/summary?s={urllib.parse.quote_plus(ticker)}'
try:
text = self.urlopen(url, redirect=True, data=None, headers=None)
text = self.urlopen(url)
except BaseException as e:
logger.exception("BaseException ticker=%s datacode=%s last_url=%s redirect_count=%s", ticker, datacode, self.last_url, self.redirect_count)
del self.realtime[ticker]
+65 -34
View File
@@ -8,7 +8,6 @@
# version 3 of the License, or (at your option) any later version.
import csv
import datetime
import json
import logging
@@ -20,7 +19,6 @@ import urllib.parse
import dateutil.parser
import jsonParser
from baseclient import BaseClient, HttpException
from datacode import Datacode
from naivehtmlparser import NaiveHTMLParser
@@ -65,42 +63,63 @@ class Yahoo(BaseClient):
self.crumb = None
self.realtime = {}
self.historicdata = {}
self.js = jsonParser.jsonObject
def _read_ticker_csv_file(self, ticker):
def _read_ticker_json_file(self, ticker):
fn = os.path.join(self.basedir, 'yahoo-{}.csv'.format(ticker))
fn = os.path.join(self.basedir, 'yahoo-hist-{}.json'.format(ticker))
if not os.path.isfile(fn):
return
with open(fn, newline='', encoding="utf-8") as csvfile:
reader = csv.DictReader(csvfile)
with open(fn, newline='', encoding="utf-8") as jsonfile:
js = jsonfile.read()
ticks = {}
parsed = json.loads(js)
parsed = parsed['chart']['result'][0]
for row in reader:
tick = self.get_ticker()
try:
tick[Datacode.OPEN] = float(row['Open'])
tick[Datacode.LOW] = float(row['Low'])
tick[Datacode.HIGH] = float(row['High'])
tick[Datacode.VOLUME] = float(row['Volume'])
tick[Datacode.CLOSE] = float(row['Close'])
tick[Datacode.ADJ_CLOSE] = float(row['Adj Close'])
except:
pass
price_hint = 2
if 'priceHint' in parsed['meta']:
price_hint = str(parsed['meta']['priceHint'])
if price_hint and price_hint.isnumeric():
price_hint = int(price_hint)
else:
price_hint = 2
if len(tick) > 0:
ticks[row['Date']] = tick
tz = datetime.timezone(datetime.timedelta(seconds=parsed['meta']['gmtoffset']), parsed['meta']['exchangeTimezoneName'])
self.historicdata[ticker] = ticks
rows = list(
zip((datetime.datetime.fromtimestamp(ts, tz).date() for ts in parsed['timestamp']),
parsed['indicators']['quote'][0]['open'],
parsed['indicators']['quote'][0]['low'],
parsed['indicators']['quote'][0]['high'],
parsed['indicators']['quote'][0]['volume'],
parsed['indicators']['quote'][0]['close'],
parsed['indicators']['adjclose'][0]['adjclose']))
ticks = {}
for row in rows:
tick = self.get_ticker()
try:
tick[Datacode.OPEN] = round(float(row[1]), price_hint)
tick[Datacode.LOW] = round(float(row[2]), price_hint)
tick[Datacode.HIGH] = round(float(row[3]), price_hint)
tick[Datacode.VOLUME] = round(float(row[4]), price_hint)
tick[Datacode.CLOSE] = round(float(row[5]), price_hint)
tick[Datacode.ADJ_CLOSE] = round(float(row[6]), price_hint)
except:
pass
if len(tick) > 0:
ticks[str(row[0])] = tick # Date
self.historicdata[ticker] = ticks
def handleCookiesAndConsent(self, url, ticker, datacode, html_file):
try:
text = self.urlopen(url, redirect=True)
text = self.urlopen(url)
except BaseException as e:
logger.exception("BaseException (1) ticker=%s datacode=%s last_url=%s redirect_count=%s %s",
ticker, datacode, self.last_url, self.redirect_count, e)
@@ -132,10 +151,11 @@ class Yahoo(BaseClient):
data = {'reject': 'reject'}
for d in inputs:
data[d.attrib['name']] = d.attrib['value']
if 'name' in d.attrib and 'value' in d.attrib:
data[d.attrib['name']] = d.attrib['value']
try:
text = self.urlopen(self.last_url, redirect=True, data=urllib.parse.urlencode(data))
text = self.urlopen(self.last_url, data=data)
except BaseException as e:
logger.exception("BaseException (4) ticker=%s datacode=%s last_url=%s redirect_count=%s %s",
ticker, datacode, self.last_url, self.redirect_count, e)
@@ -176,19 +196,28 @@ class Yahoo(BaseClient):
if not self.crumb:
url = 'https://finance.yahoo.com/quote/{}?p={}'.format(ticker, ticker)
url = f'https://finance.yahoo.com/quote/{ticker}'
text = self.handleCookiesAndConsent(url, ticker, datacode, f'yahoo-{ticker}.html')
if text is None:
del self.realtime[ticker]
return 'Yahoo.getRealtime({}, {}) - handleCookiesAndConsent'.format(ticker, datacode)
# crumbs like 'TKkC\u002FZBwoUA' may contain unicode _text_ (not encoded code points)
try:
r = '"crumb":"([^"]{11})"'
r = r'\bcrumb=([^"]{11,})"'
pattern = re.compile(r)
match = pattern.search(text)
if match:
self.crumb = match.group(1)
self.crumb = urllib.parse.unquote(match.group(1).encode('unicode-escape').decode('ascii'))
logger.debug(f"crumb='{match.group(1)}' self.crumb='{self.crumb}'")
else:
r = r'"crumb"\s*:\s*"([^"]{11,})"'
pattern = re.compile(r)
match = pattern.search(text)
if match:
self.crumb = urllib.parse.unquote(match.group(1).encode('unicode-escape').decode('ascii'))
logger.debug(f"crumb='{match.group(1)}' self.crumb='{self.crumb}'")
except BaseException as e:
logger.exception("BaseException ticker=%s datacode=%s", ticker, datacode)
del self.realtime[ticker]
@@ -219,7 +248,10 @@ class Yahoo(BaseClient):
parsed = json.loads(js)
parsed = parsed['quoteSummary']['result'][0]
summaryDetail = dict(sorted(parsed['summaryDetail'].items()))
summaryDetail = dict()
if 'summaryDetail' in parsed:
summaryDetail = dict(sorted(parsed['summaryDetail'].items()))
price = dict(sorted(parsed['price'].items()))
if 'defaultKeyStatistics' in parsed:
@@ -343,7 +375,7 @@ class Yahoo(BaseClient):
# the moment we are asked for ADJ_CLOSE we ignore the ticker cache to refresh
if Datacode.ADJ_CLOSE != datacode and ticker not in self.historicdata:
self._read_ticker_csv_file(ticker)
self._read_ticker_json_file(ticker)
try:
date_as_dt = dateutil.parser.parse(date, yearfirst=True, dayfirst=False)
@@ -397,16 +429,16 @@ class Yahoo(BaseClient):
try:
url = 'https://query1.finance.yahoo.com/v7/finance/download/{}' \
url = 'https://query1.finance.yahoo.com/v8/finance/chart/{}' \
'?period1={}&period2={}&interval=1d&events=history&crumb={}' \
.format(ticker, t1, t2, urllib.parse.quote_plus(self.crumb))
text = self.urlopen(url)
with open(os.path.join(self.basedir, 'yahoo-{}.csv'.format(ticker)), "w", encoding="utf-8") as csv_file:
with open(os.path.join(self.basedir, 'yahoo-hist-{}.json'.format(ticker)), "w", encoding="utf-8") as csv_file:
print(text, file=csv_file)
self._read_ticker_csv_file(ticker)
self._read_ticker_json_file(ticker)
except HttpException:
logger.exception("HttpException ticker=%s datacode=%s date=%s", ticker, datacode, date)
@@ -436,6 +468,5 @@ class Yahoo(BaseClient):
return None
def createInstance(ctx):
return Yahoo(ctx)
+1 -1
View File
@@ -14,7 +14,7 @@ import os
cur_dir = os.getcwd()
addin_id = "com.financials.getinfo"
addin_version = "3.6.0"
addin_version = "3.8.0"
addin_displayname = "Financial Market Extension"
addin_publisher_link = "https://github.com/cmallwitz/Financials-Extension"
addin_publisher_name = "The Publisher"
-115
View File
@@ -1,115 +0,0 @@
# jsonParser.py
#
# Implementation of a simple JSON parser, returning a hierarchical
# ParseResults object support both list- and dict-style data access.
#
# Copyright 2006, by Paul McGuire
#
# Updated 8 Jan 2007 - fixed dict grouping bug, and made elements and
# members optional in array and object collections
#
# Updated 9 Aug 2016 - use more current pyparsing constructs/idioms
#
# https://github.com/pyparsing/pyparsing/blob/master/examples/jsonParser.py - revision 53d1b4a on 1 Nov 2019
json_bnf = """
object
{ members }
{}
members
string : value
members , string : value
array
[ elements ]
[]
elements
value
elements , value
value
string
number
object
array
true
false
null
"""
import pyparsing as pp
from pyparsing import pyparsing_common as ppc
def make_keyword(kwd_str, kwd_value):
return pp.Keyword(kwd_str).setParseAction(pp.replaceWith(kwd_value))
TRUE = make_keyword("true", True)
FALSE = make_keyword("false", False)
NULL = make_keyword("null", None)
LBRACK, RBRACK, LBRACE, RBRACE, COLON = map(pp.Suppress, "[]{}:")
jsonString = pp.dblQuotedString().setParseAction(pp.removeQuotes)
jsonNumber = ppc.number()
jsonObject = pp.Forward()
jsonValue = pp.Forward()
jsonElements = pp.delimitedList(jsonValue)
jsonArray = pp.Group(LBRACK + pp.Optional(jsonElements, []) + RBRACK)
jsonValue << (
jsonString | jsonNumber | pp.Group(jsonObject) | jsonArray | TRUE | FALSE | NULL
)
memberDef = pp.Group(jsonString + COLON + jsonValue)
jsonMembers = pp.delimitedList(memberDef)
jsonObject << pp.Dict(LBRACE + pp.Optional(jsonMembers) + RBRACE)
jsonComment = pp.cppStyleComment
jsonObject.ignore(jsonComment)
if __name__ == "__main__":
testdata = """
{
"glossary": {
"title": "example glossary",
"GlossDiv": {
"title": "S",
"GlossList":
{
"ID": "SGML",
"SortAs": "SGML",
"GlossTerm": "Standard Generalized Markup Language",
"TrueValue": true,
"FalseValue": false,
"Gravity": -9.8,
"LargestPrimeLessThan100": 97,
"AvogadroNumber": 6.02E23,
"EvenPrimesGreaterThan2": null,
"PrimesLessThan10" : [2,3,5,7],
"Acronym": "SGML",
"Abbrev": "ISO 8879:1986",
"GlossDef": "A meta-markup language, used to create markup languages such as DocBook.",
"GlossSeeAlso": ["GML", "XML", "markup"],
"EmptyDict" : {},
"EmptyList" : []
}
}
}
}
"""
results = jsonObject.parseString(testdata)
results.pprint()
print()
def testPrint(x):
print(type(x), repr(x))
print(list(results.glossary.GlossDiv.GlossList.keys()))
testPrint(results.glossary.title)
testPrint(results.glossary.GlossDiv.GlossList.ID)
testPrint(results.glossary.GlossDiv.GlossList.FalseValue)
testPrint(results.glossary.GlossDiv.GlossList.Acronym)
testPrint(results.glossary.GlossDiv.GlossList.EvenPrimesGreaterThan2)
testPrint(results.glossary.GlossDiv.GlossList.PrimesLessThan10)
+17 -10
View File
@@ -142,31 +142,31 @@ class Test(unittest.TestCase):
def test_US_futures(self):
# https://markets.ft.com/data/commodities/tearsheet/summary?s=775326843 ESH25:IOM
# https://markets.ft.com/data/commodities/tearsheet/summary?s=823439664 ESH26:IOM - EMINI S&P MAR26
s = financials.getRealtime('775326843', Datacode.NAME.value, 'FT')
s = financials.getRealtime('823439664', Datacode.NAME.value, 'FT')
self.assertEqual(str, type(s), 'test_realtime_US_futures NAME {}'.format(s))
self.assertEqual('EMINI S&P MAR25', s, 'test_US_futures NAME {}'.format(s))
self.assertEqual('EMINI S&P MAR26', s, 'test_US_futures NAME {}'.format(s))
s = financials.getRealtime('775326843', Datacode.LAST_PRICE.value, 'FT')
s = financials.getRealtime('823439664', Datacode.LAST_PRICE.value, 'FT')
self.assertEqual(float, type(s), 'test_US_futures LAST_PRICE {}'.format(s))
# s = financials.getRealtime('775326843', Datacode.OPEN.value, 'FT')
# self.assertEqual(float, type(s), 'test_US_futures OPEN {}'.format(s))
s = financials.getRealtime('775326843', Datacode.VOLUME.value, 'FT')
s = financials.getRealtime('823439664', Datacode.VOLUME.value, 'FT')
self.assertEqual(float, type(s), 'test_US_futures VOLUME {}'.format(s))
s = financials.getRealtime('775326843', Datacode.LOW_52_WEEK.value, 'FT')
s = financials.getRealtime('823439664', Datacode.LOW_52_WEEK.value, 'FT')
self.assertEqual(float, type(s), 'test_US_futures LOW_52_WEEK {}'.format(s))
s = financials.getRealtime('775326843', Datacode.HIGH_52_WEEK.value, 'FT')
s = financials.getRealtime('823439664', Datacode.HIGH_52_WEEK.value, 'FT')
self.assertEqual(float, type(s), 'test_US_futures HIGH_52_WEEK {}'.format(s))
s = financials.getRealtime('775326843', Datacode.CHANGE.value, 'FT')
s = financials.getRealtime('823439664', Datacode.CHANGE.value, 'FT')
self.assertEqual(float, type(s), 'test_US_futures CHANGE {}'.format(s))
s = financials.getRealtime('775326843', Datacode.CHANGE_IN_PERCENT.value, 'FT')
s = financials.getRealtime('823439664', Datacode.CHANGE_IN_PERCENT.value, 'FT')
self.assertEqual(float, type(s), 'test_US_futures CHANGE_IN_PERCENT {}'.format(s))
def test_UK_ETF(self):
@@ -275,7 +275,7 @@ class Test(unittest.TestCase):
self.assertTrue(testutils.is_date(s), 'test_DE_equity EX_DIV_DATE {}'.format(s))
s = financials.getRealtime('ISHAX:GER', 'NAME', 'FT')
self.assertEqual('INTERSHOP Communications AG', s, 'test_DE_equity NAME {}'.format(s))
self.assertEqual('Intershop Communications AG', s, 'test_DE_equity NAME {}'.format(s))
s = financials.getRealtime('ISHAX:GER', 'BETA', 'FT')
self.assertTrue(testutils.is_positive_float(s), 'test_DE_equity BETA {}'.format(s))
@@ -309,6 +309,13 @@ class Test(unittest.TestCase):
self.assertEqual(str, type(s), 'test_DK_equity INDUSTRY {}'.format(s))
self.assertEqual('Pharmaceuticals and Biotechnology', s, 'test_DK_equity INDUSTRY {}'.format(s))
def test_SE_equity(self):
s = financials.getRealtime('ACRI A:STO', 'name', 'FT')
self.assertEqual('Acrinova AB (publ)', s, 'test_SE_equity NAME {}'.format(s))
s = financials.getRealtime('SE0015660014', 'name', 'FT')
self.assertEqual('Acrinova AB (publ)', s, 'test_SE_equity NAME {}'.format(s))
def test_TY_equity(self):
s = financials.getRealtime('6503:TYO', 'OPEN', 'FT')
self.assertEqual(float, type(s), 'test_TY_equity OPEN {}'.format(s))
+29 -29
View File
@@ -24,7 +24,7 @@ import testutils
financials = financials.createInstance(None)
def urlopen_fail(self, url, redirect=True, data=None, headers={}, cookies=[], **kwargs):
def urlopen_fail(self, url, data=None):
raise baseclient.HttpException(url, 'ERROR: simulated urlopen() failed')
@@ -183,53 +183,53 @@ class Test(unittest.TestCase):
# symbol from https://finance.yahoo.com/quote/IBM/options?p=IBM
s = financials.getRealtime('IBM250117C00165000', Datacode.PREV_CLOSE.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.PREV_CLOSE.value, 'YAHOO')
self.assertEqual(float, type(s), 'test_realtime_US_options PREV_CLOSE {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.NAME.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.NAME.value, 'YAHOO')
self.assertEqual(str, type(s), 'test_realtime_US_options NAME {}'.format(s))
self.assertEqual('IBM Jan 2025 165.000 call', s, 'test_realtime_US_options NAME {}'.format(s))
self.assertEqual('IBM Jan 2026 230.000 call', s, 'test_realtime_US_options NAME {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.EXPIRY_DATE.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.EXPIRY_DATE.value, 'YAHOO')
self.assertEqual(str, type(s), 'test_realtime_US_options EXPIRY_DATE {}'.format(s))
self.assertTrue(testutils.is_date(s), 'test_realtime_US_options EXPIRY_DATE {}'.format(s))
self.assertEqual("2025-01-17", s, 'test_realtime_US_options EXPIRY_DATE {}'.format(s))
self.assertEqual("2026-01-16", s, 'test_realtime_US_options EXPIRY_DATE {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.LAST_PRICE.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.LAST_PRICE.value, 'YAHOO')
self.assertEqual(float, type(s), 'test_realtime_US_options LAST_PRICE {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.OPEN.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.OPEN.value, 'YAHOO')
self.assertEqual(float, type(s), 'test_realtime_US_options OPEN {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.VOLUME.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.VOLUME.value, 'YAHOO')
self.assertEqual(float, type(s), 'test_realtime_US_options VOLUME {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.BID.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.BID.value, 'YAHOO')
self.assertEqual(float, type(s), 'test_realtime_US_options BID {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.ASK.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.ASK.value, 'YAHOO')
self.assertEqual(float, type(s), 'test_realtime_US_options ASK {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.PAYOUT_RATIO.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.PAYOUT_RATIO.value, 'YAHOO')
self.assertIsNone(s, 'test_realtime_US_options PAYOUT_RATIO {}'.format(s))
s = financials.getRealtime('IBM250117C00165000', Datacode.SECTOR.value, 'YAHOO')
s = financials.getRealtime('IBM260116C00230000', Datacode.SECTOR.value, 'YAHOO')
self.assertIsNone(s, 'test_realtime_US_options SECTOR {}'.format(s))
def test_realtime_US_futures(self):
s = financials.getRealtime('ES=F', Datacode.NAME.value, 'YAHOO')
self.assertEqual(str, type(s), 'test_realtime_US_futures NAME {}'.format(s))
self.assertEqual('E-Mini S&P 500 Jun 24', s, 'test_realtime_US_futures NAME {}'.format(s))
self.assertEqual('E-Mini S&P 500 Jun 25', s, 'test_realtime_US_futures NAME {}'.format(s))
s = financials.getRealtime('ES=F', Datacode.TICKER.value, 'YAHOO')
self.assertEqual(str, type(s), 'test_realtime_US_futures TICKER {}'.format(s))
self.assertEqual('ESM24.CME', s, 'test_realtime_US_futures TICKER {}'.format(s))
self.assertEqual('ESM25.CME', s, 'test_realtime_US_futures TICKER {}'.format(s))
s = financials.getRealtime('ES=F', Datacode.SETTLEMENT_DATE.value, 'YAHOO')
self.assertEqual(str, type(s), 'test_realtime_US_futures SETTLEMENT_DATE {}'.format(s))
self.assertTrue(testutils.is_date(s), 'test_realtime_US_futures SETTLEMENT_DATE {}'.format(s))
self.assertEqual("2024-06-21", s, 'test_realtime_US_futures SETTLEMENT_DATE {}'.format(s))
self.assertEqual("2025-06-20", s, 'test_realtime_US_futures SETTLEMENT_DATE {}'.format(s))
s = financials.getRealtime('ES=F', Datacode.LAST_PRICE.value, 'YAHOO')
self.assertEqual(float, type(s), 'test_realtime_US_futures LAST_PRICE {}'.format(s))
@@ -353,7 +353,7 @@ class Test(unittest.TestCase):
s = financials.getRealtime('NOVO-B.CO', 'industry', 'YAHOO')
self.assertEqual(str, type(s), 'test_DK_equity INDUSTRY {}'.format(s))
self.assertEqual('Biotechnology', s, 'test_DK_equity INDUSTRY {}'.format(s))
self.assertEqual('Drug Manufacturers - General', s, 'test_DK_equity INDUSTRY {}'.format(s))
s = financials.getRealtime('MAERSK-B.CO', 'currency', 'YAHOO')
self.assertEqual('DKK', s, 'test_DK_equity CURRENCY {}'.format(s))
@@ -406,15 +406,15 @@ class Test(unittest.TestCase):
self.assertIsNone(s, 'test_historic_US_equity LAST_PRICE {}'.format(s))
s = financials.getHistoric('IBM', Datacode.CLOSE.value, '2017-01-03', 'YAHOO')
self.assertEqual(159.837479, s, 'test_historic_US_equity CLOSE {}'.format(s))
self.assertAlmostEqual(159.84, s, 2, 'test_historic_US_equity CLOSE {}'.format(s))
financials.yahoo.historicdata = {}
s = financials.getHistoric('IBM', Datacode.CLOSE.value, '2017-01-03', 'YAHOO')
self.assertEqual(159.837479, s, 'test_historic_US_equity CLOSE {}'.format(s))
self.assertAlmostEqual(159.84, s, 2, 'test_historic_US_equity CLOSE {}'.format(s))
directory = os.path.join(str(pathlib.Path.home()), '.financials-extension')
ibm = os.path.join(directory, 'yahoo-IBM.csv')
ibm = os.path.join(directory, 'yahoo-hist-IBM.json')
try:
os.unlink(ibm)
except:
@@ -423,7 +423,7 @@ class Test(unittest.TestCase):
financials.yahoo.historicdata = {}
s = financials.getHistoric('IBM', Datacode.CLOSE.value, '2017-01-03', 'YAHOO')
self.assertEqual(159.837479, s, 'test_historic_US_equity CLOSE {}'.format(s))
self.assertAlmostEqual(159.84, s, 2, 'test_historic_US_equity CLOSE {}'.format(s))
s = financials.getHistoric('IBM', Datacode.ADJ_CLOSE.value, '2017-01-03', 'YAHOO')
self.assertEqual(float, type(s), 'test_historic_US_equity ADJ_CLOSE {}'.format(s))
@@ -431,7 +431,7 @@ class Test(unittest.TestCase):
def test_historic_UK_ETF(self):
directory = os.path.join(str(pathlib.Path.home()), '.financials-extension')
verx = os.path.join(directory, 'yahoo-VERX.L.csv')
verx = os.path.join(directory, 'yahoo-hist-VERX.L.json')
try:
os.unlink(verx)
except:
@@ -443,10 +443,10 @@ class Test(unittest.TestCase):
self.assertEqual(s, 'Not a trading day \'2017-01-01\'', 'test_historic_UK_ETF LAST_PRICE {}'.format(s))
s = financials.getHistoric('VERX.L', Datacode.CLOSE.value, '2017-01-03', 'YAHOO')
self.assertEqual(s, 23.24, 'test_historic_UK_ETF CLOSE {}'.format(s))
self.assertAlmostEqual(s, 23.24, 2, 'test_historic_UK_ETF CLOSE {}'.format(s))
s = financials.getHistoric('VERX.L', Datacode.CLOSE.value, '2016-10-03', 'YAHOO')
self.assertEqual(s, 22.26, 'test_historic_UK_ETF CLOSE {}'.format(s))
self.assertAlmostEqual(s, 22.26, 2, 'test_historic_UK_ETF CLOSE {}'.format(s))
# Inception Date 2014-09-30
s = financials.getHistoric('VERX.L', Datacode.CLOSE.value, '2018-04-02', 'YAHOO')
@@ -457,13 +457,13 @@ class Test(unittest.TestCase):
self.assertEqual(s, 'Not a trading day \'2015-01-01\'', 'test_historic_UK_ETF CLOSE {}'.format(s))
s = financials.getHistoric('VERX.L', Datacode.CLOSE.value, 42738, 'YAHOO') # 2017-01-03
self.assertEqual(s, 23.24, 'test_historic_UK_ETF CLOSE {}'.format(s))
self.assertAlmostEqual(s, 23.24, 2, 'test_historic_UK_ETF CLOSE {}'.format(s))
s = financials.getHistoric('VERX.L', Datacode.CLOSE.value, 42738.0, 'YAHOO') # 2017-01-03
self.assertEqual(s, 23.24, 'test_historic_UK_ETF CLOSE {}'.format(s))
self.assertAlmostEqual(s, 23.24, 2, 'test_historic_UK_ETF CLOSE {}'.format(s))
s = financials.getHistoric('VERX.L', Datacode.CLOSE.value, 42646.0, 'YAHOO') # 2016-10-03
self.assertEqual(s, 22.26, 'test_historic_UK_ETF CLOSE {}'.format(s))
self.assertAlmostEqual(s, 22.26, 2, 'test_historic_UK_ETF CLOSE {}'.format(s))
def test_historic_DE_equity(self):
@@ -471,10 +471,10 @@ class Test(unittest.TestCase):
self.assertEqual(s, 'Not a trading day \'2017-01-01\'', 'test_historic_DE_equity LAST_PRICE {}'.format(s))
s = financials.getHistoric('SAP.DE', Datacode.CLOSE.value, '2017-01-03', 'YAHOO')
self.assertEqual(s, 82.889999, 'test_historic_DE_equity CLOSE {}'.format(s))
self.assertAlmostEqual(s, 82.89, 2, 'test_historic_DE_equity CLOSE {}'.format(s))
s = financials.getHistoric('LYY8.DE', Datacode.CLOSE.value, '2017-01-03', 'YAHOO')
self.assertEqual(s, 96.010002, 'test_historic_DE_equity CLOSE {}'.format(s))
self.assertAlmostEqual(s, 96.01, 2, 'test_historic_DE_equity CLOSE {}'.format(s))
def test_realtime_errors(self):